system

The system uses a Generative Adversarial Network to efficiently generate and integrate 3D maps for urban digital twins, addressing the cost and time issues of traditional methods, facilitating rapid and accurate map creation for urban planning and management.

JP2026064572APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Creating 3D maps for urban digital twins is costly and time-consuming, hindering the expansion of urban digital twin use cases and slowing down comprehensive maintenance.

Method used

A system utilizing a generative model, specifically a Generative Adversarial Network (GAN), to efficiently generate and integrate highly accurate 3D maps, which includes data collection, preprocessing, training, integration with existing data, cloud storage, and user access, enabling rapid and efficient 3D map generation and distribution.

Benefits of technology

Enables the rapid and efficient creation of highly accurate 3D maps, making them accessible to stakeholders for timely decision-making and use case development, differentiating from existing systems like PLATEAU.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064572000001_ABST
    Figure 2026064572000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Means for collecting geographic information data, A means for preprocessing the collected geographic information data, Means for training generative models, A means for generating a 3D map using a trained generative model, Means for integrating with existing map data, A method for uploading the generated 3D map to a cloud storage service and distributing it to users, A means for users to access and view and manipulate a 3D map that has been generated, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Creating a 3D map as the basis for an urban digital twin requires a lot of manual costs, and the slowness of comprehensive maintenance has become an issue. As a result, the groundwork for using urban digital twins has not expanded, and use case development has not progressed. There is a need for a method to solve this problem, generate 3D maps more efficiently, and accelerate the use of urban digital twins.

Means for Solving the Problems

[0005] The present invention aims to solve the above problems by the following means: a system including means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, and means for users to access, view, and manipulate the generated 3D map. In particular, by using a machine learning algorithm in the generative model, and specifically including a generative adversarial network, it is possible to generate a 3D map with high accuracy and speed.

[0006] "Geographic information data" refers to data that shows the geographical characteristics of a city, and includes information such as buildings, roads, land use, and topography.

[0007] "Means of collection" refers to the processes and technologies for obtaining geographic information data and related data from the internet, databases, APIs, sensors, etc.

[0008] "Preprocessing" refers to the process of converting collected geographic information data into a format optimal for analysis and use, and includes noise reduction, format conversion, and coordinate system standardization.

[0009] A "generative model" is a model that uses machine learning algorithms to learn various features from input data and generate new data.

[0010] "Training methods" refer to the process of using collected data to train a generative model and adjusting the appropriate machine learning algorithm.

[0011] A "trained generative model" is a model in which a machine learning algorithm has been sufficiently trained for a specific task, and is capable of making accurate predictions and generating data for new data.

[0012] "Means for generating 3D maps" refers to the process or technology of creating 3D map data from 2D geographic information data using a trained generative model.

[0013] "Means of integration" refers to the process of combining the generated 3D map with existing map data and other related data to create a consistent digital twin.

[0014] A "cloud storage service" refers to an online storage service that allows data to be stored and managed via the internet, making it accessible to multiple users.

[0015] "Means of distribution to users" refers to the process of uploading the generated 3D map data to a cloud storage service and distributing it so that users can access it in a specific way.

[0016] "Means by which users access, view, and manipulate" refers to the function that allows users to access cloud storage services through a web browser or dedicated application and to display and manipulate the generated 3D maps.

[0017] A "machine learning algorithm" is a mathematical model and technique that learns from data and extracts patterns and features, and has the ability to generate or predict new data based on the patterns discovered.

[0018] A "generative adversarial network" is a technique that generates realistic data by using two networks—a generative model and a discriminative model—and having them compete with each other during training. [Brief explanation of the drawing]

[0019] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0020] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0023] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0024] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0025] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0027] [First Embodiment]

[0028] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0029] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0032] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0035] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0039] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0040] As an embodiment of this invention, the specific operation of a city digital twin construction system using generative AI will be described. This system consists of the following steps.

[0041] First, the server collects geographic information data for the city. Specifically, it retrieves data such as buildings, roads, land use, and topography from geographic information databases, and also downloads publicly available aerial and satellite imagery. This geographic information data includes formats such as Shapefile and GeoJSON.

[0042] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, a denoising program is run to filter out unnecessary information from the collected data, and all data is converted to the WGS84 coordinate system to unify the coordinate system. Furthermore, all data is converted to and unified in GeoJSON format.

[0043] Next, the server trains the generative AI model. For this, manually created 3D map data is used as the training dataset. For the generative model, for example, a Generative Adversarial Network (GAN) is used. During training, the AI ​​model learns features such as the shapes of buildings and roads, enabling the generation of highly accurate 3D maps.

[0044] Next, the server generates a 3D map using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model generates a predicted 3D shape from this data. The generated 3D map is stored in the server's internal database.

[0045] Next, the server performs a process to integrate the generated 3D map with existing map data. This process involves running scripts to retrieve supplementary data (e.g., trees and walkways) from the existing database and integrate it with the generated 3D map. The integrated data undergoes quality checks, and after correcting inconsistencies and duplicate data, a single, consistent digital twin is created.

[0046] Next, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. In this step, the integrated 3D map is uploaded to the cloud, and a REST API is configured to expose the endpoint. Stakeholders are notified of the API endpoint and how to access it.

[0047] Finally, the terminal (user) accesses the cloud server and uses the generated 3D map. For example, the user can open a web browser, access the provided URL, enter their authentication information, and access the 3D map from the interface. The user can interactively manipulate the 3D map to check the height of buildings and the width of roads in a specific area.

[0048] This system enables the efficient creation of digital twins of cities, making them accessible to a wide range of stakeholders. Furthermore, rapid 3D map generation allows stakeholders to quickly begin exploring use cases. This differentiates the system from the Ministry of Land, Infrastructure, Transport and Tourism's PLATEAU system.

[0049] The following describes the processing flow.

[0050] Step 1:

[0051] The server collects urban geographic information data. Specifically, it accesses geographic information databases and uses APIs to obtain urban building data, road data, and topographic data. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools.

[0052] Step 2:

[0053] The server preprocesses the collected geographic information data. First, it runs a denoising program to filter out unnecessary information and erroneous data points from the collected data. Next, it converts all data to the WGS84 coordinate system to unify the data's coordinate system. Finally, it converts each dataset to GeoJSON format to unify the data.

[0054] Step 3:

[0055] The server trains the generative AI model. It prepares manually created 3D map data as the training dataset and configures the Generative Adversarial Network (GAN) algorithm. The training data is input into the AI ​​model, and the model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0056] Step 4:

[0057] The server generates 3D maps using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is stored in an internal database.

[0058] Step 5:

[0059] The server integrates the generated 3D map with existing map data. It retrieves supplementary data (e.g., tree and footpath data) from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to eliminate redundancies and inconsistencies for consistency, creating the final digital twin data.

[0060] Step 6:

[0061] The server uploads the integrated digital twin data to a cloud storage service. A REST API is configured to access the uploaded data, and the endpoint is exposed. Furthermore, stakeholders are notified of the API endpoint and how to access it.

[0062] Step 7:

[0063] The terminal (user) accesses a cloud server and utilizes the generated 3D map. The user opens a web browser, accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map and, for example, investigate the height of buildings and the width of roads in a specific area for real estate development planning.

[0064] (Example 1)

[0065] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0066] Providing highly accurate geographic information is essential in modern urban planning and infrastructure management. However, traditional methods involve significant time and effort in collecting and preprocessing geographic data, and generating and integrating 3D models. As a result, providing timely data is difficult, hindering rapid decision-making and action planning by stakeholders. Furthermore, unifying data in different formats and building a consistent digital twin requires advanced technology and expertise. A new system is needed to solve these problems.

[0067] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0068] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a three-dimensional map using the trained generative model, means for integrating with existing map data, means for uploading the generated three-dimensional map to a cloud storage service and distributing it to users, and means for users to access, view, and manipulate the generated three-dimensional map. This enables the rapid and efficient collection and preprocessing of highly accurate geographic information, and the generation and integration of highly accurate three-dimensional maps in a short time using a generative AI model. Furthermore, by utilizing cloud storage, stakeholders can easily access and interact with the data.

[0069] "Geographic information data" refers to data that includes spatial and locational information about urban and regional features (buildings, roads, land use, topography, etc.).

[0070] "Preprocessing" refers to the process of removing noise from collected geographic information data, checking for consistency, unifying coordinate systems, and converting formats.

[0071] A "generative model" is a model that uses machine learning algorithms to generate a desired output (for example, a three-dimensional map) from input data.

[0072] A "three-dimensional map" is map data that visualizes the shape and location information of features in a city or region in three-dimensional space.

[0073] A "cloud storage service" is a service that stores and manages data online, allowing users to access it via the internet.

[0074] A "user" is a person or organization that has the authority to access the generated three-dimensional map and to view and manipulate its information.

[0075] A Generative Adversarial Network (GAN) is a type of generative model in which two paired neural networks compete to generate data.

[0076] A "digital twin" is a virtual model that accurately replicates a physical, real-world object or system, reflecting real-time data.

[0077] A "REST API" is a type of interface for providing web services, and it is a set of design principles for manipulating resources using the HTTP protocol.

[0078] An "API endpoint" is a URL or URI used to access a specific resource or function designated through an API.

[0079] Modes for carrying out the invention

[0080] As an embodiment of this invention, the specific operation of a city digital twin construction system using generative AI is shown below. This system collects and preprocesses geographic information of a city, creates and integrates a three-dimensional map using a generative model, and provides it via a cloud storage service.

[0081] First, the server collects geographic data of the city. Specifically, it uses the OpenStreetMap API to obtain data on buildings, roads, land use, and topography, and downloads the latest satellite imagery from NASA's Landsat satellites. This geographic data includes formats such as Shapefile and GeoJSON. This allows for obtaining up-to-date and detailed geographic information about the city.

[0082] Next, the server preprocesses the collected geographic data. This process includes denoising, consistency checking, and format conversion using the Python Geopandas library. Specifically, Geopandas is used to convert the coordinate system to WGS84, filter out unnecessary data, and unify all data by converting it to GeoJSON format. This process results in consistent geographic data.

[0083] Next, the server trains the generative AI model. Detailed, manually created 3D map data is used as the training dataset. A Generative Adversarial Network (GAN) is used as the generative model. During training, the AI ​​model learns features such as the shapes of buildings and roads, enabling the generation of highly accurate 3D maps. The TENSORFLOW® library is used to train the generative model.

[0084] Next, the server uses a trained generative AI model to generate a 3D map based on new geographic data. New building and road data are input into the generative model, which then generates a predicted 3D shape from this data. The generated 3D map is stored in the server's internal database. This procedure ensures that the most up-to-date 3D map data is always generated.

[0085] Subsequently, the server integrates the generated 3D map with existing map data. This process involves retrieving supplementary data such as trees and footpaths from the existing database and running an integration script. Using the integration script, the generated 3D map and the existing map data are combined into a single, consistent digital twin. After integration, quality checks are also performed to correct inconsistencies and duplicate data.

[0086] Next, the server uploads the integrated 3D map data to a cloud storage service and distributes it to users. In this step, the data is uploaded to an Amazon S3 bucket, and a REST API is configured to expose the endpoint. Stakeholders are notified of the API endpoint and how to access it. This allows stakeholders to access the updated 3D map at any time.

[0087] Finally, the terminal (user) accesses the provided API endpoint to view and interact with the generated 3D map. The user opens a web browser, accesses the provided URL, and enters their authentication information. Through the interface, they can access the 3D map and check the height of buildings and the width of roads in a specific area. This system allows the user to utilize detailed 3D map information that is updated in real time.

[0088] As a concrete example, the server retrieves city building data using the OpenStreetMap API and preprocesses it using the latest satellite imagery downloaded from NASA's Landsat satellites. Then, it trains a GAN model in TensorFlow using manually created detailed 3D map data of Tokyo. Using the new geographic data as input, the trained model generates a highly accurate 3D map, which is then integrated with the existing data. The integrated data is uploaded to an Amazon S3 bucket, and an API endpoint is configured. Stakeholders can access the provided URL to view the 3D map and obtain the necessary information.

[0089] Examples of prompt statements include the following:

[0090] I want to generate 3D map data of a city. Please use the following geographic information data as input to generate detailed 3D shapes of buildings, roads, trees, etc.

[0091] Building data (Shapefile) obtained from OpenStreetMap

[0092] NASA Landsat satellite imagery (GeoTIFF)

[0093] Tree data (GeoJSON)

[0094] Based on the data above, please generate a high-precision 3D map.

[0095] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0096] Step 1:

[0097] The server collects geographic data of cities. Specifically, it uses the OpenStreetMap API to retrieve building and road data and downloads the latest satellite imagery from NASA's Landsat satellites. It takes API endpoints and query parameters as input and receives geographic data in Shapefile or GeoJSON format as output. This provides up-to-date and detailed geographic information of cities.

[0098] Step 2:

[0099] The server preprocesses the collected geographic data. Specifically, it uses the Python Geopandas library to convert the data's coordinate system to WGS84 and filters out unnecessary data. Using the diverse formats of geographic data collected as input, it obtains data in a unified GeoJSON format as output. This process ensures a consistent dataset. For example, to perform a coordinate system transformation using Geopandas, execute geo_df.to_crs(epsg=4326).

[0100] Step 3:

[0101] The server trains a generative AI model. Specifically, it trains a generative adversarial network (GAN) using TensorFlow with manually created 3D map data. Detailed, manually created 3D map data is used as input, and a trained generative AI model is obtained as output. This process allows the AI ​​model to learn the features of buildings and roads, enabling highly accurate 3D map generation. For example, the command `model.fit(training_data, epochs=50)` is used.

[0102] Step 4:

[0103] The server generates a 3D map using a trained generative AI model. Specifically, it takes new geographic data as input and generates the 3D shape predicted by the model. It uses newly collected and pre-processed geographic data as input and produces generated 3D map data as output. This output is stored in the server's internal database. For example, `model.predict(new_data)` is executed to generate the 3D shape.

[0104] Step 5:

[0105] The server integrates the generated 3D map with existing map data. Specifically, it retrieves supplementary data from the existing database and runs an integration script. It uses the new 3D map data and supplementary data (e.g., tree and sidewalk data) as input and obtains a consistent digital twin as output. By running the integration script, the data is integrated using the command integrate_data(new_3d_map, additional_data).

[0106] Step 6:

[0107] The server uploads the integrated 3D map to a cloud storage service and distributes it to stakeholders. Specifically, it uploads the data to an Amazon S3 bucket, configures a REST API, and exposes an endpoint. It uses the integrated data as input and obtains the data stored in the cloud and the API endpoint as output. For example, it uses the command `s3_client.upload_file('integrated_3d_map.json', 'bucket-name', 'path / in / bucket')`.

[0108] Step 7:

[0109] The user accesses the provided API endpoint to view and manipulate the generated 3D map. For example, the user opens a web browser, accesses the provided URL, and enters their authentication information. Using the specified URL and authentication information as input, they obtain an interactive display of the 3D map as output. This allows the user to check the height of buildings and the width of roads in a specific area.

[0110] (Application Example 1)

[0111] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0112] In recent years, in order to solve the increasingly complex traffic problems that accompany urban development, there is a need for real-time, high-precision map information and traffic data-driven operational management of autonomous vehicles. However, current systems have slow map information updates, making it difficult to respond quickly to changes in traffic conditions. Furthermore, the integration of information from multiple data sources is insufficient, making it difficult to provide optimal route guidance. This hinders the improvement of efficiency and safety of autonomous vehicles. This invention aims to solve these problems.

[0113] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0114] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, means for users to access, view, and manipulate the generated 3D map, and means for calculating the optimal route for an autonomous vehicle and performing operational management. This enables efficient operational management of autonomous vehicles based on highly accurate and real-time map information and traffic data.

[0115] "Geographic information data" refers to data that includes information about urban structure, land use, buildings, roads, topography, and so on.

[0116] "Preprocessing" is the process of removing noise from collected data, checking data consistency, and standardizing the format.

[0117] A "generative model" is an algorithm for generating 3D maps through learning, and it utilizes machine learning techniques.

[0118] A "3D map" is a digital map that represents urban structures such as buildings, roads, and terrain in three dimensions.

[0119] A "cloud storage service" is a data storage system that allows data to be stored via the internet and shared among multiple users.

[0120] A "user" is an end-user who utilizes the data and services provided by this system.

[0121] An "autonomous vehicle" is a vehicle that uses sensors and AI to automate its driving process.

[0122] An "optimal route" is a path that minimizes travel time and distance, taking into account traffic conditions and geographical features.

[0123] "Operational management" refers to the process of monitoring and adjusting the operating schedule, routes, and traffic conditions of autonomous vehicles.

[0124] The specific system configuration and operation for realizing this invention are described below.

[0125] Hardware and Software Overview

[0126] This invention utilizes the following main hardware and software.

[0127] hardware

[0128] Server: A server equipped with a high-performance CPU and large-capacity storage.

[0129] Autonomous vehicle onboard computer: Onboard computer equipped with a GPU.

[0130] Device: Smartphone (ANDROID® / iOS).

[0131] software

[0132] Cloud storage services: such as AWS® S3 and Google® Cloud Storage.

[0133] Traffic information APIs: such as Google Maps API and Here API.

[0134] Generative AI model: GAN (Generative Adversarial Network).

[0135] Data processing programs: Python, TensorFlow / PyTorch.

[0136] Explanation of program processing

[0137] The server collects geographic information data. This data is obtained from geographic information databases, aerial photographs, satellite imagery, etc. The collected geographic information data includes formats such as Shapefile and GeoJSON.

[0138] The server preprocesses the collected data. It performs processes such as noise reduction, data consistency checks, and format conversion, and converts all data into a unified format (GeoJSON).

[0139] To train the generative AI model, manually created 3D map data is used as the training dataset. A Generative Adversarial Network (GAN) is used as the generative model to learn features such as the shapes of buildings and roads.

[0140] A pre-trained generative AI model is used to generate 3D maps from new geographic data. These generated 3D maps are stored in the server's internal database.

[0141] Next, the generated 3D map is integrated with existing map data. A script is executed to retrieve supplementary data (e.g., trees and sidewalks) from the existing database and integrate it with the generated 3D map.

[0142] The integrated 3D map data is uploaded to a cloud storage service, and a REST API is configured to expose the endpoint. Users are notified of the API endpoint and how to access it.

[0143] Using a smartphone as a terminal, users access a cloud server and utilize the generated 3D map data. For example, a user can open a web browser, access the provided URL, enter authentication information, and access the 3D map from the interface.

[0144] Operation management of autonomous vehicles

[0145] Furthermore, this system acquires traffic information in real time and calculates the optimal route for autonomous vehicles. Traffic information is collected from traffic sensors, cameras, and social media.

[0146] The terminal's AI algorithm integrates generated 3D map data with real-time traffic information to calculate the optimal route. The calculated route is transmitted to the autonomous vehicle's onboard computer, and the vehicle operates according to that route.

[0147] Users can visually check route information and traffic conditions through the app's interface. For example, they can receive route guidance by entering the following prompt:

[0148] Example of a prompt

[0149] "Please tell me the best route from Tokyo Station to Shinjuku Station."

[0150] Upon receiving this prompt, the system can use the collected data and generated AI models to calculate the optimal route in real time and provide it to the user.

[0151] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0152] Step 1:

[0153] The server collects geographic information data. It retrieves data such as buildings, roads, land use, and topography from geographic information databases, and downloads publicly available aerial and satellite imagery. Input is data in Shapefile or GeoJSON format, and output is the collected raw data.

[0154] Step 2:

[0155] The server preprocesses the geographic information data it collects. A noise reduction program is run to filter out unnecessary information from the collected data, and all data is converted to the WGS84 coordinate system to unify the coordinate system. Furthermore, all data is converted to the GeoJSON format for standardization. The input is the collected raw data, and the output is the clean data after preprocessing.

[0156] Step 3:

[0157] The server trains a generative AI model. Manually created 3D map data is used as the training dataset, and a generative adversarial network (GAN) is used to train the model. The input is the training dataset, and the output is the trained generative AI model.

[0158] Step 4:

[0159] The server generates a 3D map using a trained generative AI model. New geographic data is input to the generative model, and the model generates a predicted 3D shape from this data. The input is new geographic data, and the output is the generated 3D map.

[0160] Step 5:

[0161] The server integrates the generated 3D map with existing map data. A script is executed to retrieve supplementary data (e.g., trees and sidewalks) from the existing database and integrate it with the generated 3D map. The input is the generated 3D map and supplementary data, and the output is the integrated 3D map.

[0162] Step 6:

[0163] The server uploads the integrated 3D map data to a cloud storage service, configures a REST API, and exposes an endpoint. The input is the integrated 3D map, and the output is the data on the cloud storage and the exposed API endpoint.

[0164] Step 7:

[0165] The user accesses the cloud server using their device (smartphone). They access the provided URL, enter their authentication information, and access the 3D map from the interface. The input is the user's authentication information, and the output is the 3D map display on the interface.

[0166] Step 8:

[0167] The server collects real-time traffic information from traffic sensors, cameras, social media, and other sources. The input is real-time traffic data, and the output is collected traffic information.

[0168] Step 9:

[0169] The terminal's AI algorithm integrates generated 3D map data with real-time traffic information to calculate the optimal route. The calculated route is transmitted to the autonomous vehicle's onboard computer, and the vehicle operates according to that route. The input is 3D map data and traffic information, and the output is the optimal route and its operation information.

[0170] Step 10:

[0171] Users visually check route information and traffic conditions through the app's interface. For example, they can receive route guidance by prompting them with a phrase like, "Tell me the best route from Tokyo Station to Shinjuku Station." The input is the user's prompt, and the output is the route information displayed on the screen.

[0172] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0173] As an embodiment of this invention, we will describe the specific operation of a system that combines a city digital twin construction system using generative AI with an emotion engine that recognizes user emotions. This system consists of the following steps.

[0174] First, the server collects geographic information data for cities. Specifically, it accesses geographic information databases and uses APIs to obtain building data, road data, and topographic data for cities. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools.

[0175] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, a denoising program is run to filter out unnecessary information and erroneous data points from the collected data, and all data is converted to the WGS84 coordinate system to unify the coordinate system. Furthermore, all data is converted and unified to the GeoJSON format.

[0176] Next, the server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset, and the Generative Adversarial Network (GAN) algorithm is configured. The training data is input into the AI ​​model, and the model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0177] Next, the server generates a 3D map using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is saved to an internal database.

[0178] Next, the server integrates the generated 3D map with existing map data. It retrieves supplementary data (e.g., tree and sidewalk data) from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to eliminate redundancies and inconsistencies for consistency, creating the final digital twin data.

[0179] Next, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. In this step, the integrated 3D map is uploaded to the cloud, and a REST API is configured to expose the endpoint. Stakeholders are notified of the API endpoint and how to access it.

[0180] Furthermore, this system incorporates an emotion engine that recognizes the user's emotions. The terminal (user) collects emotion data through a camera, microphone, or other sensor devices. The server analyzes the collected emotion data and uses machine learning algorithms to recognize the user's emotions in real time. For example, by analyzing the user's facial expressions, it is possible to identify the user's emotional state, such as whether they are excited or unhappy.

[0181] The recognized emotion data is reflected as feedback in the system interface. For example, if the user shows a surprised expression, the system can adjust the displayed content to emphasize information that is of interest to the user. Also, if the user shows a confused expression, the system will make dynamic changes such as displaying an operation guide.

[0182] Finally, the device (user) accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map and, for example, investigate the height of buildings and the width of roads in a specific area for real estate development planning. In addition, the emotion engine provides feedback on the user's emotional data, offering a more intuitive and personalized user experience.

[0183] This system enables the efficient construction of digital twins of cities, making them accessible to a wide range of stakeholders, while also enhancing the user experience through an emotion engine. Furthermore, rapid 3D map generation and the integration of emotion data expand the usability of urban digital twins, allowing stakeholders to quickly begin exploring use cases. This differentiates the system from the Ministry of Land, Infrastructure, Transport and Tourism's PLATEAU.

[0184] The following describes the processing flow.

[0185] Step 1:

[0186] The server collects urban geographic information data. Specifically, it accesses geographic information databases and uses APIs to obtain urban building data, road data, and topographic data. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools.

[0187] Step 2:

[0188] The server preprocesses the collected geographic information data. First, it runs a denoising program to filter out unnecessary information and erroneous data points from the collected data. Next, it converts all data to the WGS84 coordinate system to unify the data's coordinate system. Finally, it converts each dataset to GeoJSON format to unify the data.

[0189] Step 3:

[0190] The server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset. The Generative Adversarial Network (GAN) algorithm is configured, and the training data is input into the AI ​​model. The AI ​​model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0191] Step 4:

[0192] The server generates 3D maps using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is stored in an internal database.

[0193] Step 5:

[0194] The server integrates the generated 3D map with existing map data. It runs a script that retrieves supplementary data from the existing map database and integrates it with the 3D map. The integrated data is then processed to eliminate duplication and inconsistencies, creating a consistent digital twin.

[0195] Step 6:

[0196] The server uploads the integrated digital twin data to a cloud storage service. A REST API is configured to access the data, and the endpoint is exposed. Stakeholders are notified of the API endpoint and how to access it.

[0197] Step 7:

[0198] The user accesses a cloud server and utilizes the generated 3D map. The user opens a web browser, accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user interacts with the 3D map and investigates information about a specific area.

[0199] Step 8:

[0200] The device (user) collects emotional data using a camera, microphone, or other sensor devices. For example, the camera is used to capture the user's facial expressions, and the microphone is used to collect the tone of their voice.

[0201] Step 9:

[0202] The server analyzes the collected emotional data. Using machine learning algorithms, it analyzes the user's facial expressions and tone of voice to recognize the user's emotional state. For example, it can determine from facial expressions whether the user is excited or unhappy.

[0203] Step 10:

[0204] The server dynamically changes the system interface based on the analysis results. For example, if the user looks confused, it displays an operation guide. If the user looks surprised, it highlights information that might interest them.

[0205] This system efficiently constructs digital twins of cities, making them accessible to numerous stakeholders, while also enhancing the user experience through an emotion engine. Rapid 3D map generation and the integration of emotion data expand the usability of urban digital twins, allowing stakeholders to quickly begin exploring use cases. This differentiates the system from other existing systems.

[0206] (Example 2)

[0207] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0208] Conventional urban digital twin systems required improved efficiency in collecting geographic information data and generating 3D maps, but they lacked the ability to recognize user emotions in real time and reflect them in the interface. Therefore, it was difficult to provide users with an intuitive and personalized user experience. Furthermore, it was challenging to provide the generated 3D maps in real time while maintaining the consistency of the integrated data.

[0209] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0210] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, means for users to access, view, and manipulate the generated 3D map, means for collecting sentiment data, means for analyzing the collected sentiment data and recognizing the user's sentiment in real time, and means for dynamically adjusting the interface based on the recognized sentiment data. This enables the efficient construction of a city digital twin and the provision of an intuitive and personalized user experience that takes into account the user's sentiment.

[0211] "Geographic information data" refers to a dataset containing geographical location information, providing information about elements such as buildings, roads, and topography.

[0212] "Preprocessing" is the process of ensuring the consistency of collected data by performing tasks such as noise reduction, format conversion, and coordinate system standardization.

[0213] A "generative model" is an algorithm that generates new data based on input data, and often refers specifically to machine learning algorithms.

[0214] A "trained generative model" is a generative model that has been learned using training data and optimized to perform a specific task.

[0215] A "3D map" is a visualization of geographic information data in three-dimensional space, and is a data structure designed to realistically reproduce real-world geographical features.

[0216] A "cloud storage service" is a remote storage service that allows you to store and manage data via the internet.

[0217] "Emotional data" refers to data that indicates a user's emotional state, collected based on factors such as the user's facial expressions and tone of voice.

[0218] "Recognizing in real time" means performing the entire process from data collection to analysis immediately, and providing results without any time delay.

[0219] An "interface" refers to the screen or control panel that a user uses to interact with a system, and is the point of contact that provides a user experience.

[0220] "Dynamic adjustment" refers to changing the system's display content and behavior in real time according to the user's situation and actions.

[0221] Modes for carrying out the invention

[0222] As an embodiment of this invention, we will describe the specific operation of a system that combines a city digital twin construction system using generative AI with an emotion engine that recognizes user emotions. This system consists of the following steps.

[0223] First, the server collects geographic information data for cities. Specifically, it accesses geographic information databases and uses APIs to obtain building data, road data, and terrain data for cities. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools. For example, it can use open-source geographic information databases (e.g., OpenStreetMap API). It can also obtain the latest satellite imagery through NASA's satellite imagery service. Furthermore, it is possible to scrape the necessary image data using the Python library BeautifulSoup.

[0224] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, it uses Python's NumPy and Pandas libraries to check data consistency and remove unnecessary information or incorrect data points. It also uses the GeoPandas library to unify the coordinate system of the collected data to WGS84 (e.g., EPSG:4326) and convert it to GeoJSON format.

[0225] Next, the server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset, and the model is built using the Generative Adversarial Network (GAN) algorithm. The server uses libraries such as TensorFlow and PyTorch to input the manually created 3D map data into the AI ​​model and runs a program to learn the shapes of buildings and roads. This process is repeated until the model's accuracy is sufficiently high.

[0226] Next, the server generates a 3D map using a trained generative AI model. For example, new geographic information data (building data, road data, etc.) is input into the generative AI model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is stored in an internal database. By using geographic information databases such as PostGIS, efficient data storage and retrieval are possible.

[0227] Next, the server integrates the generated 3D map with existing map data. It retrieves supplementary data (e.g., tree and sidewalk data) from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to eliminate duplication and inconsistencies, completing it as the final digital twin data.

[0228] Next, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. In this step, the integrated 3D map is uploaded to the cloud (for example, Amazon S3). To set up a REST API and notify stakeholders how to access it, an API endpoint can be built using the Flask framework.

[0229] Furthermore, this system incorporates an emotion engine that recognizes the user's emotions. The terminal (user) collects emotion data through a camera, microphone, or other sensor devices. For example, the camera is used to capture the user's facial expressions and the image data is sent to the server. The microphone is also used to collect audio data, capturing the user's speech and tone of voice.

[0230] The server analyzes collected emotional data and uses machine learning algorithms to recognize the user's emotions in real time. Specifically, it analyzes facial expressions using OpenCV and Dlib libraries, and analyzes audio data using TensorFlow, etc., to determine the user's emotional state. For example, if the user has a surprised expression, it recognizes an excited state, and if they have a grumpy expression, it understands the situation.

[0231] The recognized emotion data is reflected as feedback in the system interface. For example, if the user shows a surprised expression, the system adjusts the displayed content to emphasize information that will interest the user. Also, if the user shows a confused expression, the system makes dynamic changes such as displaying an operation guide.

[0232] Finally, the terminal (user) accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user accesses the provided URL through a browser and logs in using the authentication information. It provides interactive functionality to manipulate the 3D map, making it possible, for example, to survey the height of buildings and the width of roads in a specific area for real estate development planning. The emotion engine provides real-time feedback of the user's emotional data, offering an intuitive and personalized user experience.

[0233] As a concrete example, the following is an example of a prompt statement.

[0234] Example of a prompt:

[0235] Please describe a program that acquires satellite imagery and geographic information data, preprocesses it with noise reduction and format conversion, generates a 3D map using an AI model, integrates it with existing data, and uploads it to cloud storage. Also, please describe an emotion engine that recognizes user emotions and provides real-time feedback.

[0236] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0237] Step 1:

[0238] The server collects urban geographic information data. Specifically, the server accesses a geographic information database and retrieves building data, road data, and terrain data via an API. The input requires the database's API endpoint and request parameters, while the output is datasets of buildings, roads, and terrain. In addition, the server downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools. The BeautifulSoup library in Python is used for these operations.

[0239] Step 2:

[0240] The server preprocesses the collected geographic information data. The specific inputs are building data, road data, terrain data, and image data collected in Step 1. First, the data consistency is checked using Python's NumPy and Pandas libraries, and unnecessary information and incorrect data points are filtered out. Next, the GeoPandas library is used to unify the data coordinate system to WGS84 (EPSG:4326) and convert it to GeoJSON format. The output is preprocessed geographic information data in a unified format.

[0241] Step 3:

[0242] The server trains a generative AI model. The specific inputs are collected and pre-processed geographic data and manually created 3D map datasets. The server uses TensorFlow and PyTorch libraries to train a generative adversarial network (GAN) algorithm with these inputs. The datasets are fed into the AI ​​model to learn the shapes of buildings and roads. This process is repeated iteratively until the model's accuracy is sufficiently high. The output is the trained AI model.

[0243] Step 4:

[0244] The server generates 3D maps using a trained generative AI model. The specific inputs are the trained AI model and new geographic data (building data, road data, etc.). This data is input to the model, which then predicts and generates the 3D shape. The output is the generated 3D map data, which is stored in an internal database (e.g., PostGIS).

[0245] Step 5:

[0246] The server integrates the generated 3D map with existing map data. The specific inputs are the generated 3D map data and supplementary data (tree data, footpath data, etc.) obtained from existing map databases. The server executes scripts to integrate these, resolving duplicate data and inconsistencies. The output is integrated, consistent digital twin data.

[0247] Step 6:

[0248] The server uploads digital twin data to a cloud storage service and distributes it to stakeholders. The specific input is integrated digital twin data. The server uploads the data to a cloud storage service such as Amazon S3. Furthermore, it configures a REST API using the Flask framework and exposes an API endpoint. The output is the digital twin data on cloud storage and the API endpoint.

[0249] Step 7:

[0250] The device (user) collects emotional data using sensor devices such as cameras and microphones. Specific inputs include the user's facial expressions and voice data. The device captures the user's facial expressions through the camera and collects image data. It also collects voice data using the microphone. The output is the collected emotional data.

[0251] Step 8:

[0252] The server analyzes collected emotion data and recognizes the user's emotions in real time. The specific input is emotion data (facial expressions, voice data, etc.) sent from the terminal. The server analyzes facial expressions using OpenCV and Dlib libraries, and analyzes voice data using TensorFlow, etc., to determine the user's emotional state. The output is the recognized emotion data.

[0253] Step 9:

[0254] The server dynamically adjusts the interface based on recognized emotion data. The specific inputs are the recognized emotion data and the data of the currently displayed interface. The server executes a program that adjusts the displayed content based on the user's emotional state. For example, if the user shows a surprised expression, the system highlights interesting information. If the user is confused, it displays an operation guide. The output is the adjusted interface.

[0255] Step 10:

[0256] The user accesses the provided URL and enters their authentication information to access the 3D map. The specific inputs are the provided URL and authentication information. The user accesses the URL using a web browser and logs in using the authentication information. The user can interactively manipulate the 3D map to investigate building heights and road widths in specific areas for real estate development planning. The output is the interactively manipulated 3D map.

[0257] (Application Example 2)

[0258] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0259] In modern industrial settings, it is crucial to monitor the operation of factory robots in real time and direct them to perform optimal actions. However, current systems have limited means of accurately understanding robot operation, and there is a lack of means to improve the user experience. Furthermore, there are currently no systems that improve work efficiency by analyzing the user's emotional state and providing feedback. It is desirable to provide a system that solves these problems and manages robot operations more effectively.

[0260] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0261] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, means for users to access, view, and manipulate the generated 3D map, means for collecting and analyzing user sentiment data in real time, and means for adjusting the user experience based on the analyzed sentiment data. This makes it possible to monitor the operating status of robots in a factory in real time, analyze the emotional state of users, and provide optimal feedback.

[0262] "Geographic information data" refers to data that includes geographical information such as buildings, roads, and topography within cities and factories.

[0263] "Preprocessing" refers to data preparation tasks such as denoising, consistency checking, and format conversion of collected data.

[0264] A "generative model" is an artificial intelligence model that is trained to generate new data based on collected data.

[0265] A "3D map" is a three-dimensional map generated based on geographic information data, representing buildings, roads, and other topographic elements in 3D.

[0266] A "cloud storage service" is a service that allows you to store and manage data via the internet and access it when needed.

[0267] "User" refers to an individual or organization that uses the system to view and manipulate 3D maps.

[0268] "Emotional data" refers to data about a user's emotional state obtained from their facial expressions, voice, and other sources.

[0269] "Analysis" is the process of extracting and understanding information based on collected data.

[0270] An "algorithm" is a method for solving a problem through a series of steps or calculations.

[0271] "Integration" is the process of combining multiple datasets into a single, consistent dataset.

[0272] As an embodiment of this invention, a system for real-time monitoring and control of robot movements within a factory is provided. The system combines a factory digital twin construction system using a generative AI model with an emotion engine that recognizes user emotions. The system operates in the following specific steps.

[0273] First, the server collects geographical information data within the factory. Specifically, it acquires robot location data, operating status, and 3D scan data of the environment. This includes collecting data from sensor devices and cameras within the factory.

[0274] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, it runs a denoising program to filter out unnecessary information from the collected data and converts all data to the WGS84 coordinate system to unify the coordinate system. Furthermore, it converts all data to the GeoJSON format for unification.

[0275] Next, the server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset, and the Generative Adversarial Network (GAN) algorithm is configured. The training data is input into the AI ​​model, and the model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0276] Next, the server generates a 3D map using a trained generative AI model. New geographic data is input into the generative model, which then predicts and generates the 3D shape from this data. The generated 3D map data is stored in an internal database.

[0277] Next, the server integrates the generated 3D map with existing map data. It retrieves supplementary data from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to remove duplicates and inconsistencies for consistency, creating the final digital twin data.

[0278] Furthermore, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. The data uploaded to the cloud is then configured with a REST API to expose an endpoint. Stakeholders are notified of the API endpoint and how to access it.

[0279] In addition, this system incorporates an emotion engine that recognizes the user's emotions. The terminal (user) collects emotion data through cameras, microphones, and other sensor devices. For example, it uses the smartphone's camera and microphone to capture facial expressions and voice. The server analyzes the collected emotion data and uses machine learning algorithms to recognize the user's emotions in real time. Specifically, by analyzing the user's facial expressions and voice, it can identify emotional states such as whether the user is excited or unhappy.

[0280] The recognized emotion data is reflected as feedback to the system interface. For example, when the user shows a surprised expression, the system can adjust the display content to emphasize information that attracts the user's interest. Also, when the user has a confused expression, dynamic changes such as displaying an operation guide are made by the system.

[0281] Finally, the terminal (user) accesses the provided URL, enters the authentication information, and accesses the 3D map from the interface. The user can interactively operate the 3D map, for example, monitor and control the operation of robots working in a specific area of the factory. In addition, the user's emotion data is fed back by the emotion engine, providing a more intuitive and personalized operation experience.

[0282] This system enables real-time monitoring of the operation of robots in the factory, analysis of the user's emotional state, and provision of optimal feedback, improving both work efficiency and the operation experience.

[0283] Example of prompt text

[0284] "Please develop an application that generates a 3D digital twin model of factory robots and displays and controls it in real time on a smartphone. Also, add a function to analyze the user's facial expressions and voice emotions and provide feedback."

[0285] The flow of specific processing in Application Example 2 will be described using FIG. 14.

[0286] Step 1:

[0287] The server collects geographical information data in the factory. The input data is the position data, operating status, and 3D scan data of the environment of the robots obtained from sensor devices and cameras. The server receives these data and stores them in the database.

[0288] Step 2:

[0289] The server preprocesses the collected geographic information data. Specifically, it denoises the input data, checks for data consistency, and converts it to GeoJSON format. This process removes noise and outputs a dataset in a consistent format.

[0290] Step 3:

[0291] The server trains a generative AI model. It takes manually created 3D map data as the training dataset and uses a generative adversarial network (GAN) algorithm to train the model. This training process is repeated until the model's accuracy reaches a sufficient level, and a trained model with optimal parameters is output.

[0292] Step 4:

[0293] The server generates 3D maps using a trained generative AI model. New geographic data is input into the model, which then predicts and generates 3D shapes from this data. The generated 3D map data is stored in a database.

[0294] Step 5:

[0295] The server integrates the generated 3D map with existing map data. It retrieves supplementary data from the existing map data and runs an integration script to integrate it with the 3D map. To ensure consistency, duplicates and inconsistencies are resolved, and the final digital twin data is output.

[0296] Step 6:

[0297] The server uploads the digital twin data to a cloud storage service. A REST API is configured for the uploaded data, and an endpoint is exposed. Stakeholders can access the digital twin data using the notified API endpoint.

[0298] Step 7:

[0299] The device (user) uses a camera and microphone to collect emotional data. The user's facial expressions and voice are captured as input data and sent to the server.

[0300] Step 8:

[0301] The server analyzes the received emotional data. Using machine learning algorithms, it recognizes the user's emotional state from the input data. As a result of the analysis, it outputs the emotional state, such as whether the user is excited or unhappy.

[0302] Step 9:

[0303] The server adjusts the system interface based on the analyzed emotion data. For example, if the user has a surprised expression, the system displays information to highlight information that will interest the user. If the user appears confused, dynamic changes are made, such as displaying an operation guide.

[0304] Step 10:

[0305] The user accesses the provided URL, enters their authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map and monitor and control the movements of robots within the factory in real time. Furthermore, they can enjoy a more intuitive and personalized operating experience by utilizing emotional data fed back by the emotion engine.

[0306] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0307] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">)Generative AIs such as etc. can be mentioned. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is input. The data generation model 58 infers the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization, etc.

[0308] In the above embodiment, an example form in which specific processing is performed by the data processing device 12 is given, but the technology of the present disclosure is not limited to this, and specific processing may be performed by the smart device 14.

[0309] [Second Embodiment]

[0310] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0311] As shown in FIG. 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0312] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of the "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network), etc.

[0313] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0314] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0315] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0316] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0317] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0318] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0319] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0320] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0321] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0322] As an embodiment of this invention, the specific operation of a city digital twin construction system using generative AI will be described. This system consists of the following steps.

[0323] First, the server collects geographic information data for the city. Specifically, it retrieves data such as buildings, roads, land use, and topography from geographic information databases, and also downloads publicly available aerial and satellite imagery. This geographic information data includes formats such as Shapefile and GeoJSON.

[0324] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, a denoising program is run to filter out unnecessary information from the collected data, and all data is converted to the WGS84 coordinate system to unify the coordinate system. Furthermore, all data is converted to and unified in GeoJSON format.

[0325] Next, the server trains the generative AI model. For this, manually created 3D map data is used as the training dataset. For the generative model, for example, a Generative Adversarial Network (GAN) is used. During training, the AI ​​model learns features such as the shapes of buildings and roads, enabling the generation of highly accurate 3D maps.

[0326] Next, the server generates a 3D map using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model generates a predicted 3D shape from this data. The generated 3D map is stored in the server's internal database.

[0327] Next, the server performs a process to integrate the generated 3D map with existing map data. This process involves running scripts to retrieve supplementary data (e.g., trees and walkways) from the existing database and integrate it with the generated 3D map. The integrated data undergoes quality checks, and after correcting inconsistencies and duplicate data, a single, consistent digital twin is created.

[0328] Next, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. In this step, the integrated 3D map is uploaded to the cloud, and a REST API is configured to expose the endpoint. Stakeholders are notified of the API endpoint and how to access it.

[0329] Finally, the terminal (user) accesses the cloud server and uses the generated 3D map. For example, the user can open a web browser, access the provided URL, enter their authentication information, and access the 3D map from the interface. The user can interactively manipulate the 3D map to check the height of buildings and the width of roads in a specific area.

[0330] This system enables the efficient creation of digital twins of cities, making them accessible to a wide range of stakeholders. Furthermore, rapid 3D map generation allows stakeholders to quickly begin exploring use cases. This differentiates the system from the Ministry of Land, Infrastructure, Transport and Tourism's PLATEAU system.

[0331] The following describes the processing flow.

[0332] Step 1:

[0333] The server collects urban geographic information data. Specifically, it accesses geographic information databases and uses APIs to obtain urban building data, road data, and topographic data. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools.

[0334] Step 2:

[0335] The server preprocesses the collected geographic information data. First, it runs a denoising program to filter out unnecessary information and erroneous data points from the collected data. Next, it converts all data to the WGS84 coordinate system to unify the data's coordinate system. Finally, it converts each dataset to GeoJSON format to unify the data.

[0336] Step 3:

[0337] The server trains the generative AI model. It prepares manually created 3D map data as the training dataset and configures the Generative Adversarial Network (GAN) algorithm. The training data is input into the AI ​​model, and the model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0338] Step 4:

[0339] The server generates 3D maps using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is stored in an internal database.

[0340] Step 5:

[0341] The server integrates the generated 3D map with existing map data. It retrieves supplementary data (e.g., tree and footpath data) from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to eliminate redundancies and inconsistencies for consistency, creating the final digital twin data.

[0342] Step 6:

[0343] The server uploads the integrated digital twin data to a cloud storage service. A REST API is configured to access the uploaded data, and the endpoint is exposed. Furthermore, stakeholders are notified of the API endpoint and how to access it.

[0344] Step 7:

[0345] The terminal (user) accesses a cloud server and utilizes the generated 3D map. The user opens a web browser, accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map and, for example, investigate the height of buildings and the width of roads in a specific area for real estate development planning.

[0346] (Example 1)

[0347] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0348] Providing highly accurate geographic information is essential in modern urban planning and infrastructure management. However, traditional methods involve significant time and effort in collecting and preprocessing geographic data, and generating and integrating 3D models. As a result, providing timely data is difficult, hindering rapid decision-making and action planning by stakeholders. Furthermore, unifying data in different formats and building a consistent digital twin requires advanced technology and expertise. A new system is needed to solve these problems.

[0349] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0350] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a three-dimensional map using the trained generative model, means for integrating with existing map data, means for uploading the generated three-dimensional map to a cloud storage service and distributing it to users, and means for users to access, view, and manipulate the generated three-dimensional map. This enables the rapid and efficient collection and preprocessing of highly accurate geographic information, and the generation and integration of highly accurate three-dimensional maps in a short time using a generative AI model. Furthermore, by utilizing cloud storage, stakeholders can easily access and interact with the data.

[0351] "Geographic information data" refers to data that includes spatial and locational information about urban and regional features (buildings, roads, land use, topography, etc.).

[0352] "Preprocessing" refers to the process of removing noise from collected geographic information data, checking for consistency, unifying coordinate systems, and converting formats.

[0353] A "generative model" is a model that uses machine learning algorithms to generate a desired output (for example, a three-dimensional map) from input data.

[0354] A "three-dimensional map" is map data that visualizes the shape and location information of features in a city or region in three-dimensional space.

[0355] A "cloud storage service" is a service that stores and manages data online, allowing users to access it via the internet.

[0356] A "user" is a person or organization that has the authority to access the generated three-dimensional map and to view and manipulate its information.

[0357] A Generative Adversarial Network (GAN) is a type of generative model in which two paired neural networks compete to generate data.

[0358] A "digital twin" is a virtual model that accurately replicates a physical, real-world object or system, reflecting real-time data.

[0359] A "REST API" is a type of interface for providing web services, and it is a set of design principles for manipulating resources using the HTTP protocol.

[0360] An "API endpoint" is a URL or URI used to access a specific resource or function designated through an API.

[0361] Modes for carrying out the invention

[0362] As an embodiment of this invention, the specific operation of a city digital twin construction system using generative AI is shown below. This system collects and preprocesses geographic information of a city, creates and integrates a three-dimensional map using a generative model, and provides it via a cloud storage service.

[0363] First, the server collects geographic data of the city. Specifically, it uses the OpenStreetMap API to obtain data on buildings, roads, land use, and topography, and downloads the latest satellite imagery from NASA's Landsat satellites. This geographic data includes formats such as Shapefile and GeoJSON. This allows for obtaining up-to-date and detailed geographic information about the city.

[0364] Next, the server preprocesses the collected geographic data. This process includes denoising, consistency checking, and format conversion using the Python Geopandas library. Specifically, Geopandas is used to convert the coordinate system to WGS84, filter out unnecessary data, and unify all data by converting it to GeoJSON format. This process results in consistent geographic data.

[0365] Next, the server trains the generative AI model. Detailed, manually created 3D map data is used as the training dataset. A Generative Adversarial Network (GAN) is used as the generative model. During training, the AI ​​model learns features such as the shapes of buildings and roads, enabling the generation of highly accurate 3D maps. The TensorFlow library is used to train the generative model.

[0366] Next, the server uses a trained generative AI model to generate a 3D map based on new geographic data. New building and road data are input into the generative model, which then generates a predicted 3D shape from this data. The generated 3D map is stored in the server's internal database. This procedure ensures that the most up-to-date 3D map data is always generated.

[0367] Subsequently, the server integrates the generated 3D map with existing map data. This process involves retrieving supplementary data such as trees and footpaths from the existing database and running an integration script. Using the integration script, the generated 3D map and the existing map data are combined into a single, consistent digital twin. After integration, quality checks are also performed to correct inconsistencies and duplicate data.

[0368] Next, the server uploads the integrated 3D map data to a cloud storage service and distributes it to users. In this step, the data is uploaded to an Amazon S3 bucket, and a REST API is configured to expose the endpoint. Stakeholders are notified of the API endpoint and how to access it. This allows stakeholders to access the updated 3D map at any time.

[0369] Finally, the terminal (user) accesses the provided API endpoint to view and interact with the generated 3D map. The user opens a web browser, accesses the provided URL, and enters their authentication information. Through the interface, they can access the 3D map and check the height of buildings and the width of roads in a specific area. This system allows the user to utilize detailed 3D map information that is updated in real time.

[0370] As a concrete example, the server retrieves city building data using the OpenStreetMap API and preprocesses it using the latest satellite imagery downloaded from NASA's Landsat satellites. Then, it trains a GAN model in TensorFlow using manually created detailed 3D map data of Tokyo. Using the new geographic data as input, the trained model generates a highly accurate 3D map, which is then integrated with the existing data. The integrated data is uploaded to an Amazon S3 bucket, and an API endpoint is configured. Stakeholders can access the provided URL to view the 3D map and obtain the necessary information.

[0371] Examples of prompt statements include the following:

[0372] I want to generate 3D map data of a city. Please use the following geographic information data as input to generate detailed 3D shapes of buildings, roads, trees, etc.

[0373] Building data (Shapefile) obtained from OpenStreetMap

[0374] NASA Landsat satellite imagery (GeoTIFF)

[0375] Tree data (GeoJSON)

[0376] Based on the data above, please generate a high-precision 3D map.

[0377] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0378] Step 1:

[0379] The server collects geographic data of cities. Specifically, it uses the OpenStreetMap API to retrieve building and road data and downloads the latest satellite imagery from NASA's Landsat satellites. It takes API endpoints and query parameters as input and receives geographic data in Shapefile or GeoJSON format as output. This provides up-to-date and detailed geographic information of cities.

[0380] Step 2:

[0381] The server preprocesses the collected geographic data. Specifically, it uses the Python Geopandas library to convert the data's coordinate system to WGS84 and filters out unnecessary data. Using the diverse formats of geographic data collected as input, it obtains data in a unified GeoJSON format as output. This process ensures a consistent dataset. For example, to perform a coordinate system transformation using Geopandas, execute geo_df.to_crs(epsg=4326).

[0382] Step 3:

[0383] The server trains a generative AI model. Specifically, it trains a generative adversarial network (GAN) using TensorFlow with manually created 3D map data. Detailed, manually created 3D map data is used as input, and a trained generative AI model is obtained as output. This process allows the AI ​​model to learn the features of buildings and roads, enabling highly accurate 3D map generation. For example, the command `model.fit(training_data, epochs=50)` is used.

[0384] Step 4:

[0385] The server generates a 3D map using a trained generative AI model. Specifically, it takes new geographic data as input and generates the 3D shape predicted by the model. It uses newly collected and pre-processed geographic data as input and produces generated 3D map data as output. This output is stored in the server's internal database. For example, `model.predict(new_data)` is executed to generate the 3D shape.

[0386] Step 5:

[0387] The server integrates the generated 3D map with existing map data. Specifically, it retrieves supplementary data from the existing database and runs an integration script. It uses the new 3D map data and supplementary data (e.g., tree and sidewalk data) as input and obtains a consistent digital twin as output. By running the integration script, the data is integrated using the command integrate_data(new_3d_map, additional_data).

[0388] Step 6:

[0389] The server uploads the integrated 3D map to a cloud storage service and distributes it to stakeholders. Specifically, it uploads the data to an Amazon S3 bucket, configures a REST API, and exposes an endpoint. It uses the integrated data as input and obtains the data stored in the cloud and the API endpoint as output. For example, it uses the command `s3_client.upload_file('integrated_3d_map.json', 'bucket-name', 'path / in / bucket')`.

[0390] Step 7:

[0391] The user accesses the provided API endpoint to view and manipulate the generated 3D map. For example, the user opens a web browser, accesses the provided URL, and enters their authentication information. Using the specified URL and authentication information as input, they obtain an interactive display of the 3D map as output. This allows the user to check the height of buildings and the width of roads in a specific area.

[0392] (Application Example 1)

[0393] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0394] In recent years, in order to solve the increasingly complex traffic problems that accompany urban development, there is a need for real-time, high-precision map information and traffic data-driven operational management of autonomous vehicles. However, current systems have slow map information updates, making it difficult to respond quickly to changes in traffic conditions. Furthermore, the integration of information from multiple data sources is insufficient, making it difficult to provide optimal route guidance. This hinders the improvement of efficiency and safety of autonomous vehicles. This invention aims to solve these problems.

[0395] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0396] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, means for users to access, view, and manipulate the generated 3D map, and means for calculating the optimal route for an autonomous vehicle and performing operational management. This enables efficient operational management of autonomous vehicles based on highly accurate and real-time map information and traffic data.

[0397] "Geographic information data" refers to data that includes information about urban structure, land use, buildings, roads, topography, and so on.

[0398] "Preprocessing" is the process of removing noise, checking data consistency, and standardizing the format of collected data.

[0399] A "generative model" is an algorithm for generating 3D maps through learning, and it utilizes machine learning techniques.

[0400] A "3D map" is a digital map that represents urban structures such as buildings, roads, and terrain in three dimensions.

[0401] A "cloud storage service" is a data storage system that allows data to be stored via the internet and shared among multiple users.

[0402] A "user" is an end-user who utilizes the data and services provided by this system.

[0403] An "autonomous vehicle" is a vehicle that uses sensors and AI to automate its driving process.

[0404] An "optimal route" is a path that minimizes travel time and distance, taking into account traffic conditions and geographical features.

[0405] "Operational management" refers to the process of monitoring and adjusting the operating schedule, routes, and traffic conditions of autonomous vehicles.

[0406] The specific system configuration and operation for realizing this invention are described below.

[0407] Hardware and Software Overview

[0408] This invention utilizes the following main hardware and software.

[0409] hardware

[0410] Server: A server equipped with a high-performance CPU and large-capacity storage.

[0411] Autonomous vehicle onboard computer: Onboard computer equipped with a GPU.

[0412] Device: Smartphone (Android / iOS).

[0413] software

[0414] Cloud storage services: such as AWS S3 and Google Cloud Storage.

[0415] Traffic information APIs: such as Google Maps API and Here API.

[0416] Generative AI model: GAN (Generative Adversarial Network).

[0417] Data processing programs: Python, TensorFlow / PyTorch.

[0418] Explanation of program processing

[0419] The server collects geographic information data. This data is obtained from geographic information databases, aerial photographs, satellite imagery, etc. The collected geographic information data includes formats such as Shapefile and GeoJSON.

[0420] The server preprocesses the collected data. It performs processes such as noise reduction, data consistency checks, and format conversion, and converts all data into a unified format (GeoJSON).

[0421] To train the generative AI model, manually created 3D map data is used as the training dataset. A Generative Adversarial Network (GAN) is used as the generative model to learn features such as the shapes of buildings and roads.

[0422] A pre-trained generative AI model is used to generate 3D maps from new geographic data. These generated 3D maps are stored in the server's internal database.

[0423] Next, the generated 3D map is integrated with existing map data. A script is executed to retrieve supplementary data (e.g., trees and sidewalks) from the existing database and integrate it with the generated 3D map.

[0424] The integrated 3D map data is uploaded to a cloud storage service, and a REST API is configured to expose the endpoint. Users are notified of the API endpoint and how to access it.

[0425] Using a smartphone as a terminal, users access a cloud server and utilize the generated 3D map data. For example, a user can open a web browser, access the provided URL, enter authentication information, and access the 3D map from the interface.

[0426] Operation management of autonomous vehicles

[0427] Furthermore, this system acquires traffic information in real time and calculates the optimal route for autonomous vehicles. Traffic information is collected from traffic sensors, cameras, and social media.

[0428] The terminal's AI algorithm integrates generated 3D map data with real-time traffic information to calculate the optimal route. The calculated route is transmitted to the autonomous vehicle's onboard computer, and the vehicle operates according to that route.

[0429] Users can visually check route information and traffic conditions through the app's interface. For example, they can receive route guidance by entering the following prompt:

[0430] Example of a prompt

[0431] "Please tell me the best route from Tokyo Station to Shinjuku Station."

[0432] Upon receiving this prompt, the system can use the collected data and generated AI models to calculate the optimal route in real time and provide it to the user.

[0433] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0434] Step 1:

[0435] The server collects geographic information data. It retrieves data such as buildings, roads, land use, and topography from geographic information databases, and downloads publicly available aerial and satellite imagery. Input is data in Shapefile or GeoJSON format, and output is the collected raw data.

[0436] Step 2:

[0437] The server preprocesses the geographic information data it collects. A noise reduction program is run to filter out unnecessary information from the collected data, and all data is converted to the WGS84 coordinate system to unify the coordinate system. Furthermore, all data is converted to the GeoJSON format for standardization. The input is the collected raw data, and the output is the clean data after preprocessing.

[0438] Step 3:

[0439] The server trains a generative AI model. Manually created 3D map data is used as the training dataset, and a generative adversarial network (GAN) is used to train the model. The input is the training dataset, and the output is the trained generative AI model.

[0440] Step 4:

[0441] The server generates a 3D map using a trained generative AI model. New geographic data is input to the generative model, and the model generates a predicted 3D shape from this data. The input is new geographic data, and the output is the generated 3D map.

[0442] Step 5:

[0443] The server integrates the generated 3D map with existing map data. A script is executed to retrieve supplementary data (e.g., trees and sidewalks) from the existing database and integrate it with the generated 3D map. The input is the generated 3D map and supplementary data, and the output is the integrated 3D map.

[0444] Step 6:

[0445] The server uploads the integrated 3D map data to a cloud storage service, configures a REST API, and exposes an endpoint. The input is the integrated 3D map, and the output is the data on the cloud storage and the exposed API endpoint.

[0446] Step 7:

[0447] The user accesses the cloud server using their device (smartphone). They access the provided URL, enter their authentication information, and access the 3D map from the interface. The input is the user's authentication information, and the output is the 3D map display on the interface.

[0448] Step 8:

[0449] The server collects real-time traffic information from traffic sensors, cameras, social media, and other sources. The input is real-time traffic data, and the output is collected traffic information.

[0450] Step 9:

[0451] The terminal's AI algorithm integrates generated 3D map data with real-time traffic information to calculate the optimal route. The calculated route is transmitted to the autonomous vehicle's onboard computer, and the vehicle operates according to that route. The input is 3D map data and traffic information, and the output is the optimal route and its operation information.

[0452] Step 10:

[0453] Users visually check route information and traffic conditions through the app's interface. For example, they can receive route guidance by prompting them with a phrase like, "Tell me the best route from Tokyo Station to Shinjuku Station." The input is the user's prompt, and the output is the route information displayed on the screen.

[0454] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0455] As an embodiment of this invention, we will describe the specific operation of a system that combines a city digital twin construction system using generative AI with an emotion engine that recognizes user emotions. This system consists of the following steps.

[0456] First, the server collects geographic information data for cities. Specifically, it accesses geographic information databases and uses APIs to obtain building data, road data, and topographic data for cities. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools.

[0457] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, a denoising program is run to filter out unnecessary information and erroneous data points from the collected data, and all data is converted to the WGS84 coordinate system to unify the coordinate system. Furthermore, all data is converted and unified to the GeoJSON format.

[0458] Next, the server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset, and the Generative Adversarial Network (GAN) algorithm is configured. The training data is input into the AI ​​model, and the model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0459] Next, the server generates a 3D map using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is saved to an internal database.

[0460] Next, the server integrates the generated 3D map with existing map data. It retrieves supplementary data (e.g., tree and sidewalk data) from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to eliminate redundancies and inconsistencies for consistency, creating the final digital twin data.

[0461] Next, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. In this step, the integrated 3D map is uploaded to the cloud, and a REST API is configured to expose the endpoint. Stakeholders are notified of the API endpoint and how to access it.

[0462] Furthermore, this system incorporates an emotion engine that recognizes the user's emotions. The terminal (user) collects emotion data through a camera, microphone, or other sensor devices. The server analyzes the collected emotion data and uses machine learning algorithms to recognize the user's emotions in real time. For example, by analyzing the user's facial expressions, it is possible to identify the user's emotional state, such as whether they are excited or unhappy.

[0463] The recognized emotion data is reflected as feedback in the system interface. For example, if the user shows a surprised expression, the system can adjust the displayed content to emphasize information that is of interest to the user. Also, if the user shows a confused expression, the system will make dynamic changes such as displaying an operation guide.

[0464] Finally, the device (user) accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map and, for example, investigate the height of buildings and the width of roads in a specific area for real estate development planning. In addition, the emotion engine provides feedback on the user's emotional data, offering a more intuitive and personalized user experience.

[0465] This system enables the efficient construction of digital twins of cities, making them accessible to a wide range of stakeholders, while also enhancing the user experience through an emotion engine. Furthermore, rapid 3D map generation and the integration of emotion data expand the usability of urban digital twins, allowing stakeholders to quickly begin exploring use cases. This differentiates the system from the Ministry of Land, Infrastructure, Transport and Tourism's PLATEAU.

[0466] The following describes the processing flow.

[0467] Step 1:

[0468] The server collects urban geographic information data. Specifically, it accesses geographic information databases and uses APIs to obtain urban building data, road data, and topographic data. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools.

[0469] Step 2:

[0470] The server preprocesses the collected geographic information data. First, it runs a denoising program to filter out unnecessary information and erroneous data points from the collected data. Next, it converts all data to the WGS84 coordinate system to unify the data's coordinate system. Finally, it converts each dataset to GeoJSON format to unify the data.

[0471] Step 3:

[0472] The server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset. The Generative Adversarial Network (GAN) algorithm is configured, and the training data is input into the AI ​​model. The AI ​​model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0473] Step 4:

[0474] The server generates 3D maps using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is stored in an internal database.

[0475] Step 5:

[0476] The server integrates the generated 3D map with existing map data. It runs a script that retrieves supplementary data from the existing map database and integrates it with the 3D map. The integrated data is then processed to eliminate duplication and inconsistencies, creating a consistent digital twin.

[0477] Step 6:

[0478] The server uploads the integrated digital twin data to a cloud storage service. A REST API is configured to access the data, and the endpoint is exposed. Stakeholders are notified of the API endpoint and how to access it.

[0479] Step 7:

[0480] The user accesses a cloud server and utilizes the generated 3D map. The user opens a web browser, accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user interacts with the 3D map and investigates information about a specific area.

[0481] Step 8:

[0482] The device (user) collects emotional data using a camera, microphone, or other sensor devices. For example, the camera is used to capture the user's facial expressions, and the microphone is used to collect the tone of their voice.

[0483] Step 9:

[0484] The server analyzes the collected emotional data. Using machine learning algorithms, it analyzes the user's facial expressions and tone of voice to recognize the user's emotional state. For example, it can determine from facial expressions whether the user is excited or unhappy.

[0485] Step 10:

[0486] The server dynamically changes the system interface based on the analysis results. For example, if the user looks confused, it displays an operation guide. If the user looks surprised, it highlights information that might interest them.

[0487] This system efficiently constructs digital twins of cities, making them accessible to numerous stakeholders, while also enhancing the user experience through an emotion engine. Rapid 3D map generation and the integration of emotion data expand the usability of urban digital twins, allowing stakeholders to quickly begin exploring use cases. This differentiates the system from other existing systems.

[0488] (Example 2)

[0489] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0490] Conventional urban digital twin systems required improved efficiency in collecting geographic information data and generating 3D maps, but they lacked the ability to recognize user emotions in real time and reflect them in the interface. Therefore, it was difficult to provide users with an intuitive and personalized user experience. Furthermore, it was challenging to provide the generated 3D maps in real time while maintaining the consistency of the integrated data.

[0491] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0492] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, means for users to access, view, and manipulate the generated 3D map, means for collecting sentiment data, means for analyzing the collected sentiment data and recognizing the user's sentiment in real time, and means for dynamically adjusting the interface based on the recognized sentiment data. This enables the efficient construction of a city digital twin and the provision of an intuitive and personalized user experience that takes into account the user's sentiment.

[0493] "Geographic information data" refers to a dataset containing geographical location information, providing information about elements such as buildings, roads, and topography.

[0494] "Preprocessing" is the process of ensuring the consistency of collected data by performing tasks such as noise reduction, format conversion, and coordinate system standardization.

[0495] A "generative model" is an algorithm that generates new data based on input data, and often refers specifically to machine learning algorithms.

[0496] A "trained generative model" is a generative model that has been learned using training data and optimized to perform a specific task.

[0497] A "3D map" is a visualization of geographic information data in three-dimensional space, and is a data structure designed to realistically reproduce real-world geographical features.

[0498] A "cloud storage service" is a remote storage service that allows you to store and manage data via the internet.

[0499] "Emotional data" refers to data that indicates a user's emotional state, collected based on factors such as the user's facial expressions and tone of voice.

[0500] "Recognizing in real time" means performing the entire process from data collection to analysis immediately, and providing results without any time delay.

[0501] An "interface" refers to the screen or control panel that a user uses to interact with a system, and is the point of contact that provides a user experience.

[0502] "Dynamic adjustment" refers to changing the system's display content and behavior in real time according to the user's situation and actions.

[0503] Modes for carrying out the invention

[0504] As an embodiment of this invention, we will describe the specific operation of a system that combines a city digital twin construction system using generative AI with an emotion engine that recognizes user emotions. This system consists of the following steps.

[0505] First, the server collects geographic information data for cities. Specifically, it accesses geographic information databases and uses APIs to obtain building data, road data, and terrain data for cities. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools. For example, it can use open-source geographic information databases (e.g., OpenStreetMap API). It can also obtain the latest satellite imagery through NASA's satellite imagery service. Furthermore, it is possible to scrape the necessary image data using the Python library BeautifulSoup.

[0506] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, it uses Python's NumPy and Pandas libraries to check data consistency and remove unnecessary information or incorrect data points. It also uses the GeoPandas library to unify the coordinate system of the collected data to WGS84 (e.g., EPSG:4326) and convert it to GeoJSON format.

[0507] Next, the server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset, and the model is built using the Generative Adversarial Network (GAN) algorithm. The server uses libraries such as TensorFlow and PyTorch to input the manually created 3D map data into the AI ​​model and runs a program to learn the shapes of buildings and roads. This process is repeated until the model's accuracy is sufficiently high.

[0508] Next, the server generates a 3D map using a trained generative AI model. For example, new geographic information data (building data, road data, etc.) is input into the generative AI model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is stored in an internal database. By using geographic information databases such as PostGIS, efficient data storage and retrieval are possible.

[0509] Next, the server integrates the generated 3D map with existing map data. It retrieves supplementary data (e.g., tree and sidewalk data) from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to eliminate duplication and inconsistencies, completing it as the final digital twin data.

[0510] Next, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. In this step, the integrated 3D map is uploaded to the cloud (for example, Amazon S3). To set up a REST API and notify stakeholders how to access it, an API endpoint can be built using the Flask framework.

[0511] Furthermore, this system incorporates an emotion engine that recognizes the user's emotions. The terminal (user) collects emotion data through a camera, microphone, or other sensor devices. For example, the camera is used to capture the user's facial expressions and the image data is sent to the server. The microphone is also used to collect audio data, capturing the user's speech and tone of voice.

[0512] The server analyzes collected emotional data and uses machine learning algorithms to recognize the user's emotions in real time. Specifically, it analyzes facial expressions using OpenCV and Dlib libraries, and analyzes audio data using TensorFlow, etc., to determine the user's emotional state. For example, if the user has a surprised expression, it recognizes an excited state, and if they have a grumpy expression, it understands the situation.

[0513] The recognized emotion data is reflected as feedback in the system interface. For example, if the user shows a surprised expression, the system adjusts the displayed content to emphasize information that will interest the user. Also, if the user shows a confused expression, the system makes dynamic changes such as displaying an operation guide.

[0514] Finally, the terminal (user) accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user accesses the provided URL through a browser and logs in using the authentication information. It provides interactive functionality to manipulate the 3D map, making it possible, for example, to survey the height of buildings and the width of roads in a specific area for real estate development planning. The emotion engine provides real-time feedback of the user's emotional data, offering an intuitive and personalized user experience.

[0515] As a concrete example, the following is an example of a prompt statement.

[0516] Example of a prompt:

[0517] Please describe a program that acquires satellite imagery and geographic information data, preprocesses it with noise reduction and format conversion, generates a 3D map using an AI model, integrates it with existing data, and uploads it to cloud storage. Also, please describe an emotion engine that recognizes user emotions and provides real-time feedback.

[0518] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0519] Step 1:

[0520] The server collects urban geographic information data. Specifically, the server accesses a geographic information database and retrieves building data, road data, and terrain data via an API. The input requires the database's API endpoint and request parameters, while the output is datasets of buildings, roads, and terrain. In addition, the server downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools. The BeautifulSoup library in Python is used for these operations.

[0521] Step 2:

[0522] The server preprocesses the collected geographic information data. The specific inputs are building data, road data, terrain data, and image data collected in Step 1. First, the data consistency is checked using Python's NumPy and Pandas libraries, and unnecessary information and incorrect data points are filtered out. Next, the GeoPandas library is used to unify the data coordinate system to WGS84 (EPSG:4326) and convert it to GeoJSON format. The output is preprocessed geographic information data in a unified format.

[0523] Step 3:

[0524] The server trains a generative AI model. The specific inputs are collected and pre-processed geographic data and manually created 3D map datasets. The server uses TensorFlow and PyTorch libraries to train a generative adversarial network (GAN) algorithm with these inputs. The datasets are fed into the AI ​​model to learn the shapes of buildings and roads. This process is repeated iteratively until the model's accuracy is sufficiently high. The output is the trained AI model.

[0525] Step 4:

[0526] The server generates 3D maps using a trained generative AI model. The specific inputs are the trained AI model and new geographic data (building data, road data, etc.). This data is input to the model, which then predicts and generates the 3D shape. The output is the generated 3D map data, which is stored in an internal database (e.g., PostGIS).

[0527] Step 5:

[0528] The server integrates the generated 3D map with existing map data. The specific inputs are the generated 3D map data and supplementary data (tree data, footpath data, etc.) obtained from existing map databases. The server executes scripts to integrate these, resolving duplicate data and inconsistencies. The output is integrated, consistent digital twin data.

[0529] Step 6:

[0530] The server uploads digital twin data to a cloud storage service and distributes it to stakeholders. The specific input is integrated digital twin data. The server uploads the data to a cloud storage service such as Amazon S3. Furthermore, it configures a REST API using the Flask framework and exposes an API endpoint. The output is the digital twin data on cloud storage and the API endpoint.

[0531] Step 7:

[0532] The device (user) collects emotional data using sensor devices such as cameras and microphones. Specific inputs include the user's facial expressions and voice data. The device captures the user's facial expressions through the camera and collects image data. It also collects voice data using the microphone. The output is the collected emotional data.

[0533] Step 8:

[0534] The server analyzes collected emotion data and recognizes the user's emotions in real time. The specific input is emotion data (facial expressions, voice data, etc.) sent from the terminal. The server analyzes facial expressions using OpenCV and Dlib libraries, and analyzes voice data using TensorFlow, etc., to determine the user's emotional state. The output is the recognized emotion data.

[0535] Step 9:

[0536] The server dynamically adjusts the interface based on recognized emotion data. The specific inputs are the recognized emotion data and the data of the currently displayed interface. The server executes a program that adjusts the displayed content based on the user's emotional state. For example, if the user shows a surprised expression, the system highlights interesting information. If the user is confused, it displays an operation guide. The output is the adjusted interface.

[0537] Step 10:

[0538] The user accesses the provided URL and enters their authentication information to access the 3D map. The specific inputs are the provided URL and authentication information. The user accesses the URL using a web browser and logs in using the authentication information. The user can interactively manipulate the 3D map to investigate building heights and road widths in specific areas for real estate development planning. The output is the interactively manipulated 3D map.

[0539] (Application Example 2)

[0540] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0541] In modern industrial settings, it is crucial to monitor the operation of factory robots in real time and direct them to perform optimal actions. However, current systems have limited means of accurately understanding robot operation, and there is a lack of means to improve the user experience. Furthermore, there are currently no systems that improve work efficiency by analyzing the user's emotional state and providing feedback. It is desirable to provide a system that solves these problems and manages robot operations more effectively.

[0542] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0543] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, means for users to access, view, and manipulate the generated 3D map, means for collecting and analyzing user sentiment data in real time, and means for adjusting the user experience based on the analyzed sentiment data. This makes it possible to monitor the operating status of robots in a factory in real time, analyze the emotional state of users, and provide optimal feedback.

[0544] "Geographic information data" refers to data that includes geographical information such as buildings, roads, and topography within cities and factories.

[0545] "Preprocessing" refers to data preparation tasks such as denoising, consistency checking, and format conversion of collected data.

[0546] A "generative model" is an artificial intelligence model that is trained to generate new data based on collected data.

[0547] A "3D map" is a three-dimensional map generated based on geographic information data, representing buildings, roads, and other topographic elements in 3D.

[0548] A "cloud storage service" is a service that allows you to store and manage data via the internet and access it when needed.

[0549] "User" refers to an individual or organization that uses the system to view and manipulate 3D maps.

[0550] "Emotional data" refers to data about a user's emotional state obtained from their facial expressions, voice, and other sources.

[0551] "Analysis" is the process of extracting and understanding information based on collected data.

[0552] An "algorithm" is a method for solving a problem through a series of steps or calculations.

[0553] "Integration" is the process of combining multiple datasets into a single, consistent dataset.

[0554] As an embodiment of this invention, a system for real-time monitoring and control of robot movements within a factory is provided. The system combines a factory digital twin construction system using a generative AI model with an emotion engine that recognizes user emotions. The system operates in the following specific steps.

[0555] First, the server collects geographical information data within the factory. Specifically, it acquires robot location data, operating status, and 3D scan data of the environment. This includes collecting data from sensor devices and cameras within the factory.

[0556] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, it runs a denoising program to filter out unnecessary information from the collected data and converts all data to the WGS84 coordinate system to unify the coordinate system. Furthermore, it converts all data to the GeoJSON format for unification.

[0557] Next, the server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset, and the Generative Adversarial Network (GAN) algorithm is configured. The training data is input into the AI ​​model, and the model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0558] Next, the server generates a 3D map using a trained generative AI model. New geographic data is input into the generative model, which then predicts and generates the 3D shape from this data. The generated 3D map data is stored in an internal database.

[0559] Next, the server integrates the generated 3D map with existing map data. It retrieves supplementary data from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to remove duplicates and inconsistencies for consistency, creating the final digital twin data.

[0560] Furthermore, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. The data uploaded to the cloud is then configured with a REST API to expose an endpoint. Stakeholders are notified of the API endpoint and how to access it.

[0561] In addition, this system incorporates an emotion engine that recognizes the user's emotions. The terminal (user) collects emotion data through cameras, microphones, and other sensor devices. For example, it uses the smartphone's camera and microphone to capture facial expressions and voice. The server analyzes the collected emotion data and uses machine learning algorithms to recognize the user's emotions in real time. Specifically, by analyzing the user's facial expressions and voice, it can identify emotional states such as whether the user is excited or unhappy.

[0562] The recognized emotion data is reflected as feedback in the system interface. For example, if the user shows a surprised expression, the system can adjust the displayed content to emphasize information that is of interest to the user. Also, if the user shows a confused expression, the system will make dynamic changes such as displaying an operation guide.

[0563] Finally, the terminal (user) accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map, for example, to monitor and control the movements of robots working in a specific area of ​​a factory. In addition, the emotion engine provides feedback on the user's emotional data, offering a more intuitive and personalized user experience.

[0564] This system allows for real-time monitoring of robot movements within the factory, analysis of user emotional states, and provision of optimal feedback, thereby improving both work efficiency and the user experience.

[0565] Example of a prompt

[0566] "Develop an application that generates a 3D digital twin model of a factory robot and displays and controls it in real time on a smartphone. Additionally, add a function to analyze the user's facial expressions and voice to provide feedback."

[0567] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0568] Step 1:

[0569] The server collects geographical information data within the factory. Input data includes robot position data, operating status, and 3D scan data of the environment, obtained from sensor devices and cameras. The server receives this data and stores it in a database.

[0570] Step 2:

[0571] The server preprocesses the collected geographic information data. Specifically, it denoises the input data, checks for data consistency, and converts it to GeoJSON format. This process removes noise and outputs a dataset in a consistent format.

[0572] Step 3:

[0573] The server trains a generative AI model. It takes manually created 3D map data as the training dataset and uses a generative adversarial network (GAN) algorithm to train the model. This training process is repeated until the model's accuracy reaches a sufficient level, and a trained model with optimal parameters is output.

[0574] Step 4:

[0575] The server generates 3D maps using a trained generative AI model. New geographic data is input into the model, which then predicts and generates 3D shapes from this data. The generated 3D map data is stored in a database.

[0576] Step 5:

[0577] The server integrates the generated 3D map with existing map data. It retrieves supplementary data from the existing map data and runs an integration script to integrate it with the 3D map. To ensure consistency, duplicates and inconsistencies are resolved, and the final digital twin data is output.

[0578] Step 6:

[0579] The server uploads the digital twin data to a cloud storage service. A REST API is configured for the uploaded data, and an endpoint is exposed. Stakeholders can access the digital twin data using the notified API endpoint.

[0580] Step 7:

[0581] The device (user) uses a camera and microphone to collect emotional data. The user's facial expressions and voice are captured as input data and sent to the server.

[0582] Step 8:

[0583] The server analyzes the received emotional data. Using machine learning algorithms, it recognizes the user's emotional state from the input data. As a result of the analysis, it outputs the emotional state, such as whether the user is excited or unhappy.

[0584] Step 9:

[0585] The server adjusts the system interface based on the analyzed emotion data. For example, if the user has a surprised expression, the system displays information to highlight information that will interest the user. If the user appears confused, dynamic changes are made, such as displaying an operation guide.

[0586] Step 10:

[0587] The user accesses the provided URL, enters their authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map and monitor and control the movements of robots within the factory in real time. Furthermore, they can enjoy a more intuitive and personalized operating experience by utilizing emotional data fed back by the emotion engine.

[0588] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0589] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0590] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0591] [Third Embodiment]

[0592] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0593] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0594] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0595] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0596] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0597] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0598] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0599] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0600] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0601] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0602] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0603] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0604] As an embodiment of this invention, the specific operation of a city digital twin construction system using generative AI will be described. This system consists of the following steps.

[0605] First, the server collects geographic information data for the city. Specifically, it retrieves data such as buildings, roads, land use, and topography from geographic information databases, and also downloads publicly available aerial and satellite imagery. This geographic information data includes formats such as Shapefile and GeoJSON.

[0606] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, a denoising program is run to filter out unnecessary information from the collected data, and all data is converted to the WGS84 coordinate system to unify the coordinate system. Furthermore, all data is converted to and unified in GeoJSON format.

[0607] Next, the server trains the generative AI model. For this, manually created 3D map data is used as the training dataset. For the generative model, for example, a Generative Adversarial Network (GAN) is used. During training, the AI ​​model learns features such as the shapes of buildings and roads, enabling the generation of highly accurate 3D maps.

[0608] Next, the server generates a 3D map using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model generates a predicted 3D shape from this data. The generated 3D map is stored in the server's internal database.

[0609] Next, the server performs a process to integrate the generated 3D map with existing map data. This process involves running scripts to retrieve supplementary data (e.g., trees and walkways) from the existing database and integrate it with the generated 3D map. The integrated data undergoes quality checks, and after correcting inconsistencies and duplicate data, a single, consistent digital twin is created.

[0610] Next, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. In this step, the integrated 3D map is uploaded to the cloud, and a REST API is configured to expose the endpoint. Stakeholders are notified of the API endpoint and how to access it.

[0611] Finally, the terminal (user) accesses the cloud server and uses the generated 3D map. For example, the user can open a web browser, access the provided URL, enter their authentication information, and access the 3D map from the interface. The user can interactively manipulate the 3D map to check the height of buildings and the width of roads in a specific area.

[0612] This system enables the efficient creation of digital twins of cities, making them accessible to a wide range of stakeholders. Furthermore, rapid 3D map generation allows stakeholders to quickly begin exploring use cases. This differentiates the system from the Ministry of Land, Infrastructure, Transport and Tourism's PLATEAU system.

[0613] The following describes the processing flow.

[0614] Step 1:

[0615] The server collects urban geographic information data. Specifically, it accesses geographic information databases and uses APIs to obtain urban building data, road data, and topographic data. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools.

[0616] Step 2:

[0617] The server preprocesses the collected geographic information data. First, it runs a denoising program to filter out unnecessary information and erroneous data points from the collected data. Next, it converts all data to the WGS84 coordinate system to unify the data's coordinate system. Finally, it converts each dataset to GeoJSON format to unify the data.

[0618] Step 3:

[0619] The server trains the generative AI model. It prepares manually created 3D map data as the training dataset and configures the Generative Adversarial Network (GAN) algorithm. The training data is input into the AI ​​model, and the model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0620] Step 4:

[0621] The server generates 3D maps using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is stored in an internal database.

[0622] Step 5:

[0623] The server integrates the generated 3D map with existing map data. It retrieves supplementary data (e.g., tree and footpath data) from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to eliminate redundancies and inconsistencies for consistency, creating the final digital twin data.

[0624] Step 6:

[0625] The server uploads the integrated digital twin data to a cloud storage service. A REST API is configured to access the uploaded data, and the endpoint is exposed. Furthermore, stakeholders are notified of the API endpoint and how to access it.

[0626] Step 7:

[0627] The terminal (user) accesses a cloud server and utilizes the generated 3D map. The user opens a web browser, accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map and, for example, investigate the height of buildings and the width of roads in a specific area for real estate development planning.

[0628] (Example 1)

[0629] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0630] Providing highly accurate geographic information is essential in modern urban planning and infrastructure management. However, traditional methods involve significant time and effort in collecting and preprocessing geographic data, and generating and integrating 3D models. As a result, providing timely data is difficult, hindering rapid decision-making and action planning by stakeholders. Furthermore, unifying data in different formats and building a consistent digital twin requires advanced technology and expertise. A new system is needed to solve these problems.

[0631] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0632] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a three-dimensional map using the trained generative model, means for integrating with existing map data, means for uploading the generated three-dimensional map to a cloud storage service and distributing it to users, and means for users to access, view, and manipulate the generated three-dimensional map. This enables the rapid and efficient collection and preprocessing of highly accurate geographic information, and the generation and integration of highly accurate three-dimensional maps in a short time using a generative AI model. Furthermore, by utilizing cloud storage, stakeholders can easily access and interact with the data.

[0633] "Geographic information data" refers to data that includes spatial and locational information about urban and regional features (buildings, roads, land use, topography, etc.).

[0634] "Preprocessing" refers to the process of removing noise from collected geographic information data, checking for consistency, unifying coordinate systems, and converting formats.

[0635] A "generative model" is a model that uses machine learning algorithms to generate a desired output (for example, a three-dimensional map) from input data.

[0636] A "three-dimensional map" is map data that visualizes the shape and location information of features in a city or region in three-dimensional space.

[0637] A "cloud storage service" is a service that stores and manages data online, allowing users to access it via the internet.

[0638] A "user" is a person or organization that has the authority to access the generated three-dimensional map and to view and manipulate its information.

[0639] A Generative Adversarial Network (GAN) is a type of generative model in which two paired neural networks compete to generate data.

[0640] A "digital twin" is a virtual model that accurately replicates a physical, real-world object or system, reflecting real-time data.

[0641] A "REST API" is a type of interface for providing web services, and it is a set of design principles for manipulating resources using the HTTP protocol.

[0642] An "API endpoint" is a URL or URI used to access a specific resource or function designated through an API.

[0643] Modes for carrying out the invention

[0644] As an embodiment of this invention, the specific operation of a city digital twin construction system using generative AI is shown below. This system collects and preprocesses geographic information of a city, creates and integrates a three-dimensional map using a generative model, and provides it via a cloud storage service.

[0645] First, the server collects geographic data of the city. Specifically, it uses the OpenStreetMap API to obtain data on buildings, roads, land use, and topography, and downloads the latest satellite imagery from NASA's Landsat satellites. This geographic data includes formats such as Shapefile and GeoJSON. This allows for obtaining up-to-date and detailed geographic information about the city.

[0646] Next, the server preprocesses the collected geographic data. This process includes denoising, consistency checking, and format conversion using the Python Geopandas library. Specifically, Geopandas is used to convert the coordinate system to WGS84, filter out unnecessary data, and unify all data by converting it to GeoJSON format. This process results in consistent geographic data.

[0647] Next, the server trains the generative AI model. Detailed, manually created 3D map data is used as the training dataset. A Generative Adversarial Network (GAN) is used as the generative model. During training, the AI ​​model learns features such as the shapes of buildings and roads, enabling the generation of highly accurate 3D maps. The TensorFlow library is used to train the generative model.

[0648] Next, the server uses a trained generative AI model to generate a 3D map based on new geographic data. New building and road data are input into the generative model, which then generates a predicted 3D shape from this data. The generated 3D map is stored in the server's internal database. This procedure ensures that the most up-to-date 3D map data is always generated.

[0649] Subsequently, the server integrates the generated 3D map with existing map data. This process involves retrieving supplementary data such as trees and footpaths from the existing database and running an integration script. Using the integration script, the generated 3D map and the existing map data are combined into a single, consistent digital twin. After integration, quality checks are also performed to correct inconsistencies and duplicate data.

[0650] Next, the server uploads the integrated 3D map data to a cloud storage service and distributes it to users. In this step, the data is uploaded to an Amazon S3 bucket, and a REST API is configured to expose the endpoint. Stakeholders are notified of the API endpoint and how to access it. This allows stakeholders to access the updated 3D map at any time.

[0651] Finally, the terminal (user) accesses the provided API endpoint to view and interact with the generated 3D map. The user opens a web browser, accesses the provided URL, and enters their authentication information. Through the interface, they can access the 3D map and check the height of buildings and the width of roads in a specific area. This system allows the user to utilize detailed 3D map information that is updated in real time.

[0652] As a concrete example, the server retrieves city building data using the OpenStreetMap API and preprocesses it using the latest satellite imagery downloaded from NASA's Landsat satellites. Then, it trains a GAN model in TensorFlow using manually created detailed 3D map data of Tokyo. Using the new geographic data as input, the trained model generates a highly accurate 3D map, which is then integrated with the existing data. The integrated data is uploaded to an Amazon S3 bucket, and an API endpoint is configured. Stakeholders can access the provided URL to view the 3D map and obtain the necessary information.

[0653] Examples of prompt statements include the following:

[0654] I want to generate 3D map data of a city. Please use the following geographic information data as input to generate detailed 3D shapes of buildings, roads, trees, etc.

[0655] Building data (Shapefile) obtained from OpenStreetMap

[0656] NASA Landsat satellite imagery (GeoTIFF)

[0657] Tree data (GeoJSON)

[0658] Based on the data above, please generate a high-precision 3D map.

[0659] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0660] Step 1:

[0661] The server collects geographic data of cities. Specifically, it uses the OpenStreetMap API to retrieve building and road data and downloads the latest satellite imagery from NASA's Landsat satellites. It takes API endpoints and query parameters as input and receives geographic data in Shapefile or GeoJSON format as output. This provides up-to-date and detailed geographic information of cities.

[0662] Step 2:

[0663] The server preprocesses the collected geographic data. Specifically, it uses the Python Geopandas library to convert the data's coordinate system to WGS84 and filters out unnecessary data. Using the diverse formats of geographic data collected as input, it obtains data in a unified GeoJSON format as output. This process ensures a consistent dataset. For example, to perform a coordinate system transformation using Geopandas, execute geo_df.to_crs(epsg=4326).

[0664] Step 3:

[0665] The server trains a generative AI model. Specifically, it trains a generative adversarial network (GAN) using TensorFlow with manually created 3D map data. Detailed, manually created 3D map data is used as input, and a trained generative AI model is obtained as output. This process allows the AI ​​model to learn the features of buildings and roads, enabling highly accurate 3D map generation. For example, the command `model.fit(training_data, epochs=50)` is used.

[0666] Step 4:

[0667] The server generates a 3D map using a trained generative AI model. Specifically, it takes new geographic data as input and generates the 3D shape predicted by the model. It uses newly collected and pre-processed geographic data as input and produces generated 3D map data as output. This output is stored in the server's internal database. For example, `model.predict(new_data)` is executed to generate the 3D shape.

[0668] Step 5:

[0669] The server integrates the generated 3D map with existing map data. Specifically, it retrieves supplementary data from the existing database and runs an integration script. It uses the new 3D map data and supplementary data (e.g., tree and sidewalk data) as input and obtains a consistent digital twin as output. By running the integration script, the data is integrated using the command integrate_data(new_3d_map, additional_data).

[0670] Step 6:

[0671] The server uploads the integrated 3D map to a cloud storage service and distributes it to stakeholders. Specifically, it uploads the data to an Amazon S3 bucket, configures a REST API, and exposes an endpoint. It uses the integrated data as input and obtains the data stored in the cloud and the API endpoint as output. For example, it uses the command `s3_client.upload_file('integrated_3d_map.json', 'bucket-name', 'path / in / bucket')`.

[0672] Step 7:

[0673] The user accesses the provided API endpoint to view and manipulate the generated 3D map. For example, the user opens a web browser, accesses the provided URL, and enters their authentication information. Using the specified URL and authentication information as input, they obtain an interactive display of the 3D map as output. This allows the user to check the height of buildings and the width of roads in a specific area.

[0674] (Application Example 1)

[0675] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0676] In recent years, in order to solve the increasingly complex traffic problems that accompany urban development, there is a need for real-time, high-precision map information and traffic data-driven operational management of autonomous vehicles. However, current systems have slow map information updates, making it difficult to respond quickly to changes in traffic conditions. Furthermore, the integration of information from multiple data sources is insufficient, making it difficult to provide optimal route guidance. This hinders the improvement of efficiency and safety of autonomous vehicles. This invention aims to solve these problems.

[0677] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0678] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, means for users to access, view, and manipulate the generated 3D map, and means for calculating the optimal route for an autonomous vehicle and performing operational management. This enables efficient operational management of autonomous vehicles based on highly accurate and real-time map information and traffic data.

[0679] "Geographic information data" refers to data that includes information about urban structure, land use, buildings, roads, topography, and so on.

[0680] "Preprocessing" is the process of removing noise from collected data, checking data consistency, and standardizing the format.

[0681] A "generative model" is an algorithm for generating 3D maps through learning, and it utilizes machine learning techniques.

[0682] A "3D map" is a digital map that represents urban structures such as buildings, roads, and terrain in three dimensions.

[0683] A "cloud storage service" is a data storage system that allows data to be stored via the internet and shared among multiple users.

[0684] A "user" is an end-user who utilizes the data and services provided by this system.

[0685] An "autonomous vehicle" is a vehicle that uses sensors and AI to automate its driving process.

[0686] An "optimal route" is a path that minimizes travel time and distance, taking into account traffic conditions and geographical features.

[0687] "Operational management" refers to the process of monitoring and adjusting the operating schedule, routes, and traffic conditions of autonomous vehicles.

[0688] The specific system configuration and operation for realizing this invention are described below.

[0689] Hardware and Software Overview

[0690] This invention utilizes the following main hardware and software.

[0691] hardware

[0692] Server: A server equipped with a high-performance CPU and large-capacity storage.

[0693] Autonomous vehicle onboard computer: Onboard computer equipped with a GPU.

[0694] Device: Smartphone (Android / iOS).

[0695] software

[0696] Cloud storage services: such as AWS S3 and Google Cloud Storage.

[0697] Traffic information APIs: such as Google Maps API and Here API.

[0698] Generative AI model: GAN (Generative Adversarial Network).

[0699] Data processing programs: Python, TensorFlow / PyTorch.

[0700] Explanation of program processing

[0701] The server collects geographic information data. This data is obtained from geographic information databases, aerial photographs, satellite imagery, etc. The collected geographic information data includes formats such as Shapefile and GeoJSON.

[0702] The server preprocesses the collected data. It performs processes such as noise reduction, data consistency checks, and format conversion, and converts all data into a unified format (GeoJSON).

[0703] To train the generative AI model, manually created 3D map data is used as the training dataset. A Generative Adversarial Network (GAN) is used as the generative model to learn features such as the shapes of buildings and roads.

[0704] A pre-trained generative AI model is used to generate 3D maps from new geographic data. These generated 3D maps are stored in the server's internal database.

[0705] Next, the generated 3D map is integrated with existing map data. A script is executed to retrieve supplementary data (e.g., trees and sidewalks) from the existing database and integrate it with the generated 3D map.

[0706] The integrated 3D map data is uploaded to a cloud storage service, and a REST API is configured to expose the endpoint. Users are notified of the API endpoint and how to access it.

[0707] Using a smartphone as a terminal, users access a cloud server and utilize the generated 3D map data. For example, a user can open a web browser, access the provided URL, enter authentication information, and access the 3D map from the interface.

[0708] Operation management of autonomous vehicles

[0709] Furthermore, this system acquires traffic information in real time and calculates the optimal route for autonomous vehicles. Traffic information is collected from traffic sensors, cameras, and social media.

[0710] The terminal's AI algorithm integrates generated 3D map data with real-time traffic information to calculate the optimal route. The calculated route is transmitted to the autonomous vehicle's onboard computer, and the vehicle operates according to that route.

[0711] Users can visually check route information and traffic conditions through the app's interface. For example, they can receive route guidance by entering the following prompt:

[0712] Example of a prompt

[0713] "Please tell me the best route from Tokyo Station to Shinjuku Station."

[0714] Upon receiving this prompt, the system can use the collected data and generated AI models to calculate the optimal route in real time and provide it to the user.

[0715] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0716] Step 1:

[0717] The server collects geographic information data. It retrieves data such as buildings, roads, land use, and topography from geographic information databases, and downloads publicly available aerial and satellite imagery. Input is data in Shapefile or GeoJSON format, and output is the collected raw data.

[0718] Step 2:

[0719] The server preprocesses the geographic information data it collects. A noise reduction program is run to filter out unnecessary information from the collected data, and all data is converted to the WGS84 coordinate system to unify the coordinate system. Furthermore, all data is converted to the GeoJSON format for standardization. The input is the collected raw data, and the output is the clean data after preprocessing.

[0720] Step 3:

[0721] The server trains a generative AI model. Manually created 3D map data is used as the training dataset, and a generative adversarial network (GAN) is used to train the model. The input is the training dataset, and the output is the trained generative AI model.

[0722] Step 4:

[0723] The server generates a 3D map using a trained generative AI model. New geographic data is input to the generative model, and the model generates a predicted 3D shape from this data. The input is new geographic data, and the output is the generated 3D map.

[0724] Step 5:

[0725] The server integrates the generated 3D map with existing map data. A script is executed to retrieve supplementary data (e.g., trees and sidewalks) from the existing database and integrate it with the generated 3D map. The input is the generated 3D map and supplementary data, and the output is the integrated 3D map.

[0726] Step 6:

[0727] The server uploads the integrated 3D map data to a cloud storage service, configures a REST API, and exposes an endpoint. The input is the integrated 3D map, and the output is the data on the cloud storage and the exposed API endpoint.

[0728] Step 7:

[0729] The user accesses the cloud server using their device (smartphone). They access the provided URL, enter their authentication information, and access the 3D map from the interface. The input is the user's authentication information, and the output is the 3D map display on the interface.

[0730] Step 8:

[0731] The server collects real-time traffic information from traffic sensors, cameras, social media, and other sources. The input is real-time traffic data, and the output is collected traffic information.

[0732] Step 9:

[0733] The terminal's AI algorithm integrates generated 3D map data with real-time traffic information to calculate the optimal route. The calculated route is transmitted to the autonomous vehicle's onboard computer, and the vehicle operates according to that route. The input is 3D map data and traffic information, and the output is the optimal route and its operation information.

[0734] Step 10:

[0735] Users visually check route information and traffic conditions through the app's interface. For example, they can receive route guidance by prompting them with a phrase like, "Tell me the best route from Tokyo Station to Shinjuku Station." The input is the user's prompt, and the output is the route information displayed on the screen.

[0736] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0737] As an embodiment of this invention, we will describe the specific operation of a system that combines a city digital twin construction system using generative AI with an emotion engine that recognizes user emotions. This system consists of the following steps.

[0738] First, the server collects geographic information data for cities. Specifically, it accesses geographic information databases and uses APIs to obtain building data, road data, and topographic data for cities. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools.

[0739] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, a denoising program is run to filter out unnecessary information and erroneous data points from the collected data, and all data is converted to the WGS84 coordinate system to unify the coordinate system. Furthermore, all data is converted and unified to the GeoJSON format.

[0740] Next, the server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset, and the Generative Adversarial Network (GAN) algorithm is configured. The training data is input into the AI ​​model, and the model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0741] Next, the server generates a 3D map using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is saved to an internal database.

[0742] Next, the server integrates the generated 3D map with existing map data. It retrieves supplementary data (e.g., tree and sidewalk data) from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to eliminate redundancies and inconsistencies for consistency, creating the final digital twin data.

[0743] Next, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. In this step, the integrated 3D map is uploaded to the cloud, and a REST API is configured to expose the endpoint. Stakeholders are notified of the API endpoint and how to access it.

[0744] Furthermore, this system incorporates an emotion engine that recognizes the user's emotions. The terminal (user) collects emotion data through a camera, microphone, or other sensor devices. The server analyzes the collected emotion data and uses machine learning algorithms to recognize the user's emotions in real time. For example, by analyzing the user's facial expressions, it is possible to identify the user's emotional state, such as whether they are excited or unhappy.

[0745] The recognized emotion data is reflected as feedback in the system interface. For example, if the user shows a surprised expression, the system can adjust the displayed content to emphasize information that is of interest to the user. Also, if the user shows a confused expression, the system will make dynamic changes such as displaying an operation guide.

[0746] Finally, the device (user) accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map and, for example, investigate the height of buildings and the width of roads in a specific area for real estate development planning. In addition, the emotion engine provides feedback on the user's emotional data, offering a more intuitive and personalized user experience.

[0747] This system enables the efficient construction of digital twins of cities, making them accessible to a wide range of stakeholders, while also enhancing the user experience through an emotion engine. Furthermore, rapid 3D map generation and the integration of emotion data expand the usability of urban digital twins, allowing stakeholders to quickly begin exploring use cases. This differentiates the system from the Ministry of Land, Infrastructure, Transport and Tourism's PLATEAU.

[0748] The following describes the processing flow.

[0749] Step 1:

[0750] The server collects urban geographic information data. Specifically, it accesses geographic information databases and uses APIs to obtain urban building data, road data, and topographic data. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools.

[0751] Step 2:

[0752] The server preprocesses the collected geographic information data. First, it runs a denoising program to filter out unnecessary information and erroneous data points from the collected data. Next, it converts all data to the WGS84 coordinate system to unify the data's coordinate system. Finally, it converts each dataset to GeoJSON format to unify the data.

[0753] Step 3:

[0754] The server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset. The Generative Adversarial Network (GAN) algorithm is configured, and the training data is input into the AI ​​model. The AI ​​model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0755] Step 4:

[0756] The server generates 3D maps using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is stored in an internal database.

[0757] Step 5:

[0758] The server integrates the generated 3D map with existing map data. It runs a script that retrieves supplementary data from the existing map database and integrates it with the 3D map. The integrated data is then processed to eliminate duplication and inconsistencies, creating a consistent digital twin.

[0759] Step 6:

[0760] The server uploads the integrated digital twin data to a cloud storage service. A REST API is configured to access the data, and the endpoint is exposed. Stakeholders are notified of the API endpoint and how to access it.

[0761] Step 7:

[0762] The user accesses a cloud server and utilizes the generated 3D map. The user opens a web browser, accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user interacts with the 3D map and investigates information about a specific area.

[0763] Step 8:

[0764] The device (user) collects emotional data using a camera, microphone, or other sensor devices. For example, the camera is used to capture the user's facial expressions, and the microphone is used to collect the tone of their voice.

[0765] Step 9:

[0766] The server analyzes the collected emotional data. Using machine learning algorithms, it analyzes the user's facial expressions and tone of voice to recognize the user's emotional state. For example, it can determine from facial expressions whether the user is excited or unhappy.

[0767] Step 10:

[0768] The server dynamically changes the system interface based on the analysis results. For example, if the user looks confused, it displays an operation guide. If the user looks surprised, it highlights information that might interest them.

[0769] This system efficiently constructs digital twins of cities, making them accessible to numerous stakeholders, while also enhancing the user experience through an emotion engine. Rapid 3D map generation and the integration of emotion data expand the usability of urban digital twins, allowing stakeholders to quickly begin exploring use cases. This differentiates the system from other existing systems.

[0770] (Example 2)

[0771] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0772] Conventional urban digital twin systems required improved efficiency in collecting geographic information data and generating 3D maps, but they lacked the ability to recognize user emotions in real time and reflect them in the interface. Therefore, it was difficult to provide users with an intuitive and personalized user experience. Furthermore, it was challenging to provide the generated 3D maps in real time while maintaining the consistency of the integrated data.

[0773] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0774] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, means for users to access, view, and manipulate the generated 3D map, means for collecting sentiment data, means for analyzing the collected sentiment data and recognizing the user's sentiment in real time, and means for dynamically adjusting the interface based on the recognized sentiment data. This enables the efficient construction of a city digital twin and the provision of an intuitive and personalized user experience that takes into account the user's sentiment.

[0775] "Geographic information data" refers to a dataset containing geographical location information, providing information about elements such as buildings, roads, and topography.

[0776] "Preprocessing" is the process of ensuring the consistency of collected data by performing tasks such as noise reduction, format conversion, and coordinate system standardization.

[0777] A "generative model" is an algorithm that generates new data based on input data, and often refers specifically to machine learning algorithms.

[0778] A "trained generative model" is a generative model that has been learned using training data and optimized to perform a specific task.

[0779] A "3D map" is a visualization of geographic information data in three-dimensional space, and is a data structure designed to realistically reproduce real-world geographical features.

[0780] A "cloud storage service" is a remote storage service that allows you to store and manage data via the internet.

[0781] "Emotional data" refers to data that indicates a user's emotional state, collected based on factors such as the user's facial expressions and tone of voice.

[0782] "Recognizing in real time" means performing the entire process from data collection to analysis immediately, and providing results without any time delay.

[0783] An "interface" refers to the screen or control panel that a user uses to interact with a system, and is the point of contact that provides a user experience.

[0784] "Dynamic adjustment" refers to changing the system's display content and behavior in real time according to the user's situation and actions.

[0785] Modes for carrying out the invention

[0786] As an embodiment of this invention, we will describe the specific operation of a system that combines a city digital twin construction system using generative AI with an emotion engine that recognizes user emotions. This system consists of the following steps.

[0787] First, the server collects geographic information data for cities. Specifically, it accesses geographic information databases and uses APIs to obtain building data, road data, and terrain data for cities. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools. For example, it can use open-source geographic information databases (e.g., OpenStreetMap API). It can also obtain the latest satellite imagery through NASA's satellite imagery service. Furthermore, it is possible to scrape the necessary image data using the Python library BeautifulSoup.

[0788] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, it uses Python's NumPy and Pandas libraries to check data consistency and remove unnecessary information or incorrect data points. It also uses the GeoPandas library to unify the coordinate system of the collected data to WGS84 (e.g., EPSG:4326) and convert it to GeoJSON format.

[0789] Next, the server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset, and the model is built using the Generative Adversarial Network (GAN) algorithm. The server uses libraries such as TensorFlow and PyTorch to input the manually created 3D map data into the AI ​​model and runs a program to learn the shapes of buildings and roads. This process is repeated until the model's accuracy is sufficiently high.

[0790] Next, the server generates a 3D map using a trained generative AI model. For example, new geographic information data (building data, road data, etc.) is input into the generative AI model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is stored in an internal database. By using geographic information databases such as PostGIS, efficient data storage and retrieval are possible.

[0791] Next, the server integrates the generated 3D map with existing map data. It retrieves supplementary data (e.g., tree and sidewalk data) from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to eliminate duplication and inconsistencies, completing it as the final digital twin data.

[0792] Next, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. In this step, the integrated 3D map is uploaded to the cloud (for example, Amazon S3). To set up a REST API and notify stakeholders how to access it, an API endpoint can be built using the Flask framework.

[0793] Furthermore, this system incorporates an emotion engine that recognizes the user's emotions. The terminal (user) collects emotion data through a camera, microphone, or other sensor devices. For example, the camera is used to capture the user's facial expressions and the image data is sent to the server. The microphone is also used to collect audio data, capturing the user's speech and tone of voice.

[0794] The server analyzes collected emotional data and uses machine learning algorithms to recognize the user's emotions in real time. Specifically, it analyzes facial expressions using OpenCV and Dlib libraries, and analyzes audio data using TensorFlow, etc., to determine the user's emotional state. For example, if the user has a surprised expression, it recognizes an excited state, and if they have a grumpy expression, it understands the situation.

[0795] The recognized emotion data is reflected as feedback in the system interface. For example, if the user shows a surprised expression, the system adjusts the displayed content to emphasize information that will interest the user. Also, if the user shows a confused expression, the system makes dynamic changes such as displaying an operation guide.

[0796] Finally, the terminal (user) accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user accesses the provided URL through a browser and logs in using the authentication information. It provides interactive functionality to manipulate the 3D map, making it possible, for example, to survey the height of buildings and the width of roads in a specific area for real estate development planning. The emotion engine provides real-time feedback of the user's emotional data, offering an intuitive and personalized user experience.

[0797] As a concrete example, the following is an example of a prompt statement.

[0798] Example of a prompt:

[0799] Please describe a program that acquires satellite imagery and geographic information data, preprocesses it with noise reduction and format conversion, generates a 3D map using an AI model, integrates it with existing data, and uploads it to cloud storage. Also, please describe an emotion engine that recognizes user emotions and provides real-time feedback.

[0800] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0801] Step 1:

[0802] The server collects urban geographic information data. Specifically, the server accesses a geographic information database and retrieves building data, road data, and terrain data via an API. The input requires the database's API endpoint and request parameters, while the output is datasets of buildings, roads, and terrain. In addition, the server downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools. The BeautifulSoup library in Python is used for these operations.

[0803] Step 2:

[0804] The server preprocesses the collected geographic information data. The specific inputs are building data, road data, terrain data, and image data collected in Step 1. First, the data consistency is checked using Python's NumPy and Pandas libraries, and unnecessary information and incorrect data points are filtered out. Next, the GeoPandas library is used to unify the data coordinate system to WGS84 (EPSG:4326) and convert it to GeoJSON format. The output is preprocessed geographic information data in a unified format.

[0805] Step 3:

[0806] The server trains a generative AI model. The specific inputs are collected and pre-processed geographic data and manually created 3D map datasets. The server uses TensorFlow and PyTorch libraries to train a generative adversarial network (GAN) algorithm with these inputs. The datasets are fed into the AI ​​model to learn the shapes of buildings and roads. This process is repeated iteratively until the model's accuracy is sufficiently high. The output is the trained AI model.

[0807] Step 4:

[0808] The server generates 3D maps using a trained generative AI model. The specific inputs are the trained AI model and new geographic data (building data, road data, etc.). This data is input to the model, which then predicts and generates the 3D shape. The output is the generated 3D map data, which is stored in an internal database (e.g., PostGIS).

[0809] Step 5:

[0810] The server integrates the generated 3D map with existing map data. The specific inputs are the generated 3D map data and supplementary data (tree data, footpath data, etc.) obtained from existing map databases. The server executes scripts to integrate these, resolving duplicate data and inconsistencies. The output is integrated, consistent digital twin data.

[0811] Step 6:

[0812] The server uploads digital twin data to a cloud storage service and distributes it to stakeholders. The specific input is integrated digital twin data. The server uploads the data to a cloud storage service such as Amazon S3. Furthermore, it configures a REST API using the Flask framework and exposes an API endpoint. The output is the digital twin data on cloud storage and the API endpoint.

[0813] Step 7:

[0814] The device (user) collects emotional data using sensor devices such as cameras and microphones. Specific inputs include the user's facial expressions and voice data. The device captures the user's facial expressions through the camera and collects image data. It also collects voice data using the microphone. The output is the collected emotional data.

[0815] Step 8:

[0816] The server analyzes collected emotion data and recognizes the user's emotions in real time. The specific input is emotion data (facial expressions, voice data, etc.) sent from the terminal. The server analyzes facial expressions using OpenCV and Dlib libraries, and analyzes voice data using TensorFlow, etc., to determine the user's emotional state. The output is the recognized emotion data.

[0817] Step 9:

[0818] The server dynamically adjusts the interface based on recognized emotion data. The specific inputs are the recognized emotion data and the data of the currently displayed interface. The server executes a program that adjusts the displayed content based on the user's emotional state. For example, if the user shows a surprised expression, the system highlights interesting information. If the user is confused, it displays an operation guide. The output is the adjusted interface.

[0819] Step 10:

[0820] The user accesses the provided URL and enters their authentication information to access the 3D map. The specific inputs are the provided URL and authentication information. The user accesses the URL using a web browser and logs in using the authentication information. The user can interactively manipulate the 3D map to investigate building heights and road widths in specific areas for real estate development planning. The output is the interactively manipulated 3D map.

[0821] (Application Example 2)

[0822] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0823] In modern industrial settings, it is crucial to monitor the operation of factory robots in real time and direct them to perform optimal actions. However, current systems have limited means of accurately understanding robot operation, and there is a lack of means to improve the user experience. Furthermore, there are currently no systems that improve work efficiency by analyzing the user's emotional state and providing feedback. It is desirable to provide a system that solves these problems and manages robot operations more effectively.

[0824] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0825] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, means for users to access, view, and manipulate the generated 3D map, means for collecting and analyzing user sentiment data in real time, and means for adjusting the user experience based on the analyzed sentiment data. This makes it possible to monitor the operating status of robots in a factory in real time, analyze the emotional state of users, and provide optimal feedback.

[0826] "Geographic information data" refers to data that includes geographical information such as buildings, roads, and topography within cities and factories.

[0827] "Preprocessing" refers to data preparation tasks such as denoising, consistency checking, and format conversion of collected data.

[0828] A "generative model" is an artificial intelligence model that is trained to generate new data based on collected data.

[0829] A "3D map" is a three-dimensional map generated based on geographic information data, representing buildings, roads, and other topographic elements in 3D.

[0830] A "cloud storage service" is a service that allows you to store and manage data via the internet and access it when needed.

[0831] "User" refers to an individual or organization that uses the system to view and manipulate 3D maps.

[0832] "Emotional data" refers to data about a user's emotional state obtained from their facial expressions, voice, and other sources.

[0833] "Analysis" is the process of extracting and understanding information based on collected data.

[0834] An "algorithm" is a method for solving a problem through a series of steps or calculations.

[0835] "Integration" is the process of combining multiple datasets into a single, consistent dataset.

[0836] As an embodiment of this invention, a system for real-time monitoring and control of robot movements within a factory is provided. The system combines a factory digital twin construction system using a generative AI model with an emotion engine that recognizes user emotions. The system operates in the following specific steps.

[0837] First, the server collects geographical information data within the factory. Specifically, it acquires robot location data, operating status, and 3D scan data of the environment. This includes collecting data from sensor devices and cameras within the factory.

[0838] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, it runs a denoising program to filter out unnecessary information from the collected data and converts all data to the WGS84 coordinate system to unify the coordinate system. Furthermore, it converts all data to the GeoJSON format for unification.

[0839] Next, the server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset, and the Generative Adversarial Network (GAN) algorithm is configured. The training data is input into the AI ​​model, and the model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0840] Next, the server generates a 3D map using a trained generative AI model. New geographic data is input into the generative model, which then predicts and generates the 3D shape from this data. The generated 3D map data is stored in an internal database.

[0841] Next, the server integrates the generated 3D map with existing map data. It retrieves supplementary data from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to remove duplicates and inconsistencies for consistency, creating the final digital twin data.

[0842] Furthermore, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. The data uploaded to the cloud is then configured with a REST API to expose an endpoint. Stakeholders are notified of the API endpoint and how to access it.

[0843] In addition, this system incorporates an emotion engine that recognizes the user's emotions. The terminal (user) collects emotion data through cameras, microphones, and other sensor devices. For example, it uses the smartphone's camera and microphone to capture facial expressions and voice. The server analyzes the collected emotion data and uses machine learning algorithms to recognize the user's emotions in real time. Specifically, by analyzing the user's facial expressions and voice, it can identify emotional states such as whether the user is excited or unhappy.

[0844] The recognized emotion data is reflected as feedback in the system interface. For example, if the user shows a surprised expression, the system can adjust the displayed content to emphasize information that is of interest to the user. Also, if the user shows a confused expression, the system will make dynamic changes such as displaying an operation guide.

[0845] Finally, the terminal (user) accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map, for example, to monitor and control the movements of robots working in a specific area of ​​a factory. In addition, the emotion engine provides feedback on the user's emotional data, offering a more intuitive and personalized user experience.

[0846] This system allows for real-time monitoring of robot movements within the factory, analysis of user emotional states, and provision of optimal feedback, thereby improving both work efficiency and the user experience.

[0847] Example of a prompt

[0848] "Develop an application that generates a 3D digital twin model of a factory robot and displays and controls it in real time on a smartphone. Additionally, add a function to analyze the user's facial expressions and voice to provide feedback."

[0849] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0850] Step 1:

[0851] The server collects geographical information data within the factory. Input data includes robot position data, operating status, and 3D scan data of the environment, obtained from sensor devices and cameras. The server receives this data and stores it in a database.

[0852] Step 2:

[0853] The server preprocesses the collected geographic information data. Specifically, it denoises the input data, checks for data consistency, and converts it to GeoJSON format. This process removes noise and outputs a dataset in a consistent format.

[0854] Step 3:

[0855] The server trains a generative AI model. It takes manually created 3D map data as the training dataset and uses a generative adversarial network (GAN) algorithm to train the model. This training process is repeated until the model's accuracy reaches a sufficient level, and a trained model with optimal parameters is output.

[0856] Step 4:

[0857] The server generates 3D maps using a trained generative AI model. New geographic data is input into the model, which then predicts and generates 3D shapes from this data. The generated 3D map data is stored in a database.

[0858] Step 5:

[0859] The server integrates the generated 3D map with existing map data. It retrieves supplementary data from the existing map data and runs an integration script to integrate it with the 3D map. To ensure consistency, duplicates and inconsistencies are resolved, and the final digital twin data is output.

[0860] Step 6:

[0861] The server uploads the digital twin data to a cloud storage service. A REST API is configured for the uploaded data, and an endpoint is exposed. Stakeholders can access the digital twin data using the notified API endpoint.

[0862] Step 7:

[0863] The device (user) uses a camera and microphone to collect emotional data. The user's facial expressions and voice are captured as input data and sent to the server.

[0864] Step 8:

[0865] The server analyzes the received emotional data. Using machine learning algorithms, it recognizes the user's emotional state from the input data. As a result of the analysis, it outputs the emotional state, such as whether the user is excited or unhappy.

[0866] Step 9:

[0867] The server adjusts the system interface based on the analyzed emotion data. For example, if the user has a surprised expression, the system displays information to highlight information that will interest the user. If the user appears confused, dynamic changes are made, such as displaying an operation guide.

[0868] Step 10:

[0869] The user accesses the provided URL, enters their authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map and monitor and control the movements of robots within the factory in real time. Furthermore, they can enjoy a more intuitive and personalized operating experience by utilizing emotional data fed back by the emotion engine.

[0870] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0871] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0872] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0873] [Fourth Embodiment]

[0874] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0875] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0876] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0877] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0878] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0879] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0880] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0881] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0882] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0883] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0884] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0885] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0886] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0887] As an embodiment of this invention, the specific operation of a city digital twin construction system using generative AI will be described. This system consists of the following steps.

[0888] First, the server collects geographic information data for the city. Specifically, it retrieves data such as buildings, roads, land use, and topography from geographic information databases, and also downloads publicly available aerial and satellite imagery. This geographic information data includes formats such as Shapefile and GeoJSON.

[0889] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, a denoising program is run to filter out unnecessary information from the collected data, and all data is converted to the WGS84 coordinate system to unify the coordinate system. Furthermore, all data is converted to and unified in GeoJSON format.

[0890] Next, the server trains the generative AI model. For this, manually created 3D map data is used as the training dataset. For the generative model, for example, a Generative Adversarial Network (GAN) is used. During training, the AI ​​model learns features such as the shapes of buildings and roads, enabling the generation of highly accurate 3D maps.

[0891] Next, the server generates a 3D map using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model generates a predicted 3D shape from this data. The generated 3D map is stored in the server's internal database.

[0892] Next, the server performs a process to integrate the generated 3D map with existing map data. This process involves running scripts to retrieve supplementary data (e.g., trees and walkways) from the existing database and integrate it with the generated 3D map. The integrated data undergoes quality checks, and after correcting inconsistencies and duplicate data, a single, consistent digital twin is created.

[0893] Next, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. In this step, the integrated 3D map is uploaded to the cloud, and a REST API is configured to expose the endpoint. Stakeholders are notified of the API endpoint and how to access it.

[0894] Finally, the terminal (user) accesses the cloud server and uses the generated 3D map. For example, the user can open a web browser, access the provided URL, enter their authentication information, and access the 3D map from the interface. The user can interactively manipulate the 3D map to check the height of buildings and the width of roads in a specific area.

[0895] This system enables the efficient creation of digital twins of cities, making them accessible to a wide range of stakeholders. Furthermore, rapid 3D map generation allows stakeholders to quickly begin exploring use cases. This differentiates the system from the Ministry of Land, Infrastructure, Transport and Tourism's PLATEAU system.

[0896] The following describes the processing flow.

[0897] Step 1:

[0898] The server collects urban geographic information data. Specifically, it accesses geographic information databases and uses APIs to obtain urban building data, road data, and topographic data. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools.

[0899] Step 2:

[0900] The server preprocesses the collected geographic information data. First, it runs a denoising program to filter out unnecessary information and erroneous data points from the collected data. Next, it converts all data to the WGS84 coordinate system to unify the data's coordinate system. Finally, it converts each dataset to GeoJSON format to unify the data.

[0901] Step 3:

[0902] The server trains the generative AI model. It prepares manually created 3D map data as the training dataset and configures the Generative Adversarial Network (GAN) algorithm. The training data is input into the AI ​​model, and the model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[0903] Step 4:

[0904] The server generates 3D maps using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is stored in an internal database.

[0905] Step 5:

[0906] The server integrates the generated 3D map with existing map data. It retrieves supplementary data (e.g., tree and footpath data) from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to eliminate redundancies and inconsistencies for consistency, creating the final digital twin data.

[0907] Step 6:

[0908] The server uploads the integrated digital twin data to a cloud storage service. A REST API is configured to access the uploaded data, and the endpoint is exposed. Furthermore, stakeholders are notified of the API endpoint and how to access it.

[0909] Step 7:

[0910] The terminal (user) accesses a cloud server and utilizes the generated 3D map. The user opens a web browser, accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map and, for example, investigate the height of buildings and the width of roads in a specific area for real estate development planning.

[0911] (Example 1)

[0912] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0913] Providing highly accurate geographic information is essential in modern urban planning and infrastructure management. However, traditional methods involve significant time and effort in collecting and preprocessing geographic data, and generating and integrating 3D models. As a result, providing timely data is difficult, hindering rapid decision-making and action planning by stakeholders. Furthermore, unifying data in different formats and building a consistent digital twin requires advanced technology and expertise. A new system is needed to solve these problems.

[0914] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0915] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a three-dimensional map using the trained generative model, means for integrating with existing map data, means for uploading the generated three-dimensional map to a cloud storage service and distributing it to users, and means for users to access, view, and manipulate the generated three-dimensional map. This enables the rapid and efficient collection and preprocessing of highly accurate geographic information, and the generation and integration of highly accurate three-dimensional maps in a short time using a generative AI model. Furthermore, by utilizing cloud storage, stakeholders can easily access and interact with the data.

[0916] "Geographic information data" refers to data that includes spatial and locational information about urban and regional features (buildings, roads, land use, topography, etc.).

[0917] "Preprocessing" refers to the process of removing noise from collected geographic information data, checking for consistency, unifying coordinate systems, and converting formats.

[0918] A "generative model" is a model that uses machine learning algorithms to generate a desired output (for example, a three-dimensional map) from input data.

[0919] A "three-dimensional map" is map data that visualizes the shape and location information of features in a city or region in three-dimensional space.

[0920] A "cloud storage service" is a service that stores and manages data online, allowing users to access it via the internet.

[0921] A "user" is a person or organization that has the authority to access the generated three-dimensional map and to view and manipulate its information.

[0922] A Generative Adversarial Network (GAN) is a type of generative model in which two paired neural networks compete to generate data.

[0923] A "digital twin" is a virtual model that accurately replicates a physical, real-world object or system, reflecting real-time data.

[0924] A "REST API" is a type of interface for providing web services, and it is a set of design principles for manipulating resources using the HTTP protocol.

[0925] An "API endpoint" is a URL or URI used to access a specific resource or function designated through an API.

[0926] Modes for carrying out the invention

[0927] As an embodiment of this invention, the specific operation of a city digital twin construction system using generative AI is shown below. This system collects and preprocesses geographic information of a city, creates and integrates a three-dimensional map using a generative model, and provides it via a cloud storage service.

[0928] First, the server collects geographic data of the city. Specifically, it uses the OpenStreetMap API to obtain data on buildings, roads, land use, and topography, and downloads the latest satellite imagery from NASA's Landsat satellites. This geographic data includes formats such as Shapefile and GeoJSON. This allows for obtaining up-to-date and detailed geographic information about the city.

[0929] Next, the server preprocesses the collected geographic data. This process includes denoising, consistency checking, and format conversion using the Python Geopandas library. Specifically, Geopandas is used to convert the coordinate system to WGS84, filter out unnecessary data, and unify all data by converting it to GeoJSON format. This process results in consistent geographic data.

[0930] Next, the server trains the generative AI model. Detailed, manually created 3D map data is used as the training dataset. A Generative Adversarial Network (GAN) is used as the generative model. During training, the AI ​​model learns features such as the shapes of buildings and roads, enabling the generation of highly accurate 3D maps. The TensorFlow library is used to train the generative model.

[0931] Next, the server uses a trained generative AI model to generate a 3D map based on new geographic data. New building and road data are input into the generative model, which then generates a predicted 3D shape from this data. The generated 3D map is stored in the server's internal database. This procedure ensures that the most up-to-date 3D map data is always generated.

[0932] Subsequently, the server integrates the generated 3D map with existing map data. This process involves retrieving supplementary data such as trees and footpaths from the existing database and running an integration script. Using the integration script, the generated 3D map and the existing map data are combined into a single, consistent digital twin. After integration, quality checks are also performed to correct inconsistencies and duplicate data.

[0933] Next, the server uploads the integrated 3D map data to a cloud storage service and distributes it to users. In this step, the data is uploaded to an Amazon S3 bucket, and a REST API is configured to expose the endpoint. Stakeholders are notified of the API endpoint and how to access it. This allows stakeholders to access the updated 3D map at any time.

[0934] Finally, the terminal (user) accesses the provided API endpoint to view and interact with the generated 3D map. The user opens a web browser, accesses the provided URL, and enters their authentication information. Through the interface, they can access the 3D map and check the height of buildings and the width of roads in a specific area. This system allows the user to utilize detailed 3D map information that is updated in real time.

[0935] As a concrete example, the server retrieves city building data using the OpenStreetMap API and preprocesses it using the latest satellite imagery downloaded from NASA's Landsat satellites. Then, it trains a GAN model in TensorFlow using manually created detailed 3D map data of Tokyo. Using the new geographic data as input, the trained model generates a highly accurate 3D map, which is then integrated with the existing data. The integrated data is uploaded to an Amazon S3 bucket, and an API endpoint is configured. Stakeholders can access the provided URL to view the 3D map and obtain the necessary information.

[0936] Examples of prompt statements include the following:

[0937] I want to generate 3D map data of a city. Please use the following geographic information data as input to generate detailed 3D shapes of buildings, roads, trees, etc.

[0938] Building data (Shapefile) obtained from OpenStreetMap

[0939] NASA Landsat satellite imagery (GeoTIFF)

[0940] Tree data (GeoJSON)

[0941] Based on the data above, please generate a high-precision 3D map.

[0942] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0943] Step 1:

[0944] The server collects geographic data of cities. Specifically, it uses the OpenStreetMap API to retrieve building and road data and downloads the latest satellite imagery from NASA's Landsat satellites. It takes API endpoints and query parameters as input and receives geographic data in Shapefile or GeoJSON format as output. This provides up-to-date and detailed geographic information of cities.

[0945] Step 2:

[0946] The server preprocesses the collected geographic data. Specifically, it uses the Python Geopandas library to convert the data's coordinate system to WGS84 and filters out unnecessary data. Using the diverse formats of geographic data collected as input, it obtains data in a unified GeoJSON format as output. This process ensures a consistent dataset. For example, to perform a coordinate system transformation using Geopandas, execute geo_df.to_crs(epsg=4326).

[0947] Step 3:

[0948] The server trains a generative AI model. Specifically, it trains a generative adversarial network (GAN) using TensorFlow with manually created 3D map data. Detailed, manually created 3D map data is used as input, and a trained generative AI model is obtained as output. This process allows the AI ​​model to learn the features of buildings and roads, enabling highly accurate 3D map generation. For example, the command `model.fit(training_data, epochs=50)` is used.

[0949] Step 4:

[0950] The server generates a 3D map using a trained generative AI model. Specifically, it takes new geographic data as input and generates the 3D shape predicted by the model. It uses newly collected and pre-processed geographic data as input and produces generated 3D map data as output. This output is stored in the server's internal database. For example, `model.predict(new_data)` is executed to generate the 3D shape.

[0951] Step 5:

[0952] The server integrates the generated 3D map with existing map data. Specifically, it retrieves supplementary data from the existing database and runs an integration script. It uses the new 3D map data and supplementary data (e.g., tree and sidewalk data) as input and obtains a consistent digital twin as output. By running the integration script, the data is integrated using the command integrate_data(new_3d_map, additional_data).

[0953] Step 6:

[0954] The server uploads the integrated 3D map to a cloud storage service and distributes it to stakeholders. Specifically, it uploads the data to an Amazon S3 bucket, configures a REST API, and exposes an endpoint. It uses the integrated data as input and obtains the data stored in the cloud and the API endpoint as output. For example, it uses the command `s3_client.upload_file('integrated_3d_map.json', 'bucket-name', 'path / in / bucket')`.

[0955] Step 7:

[0956] The user accesses the provided API endpoint to view and manipulate the generated 3D map. For example, the user opens a web browser, accesses the provided URL, and enters their authentication information. Using the specified URL and authentication information as input, they obtain an interactive display of the 3D map as output. This allows the user to check the height of buildings and the width of roads in a specific area.

[0957] (Application Example 1)

[0958] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0959] In recent years, in order to solve the increasingly complex traffic problems that accompany urban development, there is a need for real-time, high-precision map information and traffic data-driven operational management of autonomous vehicles. However, current systems have slow map information updates, making it difficult to respond quickly to changes in traffic conditions. Furthermore, the integration of information from multiple data sources is insufficient, making it difficult to provide optimal route guidance. This hinders the improvement of efficiency and safety of autonomous vehicles. This invention aims to solve these problems.

[0960] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0961] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, means for users to access, view, and manipulate the generated 3D map, and means for calculating the optimal route for an autonomous vehicle and performing operational management. This enables efficient operational management of autonomous vehicles based on highly accurate and real-time map information and traffic data.

[0962] "Geographic information data" refers to data that includes information about urban structure, land use, buildings, roads, topography, and so on.

[0963] "Preprocessing" is the process of removing noise from collected data, checking data consistency, and standardizing the format.

[0964] A "generative model" is an algorithm for generating 3D maps through learning, and it utilizes machine learning techniques.

[0965] A "3D map" is a digital map that represents urban structures such as buildings, roads, and terrain in three dimensions.

[0966] A "cloud storage service" is a data storage system that allows data to be stored via the internet and shared among multiple users.

[0967] A "user" is an end-user who utilizes the data and services provided by this system.

[0968] An "autonomous vehicle" is a vehicle that uses sensors and AI to automate its driving process.

[0969] An "optimal route" is a path that minimizes travel time and distance, taking into account traffic conditions and geographical features.

[0970] "Operational management" refers to the process of monitoring and adjusting the operating schedule, routes, and traffic conditions of autonomous vehicles.

[0971] The specific system configuration and operation for realizing this invention are described below.

[0972] Hardware and Software Overview

[0973] This invention utilizes the following main hardware and software.

[0974] hardware

[0975] Server: A server equipped with a high-performance CPU and large-capacity storage.

[0976] Autonomous vehicle onboard computer: Onboard computer equipped with a GPU.

[0977] Device: Smartphone (Android / iOS).

[0978] software

[0979] Cloud storage services: such as AWS S3 and Google Cloud Storage.

[0980] Traffic information APIs: such as Google Maps API and Here API.

[0981] Generative AI model: GAN (Generative Adversarial Network).

[0982] Data processing programs: Python, TensorFlow / PyTorch.

[0983] Explanation of program processing

[0984] The server collects geographic information data. This data is obtained from geographic information databases, aerial photographs, satellite imagery, etc. The collected geographic information data includes formats such as Shapefile and GeoJSON.

[0985] The server preprocesses the collected data. It performs processes such as noise reduction, data consistency checks, and format conversion, and converts all data into a unified format (GeoJSON).

[0986] To train the generative AI model, manually created 3D map data is used as the training dataset. A Generative Adversarial Network (GAN) is used as the generative model to learn features such as the shapes of buildings and roads.

[0987] A pre-trained generative AI model is used to generate 3D maps from new geographic data. These generated 3D maps are stored in the server's internal database.

[0988] Next, the generated 3D map is integrated with existing map data. A script is executed to retrieve supplementary data (e.g., trees and sidewalks) from the existing database and integrate it with the generated 3D map.

[0989] The integrated 3D map data is uploaded to a cloud storage service, and a REST API is configured to expose the endpoint. Users are notified of the API endpoint and how to access it.

[0990] Using a smartphone as a terminal, users access a cloud server and utilize the generated 3D map data. For example, a user can open a web browser, access the provided URL, enter authentication information, and access the 3D map from the interface.

[0991] Operation management of autonomous vehicles

[0992] Furthermore, this system acquires traffic information in real time and calculates the optimal route for autonomous vehicles. Traffic information is collected from traffic sensors, cameras, and social media.

[0993] The terminal's AI algorithm integrates generated 3D map data with real-time traffic information to calculate the optimal route. The calculated route is transmitted to the autonomous vehicle's onboard computer, and the vehicle operates according to that route.

[0994] Users can visually check route information and traffic conditions through the app's interface. For example, they can receive route guidance by entering the following prompt:

[0995] Example of a prompt

[0996] "Please tell me the best route from Tokyo Station to Shinjuku Station."

[0997] Upon receiving this prompt, the system can use the collected data and generated AI models to calculate the optimal route in real time and provide it to the user.

[0998] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0999] Step 1:

[1000] The server collects geographic information data. It retrieves data such as buildings, roads, land use, and topography from geographic information databases, and downloads publicly available aerial and satellite imagery. Input is data in Shapefile or GeoJSON format, and output is the collected raw data.

[1001] Step 2:

[1002] The server preprocesses the geographic information data it collects. A noise reduction program is run to filter out unnecessary information from the collected data, and all data is converted to the WGS84 coordinate system to unify the coordinate system. Furthermore, all data is converted to the GeoJSON format for standardization. The input is the collected raw data, and the output is the clean data after preprocessing.

[1003] Step 3:

[1004] The server trains a generative AI model. Manually created 3D map data is used as the training dataset, and a generative adversarial network (GAN) is used to train the model. The input is the training dataset, and the output is the trained generative AI model.

[1005] Step 4:

[1006] The server generates a 3D map using a trained generative AI model. New geographic data is input to the generative model, and the model generates a predicted 3D shape from this data. The input is new geographic data, and the output is the generated 3D map.

[1007] Step 5:

[1008] The server integrates the generated 3D map with existing map data. A script is executed to retrieve supplementary data (e.g., trees and sidewalks) from the existing database and integrate it with the generated 3D map. The input is the generated 3D map and supplementary data, and the output is the integrated 3D map.

[1009] Step 6:

[1010] The server uploads the integrated 3D map data to a cloud storage service, configures a REST API, and exposes an endpoint. The input is the integrated 3D map, and the output is the data on the cloud storage and the exposed API endpoint.

[1011] Step 7:

[1012] The user accesses the cloud server using their device (smartphone). They access the provided URL, enter their authentication information, and access the 3D map from the interface. The input is the user's authentication information, and the output is the 3D map display on the interface.

[1013] Step 8:

[1014] The server collects real-time traffic information from traffic sensors, cameras, social media, and other sources. The input is real-time traffic data, and the output is collected traffic information.

[1015] Step 9:

[1016] The terminal's AI algorithm integrates generated 3D map data with real-time traffic information to calculate the optimal route. The calculated route is transmitted to the autonomous vehicle's onboard computer, and the vehicle operates according to that route. The input is 3D map data and traffic information, and the output is the optimal route and its operation information.

[1017] Step 10:

[1018] Users visually check route information and traffic conditions through the app's interface. For example, they can receive route guidance by prompting them with a phrase like, "Tell me the best route from Tokyo Station to Shinjuku Station." The input is the user's prompt, and the output is the route information displayed on the screen.

[1019] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1020] As an embodiment of this invention, we will describe the specific operation of a system that combines a city digital twin construction system using generative AI with an emotion engine that recognizes user emotions. This system consists of the following steps.

[1021] First, the server collects geographic information data for cities. Specifically, it accesses geographic information databases and uses APIs to obtain building data, road data, and topographic data for cities. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools.

[1022] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, a denoising program is run to filter out unnecessary information and erroneous data points from the collected data, and all data is converted to the WGS84 coordinate system to unify the coordinate system. Furthermore, all data is converted and unified to the GeoJSON format.

[1023] Next, the server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset, and the Generative Adversarial Network (GAN) algorithm is configured. The training data is input into the AI ​​model, and the model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[1024] Next, the server generates a 3D map using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is saved to an internal database.

[1025] Next, the server integrates the generated 3D map with existing map data. It retrieves supplementary data (e.g., tree and sidewalk data) from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to eliminate redundancies and inconsistencies for consistency, creating the final digital twin data.

[1026] Next, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. In this step, the integrated 3D map is uploaded to the cloud, and a REST API is configured to expose the endpoint. Stakeholders are notified of the API endpoint and how to access it.

[1027] Furthermore, this system incorporates an emotion engine that recognizes the user's emotions. The terminal (user) collects emotion data through a camera, microphone, or other sensor devices. The server analyzes the collected emotion data and uses machine learning algorithms to recognize the user's emotions in real time. For example, by analyzing the user's facial expressions, it is possible to identify the user's emotional state, such as whether they are excited or unhappy.

[1028] The recognized emotion data is reflected as feedback in the system interface. For example, if the user shows a surprised expression, the system can adjust the displayed content to emphasize information that is of interest to the user. Also, if the user shows a confused expression, the system will make dynamic changes such as displaying an operation guide.

[1029] Finally, the device (user) accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map and, for example, investigate the height of buildings and the width of roads in a specific area for real estate development planning. In addition, the emotion engine provides feedback on the user's emotional data, offering a more intuitive and personalized user experience.

[1030] This system enables the efficient construction of digital twins of cities, making them accessible to a wide range of stakeholders, while also enhancing the user experience through an emotion engine. Furthermore, rapid 3D map generation and the integration of emotion data expand the usability of urban digital twins, allowing stakeholders to quickly begin exploring use cases. This differentiates the system from the Ministry of Land, Infrastructure, Transport and Tourism's PLATEAU.

[1031] The following describes the processing flow.

[1032] Step 1:

[1033] The server collects urban geographic information data. Specifically, it accesses geographic information databases and uses APIs to obtain urban building data, road data, and topographic data. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools.

[1034] Step 2:

[1035] The server preprocesses the collected geographic information data. First, it runs a denoising program to filter out unnecessary information and erroneous data points from the collected data. Next, it converts all data to the WGS84 coordinate system to unify the data's coordinate system. Finally, it converts each dataset to GeoJSON format to unify the data.

[1036] Step 3:

[1037] The server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset. The Generative Adversarial Network (GAN) algorithm is configured, and the training data is input into the AI ​​model. The AI ​​model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[1038] Step 4:

[1039] The server generates 3D maps using a trained generative AI model. New geographic data (building data, road data, etc.) is input into the generative model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is stored in an internal database.

[1040] Step 5:

[1041] The server integrates the generated 3D map with existing map data. It runs a script that retrieves supplementary data from the existing map database and integrates it with the 3D map. The integrated data is then processed to eliminate duplication and inconsistencies, creating a consistent digital twin.

[1042] Step 6:

[1043] The server uploads the integrated digital twin data to a cloud storage service. A REST API is configured to access the data, and the endpoint is exposed. Stakeholders are notified of the API endpoint and how to access it.

[1044] Step 7:

[1045] The user accesses a cloud server and utilizes the generated 3D map. The user opens a web browser, accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user interacts with the 3D map and investigates information about a specific area.

[1046] Step 8:

[1047] The device (user) collects emotional data using a camera, microphone, or other sensor devices. For example, the camera is used to capture the user's facial expressions, and the microphone is used to collect the tone of their voice.

[1048] Step 9:

[1049] The server analyzes the collected emotional data. Using machine learning algorithms, it analyzes the user's facial expressions and tone of voice to recognize the user's emotional state. For example, it can determine from facial expressions whether the user is excited or unhappy.

[1050] Step 10:

[1051] The server dynamically changes the system interface based on the analysis results. For example, if the user looks confused, it displays an operation guide. If the user looks surprised, it highlights information that might interest them.

[1052] This system efficiently constructs digital twins of cities, making them accessible to numerous stakeholders, while also enhancing the user experience through an emotion engine. Rapid 3D map generation and the integration of emotion data expand the usability of urban digital twins, allowing stakeholders to quickly begin exploring use cases. This differentiates the system from other existing systems.

[1053] (Example 2)

[1054] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1055] Conventional urban digital twin systems required improved efficiency in collecting geographic information data and generating 3D maps, but they lacked the ability to recognize user emotions in real time and reflect them in the interface. Therefore, it was difficult to provide users with an intuitive and personalized user experience. Furthermore, it was challenging to provide the generated 3D maps in real time while maintaining the consistency of the integrated data.

[1056] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1057] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, means for users to access, view, and manipulate the generated 3D map, means for collecting sentiment data, means for analyzing the collected sentiment data and recognizing the user's sentiment in real time, and means for dynamically adjusting the interface based on the recognized sentiment data. This enables the efficient construction of a city digital twin and the provision of an intuitive and personalized user experience that takes into account the user's sentiment.

[1058] "Geographic information data" refers to a dataset containing geographical location information, providing information about elements such as buildings, roads, and topography.

[1059] "Preprocessing" is the process of ensuring the consistency of collected data by performing tasks such as noise reduction, format conversion, and coordinate system standardization.

[1060] A "generative model" is an algorithm that generates new data based on input data, and often refers specifically to machine learning algorithms.

[1061] A "trained generative model" is a generative model that has been learned using training data and optimized to perform a specific task.

[1062] A "3D map" is a visualization of geographic information data in three-dimensional space, and is a data structure designed to realistically reproduce real-world geographical features.

[1063] A "cloud storage service" is a remote storage service that allows you to store and manage data via the internet.

[1064] "Emotional data" refers to data that indicates a user's emotional state, collected based on factors such as the user's facial expressions and tone of voice.

[1065] "Recognizing in real time" means performing the entire process from data collection to analysis immediately, and providing results without any time delay.

[1066] An "interface" refers to the screen or control panel that a user uses to interact with a system, and is the point of contact that provides a user experience.

[1067] "Dynamic adjustment" refers to changing the system's display content and behavior in real time according to the user's situation and actions.

[1068] Modes for carrying out the invention

[1069] As an embodiment of this invention, we will describe the specific operation of a system that combines a city digital twin construction system using generative AI with an emotion engine that recognizes user emotions. This system consists of the following steps.

[1070] First, the server collects geographic information data for cities. Specifically, it accesses geographic information databases and uses APIs to obtain building data, road data, and terrain data for cities. It also downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools. For example, it can use open-source geographic information databases (e.g., OpenStreetMap API). It can also obtain the latest satellite imagery through NASA's satellite imagery service. Furthermore, it is possible to scrape the necessary image data using the Python library BeautifulSoup.

[1071] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, it uses Python's NumPy and Pandas libraries to check data consistency and remove unnecessary information or incorrect data points. It also uses the GeoPandas library to unify the coordinate system of the collected data to WGS84 (e.g., EPSG:4326) and convert it to GeoJSON format.

[1072] Next, the server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset, and the model is built using the Generative Adversarial Network (GAN) algorithm. The server uses libraries such as TensorFlow and PyTorch to input the manually created 3D map data into the AI ​​model and runs a program to learn the shapes of buildings and roads. This process is repeated until the model's accuracy is sufficiently high.

[1073] Next, the server generates a 3D map using a trained generative AI model. For example, new geographic information data (building data, road data, etc.) is input into the generative AI model, and the model predicts and generates 3D shapes from this data. The generated 3D map data is stored in an internal database. By using geographic information databases such as PostGIS, efficient data storage and retrieval are possible.

[1074] Next, the server integrates the generated 3D map with existing map data. It retrieves supplementary data (e.g., tree and sidewalk data) from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to eliminate duplication and inconsistencies, completing it as the final digital twin data.

[1075] Next, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. In this step, the integrated 3D map is uploaded to the cloud (for example, Amazon S3). To set up a REST API and notify stakeholders how to access it, an API endpoint can be built using the Flask framework.

[1076] Furthermore, this system incorporates an emotion engine that recognizes the user's emotions. The terminal (user) collects emotion data through a camera, microphone, or other sensor devices. For example, the camera is used to capture the user's facial expressions and the image data is sent to the server. The microphone is also used to collect audio data, capturing the user's speech and tone of voice.

[1077] The server analyzes collected emotional data and uses machine learning algorithms to recognize the user's emotions in real time. Specifically, it analyzes facial expressions using OpenCV and Dlib libraries, and analyzes audio data using TensorFlow, etc., to determine the user's emotional state. For example, if the user has a surprised expression, it recognizes an excited state, and if they have a grumpy expression, it understands the situation.

[1078] The recognized emotion data is reflected as feedback in the system interface. For example, if the user shows a surprised expression, the system adjusts the displayed content to emphasize information that will interest the user. Also, if the user shows a confused expression, the system makes dynamic changes such as displaying an operation guide.

[1079] Finally, the terminal (user) accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user accesses the provided URL through a browser and logs in using the authentication information. It provides interactive functionality to manipulate the 3D map, making it possible, for example, to survey the height of buildings and the width of roads in a specific area for real estate development planning. The emotion engine provides real-time feedback of the user's emotional data, offering an intuitive and personalized user experience.

[1080] As a concrete example, the following prompt statement is shown.

[1081] Example of a prompt:

[1082] Please describe a program that acquires satellite imagery and geographic information data, preprocesses it with noise reduction and format conversion, generates a 3D map using an AI model, integrates it with existing data, and uploads it to cloud storage. Also, please describe an emotion engine that recognizes user emotions and provides real-time feedback.

[1083] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1084] Step 1:

[1085] The server collects urban geographic information data. Specifically, the server accesses a geographic information database and retrieves building data, road data, and terrain data via an API. The input requires the database's API endpoint and request parameters, while the output is datasets of buildings, roads, and terrain. In addition, the server downloads the latest image data from free satellite imagery services and aerial photo websites using scraping tools. The BeautifulSoup library in Python is used for these operations.

[1086] Step 2:

[1087] The server preprocesses the collected geographic information data. The specific inputs are building data, road data, terrain data, and image data collected in Step 1. First, the data consistency is checked using Python's NumPy and Pandas libraries, and unnecessary information and incorrect data points are filtered out. Next, the GeoPandas library is used to unify the data coordinate system to WGS84 (EPSG:4326) and convert it to GeoJSON format. The output is preprocessed geographic information data in a unified format.

[1088] Step 3:

[1089] The server trains a generative AI model. The specific inputs are collected and pre-processed geographic data and manually created 3D map datasets. The server uses TensorFlow and PyTorch libraries to train a generative adversarial network (GAN) algorithm with these inputs. The datasets are fed into the AI ​​model to learn the shapes of buildings and roads. This process is repeated iteratively until the model's accuracy is sufficiently high. The output is the trained AI model.

[1090] Step 4:

[1091] The server generates 3D maps using a trained generative AI model. The specific inputs are the trained AI model and new geographic data (building data, road data, etc.). This data is input to the model, which then predicts and generates the 3D shape. The output is the generated 3D map data, which is stored in an internal database (e.g., PostGIS).

[1092] Step 5:

[1093] The server integrates the generated 3D map with existing map data. The specific inputs are the generated 3D map data and supplementary data (tree data, footpath data, etc.) obtained from existing map databases. The server executes scripts to integrate these, resolving duplicate data and inconsistencies. The output is integrated, consistent digital twin data.

[1094] Step 6:

[1095] The server uploads digital twin data to a cloud storage service and distributes it to stakeholders. The specific input is integrated digital twin data. The server uploads the data to a cloud storage service such as Amazon S3. Furthermore, it configures a REST API using the Flask framework and exposes an API endpoint. The output is the digital twin data on cloud storage and the API endpoint.

[1096] Step 7:

[1097] The device (user) collects emotional data using sensor devices such as cameras and microphones. Specific inputs include the user's facial expressions and voice data. The device captures the user's facial expressions through the camera and collects image data. It also collects voice data using the microphone. The output is the collected emotional data.

[1098] Step 8:

[1099] The server analyzes collected emotion data and recognizes the user's emotions in real time. The specific input is emotion data (facial expressions, voice data, etc.) sent from the terminal. The server analyzes facial expressions using OpenCV and Dlib libraries, and analyzes voice data using TensorFlow, etc., to determine the user's emotional state. The output is the recognized emotion data.

[1100] Step 9:

[1101] The server dynamically adjusts the interface based on recognized emotion data. The specific inputs are the recognized emotion data and the data of the currently displayed interface. The server executes a program that adjusts the displayed content based on the user's emotional state. For example, if the user shows a surprised expression, the system highlights interesting information. If the user is confused, it displays an operation guide. The output is the adjusted interface.

[1102] Step 10:

[1103] The user accesses the provided URL and enters their authentication information to access the 3D map. The specific inputs are the provided URL and authentication information. The user accesses the URL using a web browser and logs in using the authentication information. The user can interactively manipulate the 3D map to investigate building heights and road widths in specific areas for real estate development planning. The output is the interactively manipulated 3D map.

[1104] (Application Example 2)

[1105] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1106] In modern industrial settings, it is crucial to monitor the operation of factory robots in real time and direct them to perform optimal actions. However, current systems have limited means of accurately understanding robot operation, and there is a lack of means to improve the user experience. Furthermore, there are currently no systems that improve work efficiency by analyzing the user's emotional state and providing feedback. It is desirable to provide a system that solves these problems and manages robot operations more effectively.

[1107] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1108] In this invention, the server includes means for collecting geographic information data, means for preprocessing the collected geographic information data, means for training a generative model, means for generating a 3D map using the trained generative model, means for integrating with existing map data, means for uploading the generated 3D map to a cloud storage service and distributing it to users, means for users to access, view, and manipulate the generated 3D map, means for collecting and analyzing user sentiment data in real time, and means for adjusting the user experience based on the analyzed sentiment data. This makes it possible to monitor the operating status of robots in a factory in real time, analyze the emotional state of users, and provide optimal feedback.

[1109] "Geographic information data" refers to data that includes geographical information such as buildings, roads, and topography within cities and factories.

[1110] "Preprocessing" refers to data preparation tasks such as denoising, consistency checking, and format conversion of collected data.

[1111] A "generative model" is an artificial intelligence model that is trained to generate new data based on collected data.

[1112] A "3D map" is a three-dimensional map generated based on geographic information data, representing buildings, roads, and other topographic elements in 3D.

[1113] A "cloud storage service" is a service that allows you to store and manage data via the internet and access it when needed.

[1114] "User" refers to an individual or organization that uses the system to view and manipulate 3D maps.

[1115] "Emotional data" refers to data about a user's emotional state obtained from their facial expressions, voice, and other sources.

[1116] "Analysis" is the process of extracting and understanding information based on collected data.

[1117] An "algorithm" is a method for solving a problem through a series of steps or calculations.

[1118] "Integration" is the process of combining multiple datasets into a single, consistent dataset.

[1119] As an embodiment of this invention, a system for real-time monitoring and control of robot movements within a factory is provided. The system combines a factory digital twin construction system using a generative AI model with an emotion engine that recognizes user emotions. The system operates in the following specific steps.

[1120] First, the server collects geographical information data within the factory. Specifically, it acquires robot location data, operating status, and 3D scan data of the environment. This includes collecting data from sensor devices and cameras within the factory.

[1121] Next, the server preprocesses the collected geographic data. This process includes denoising, data consistency checks, and format conversion. For example, it runs a denoising program to filter out unnecessary information from the collected data and converts all data to the WGS84 coordinate system to unify the coordinate system. Furthermore, it converts all data to the GeoJSON format for unification.

[1122] Next, the server trains the generative AI model. For this, manually created 3D map data is prepared as the training dataset, and the Generative Adversarial Network (GAN) algorithm is configured. The training data is input into the AI ​​model, and the model learns the shapes of buildings and roads. This process is repeated until the model reaches a sufficient level of accuracy.

[1123] Next, the server generates a 3D map using a trained generative AI model. New geographic data is input into the generative model, which then predicts and generates the 3D shape from this data. The generated 3D map data is stored in an internal database.

[1124] Next, the server integrates the generated 3D map with existing map data. It retrieves supplementary data from the existing map database and runs a script to integrate this data with the 3D map. The integrated data is then processed to remove duplicates and inconsistencies for consistency, creating the final digital twin data.

[1125] Furthermore, the server uploads the digital twin data to a cloud storage service and distributes it to stakeholders. The data uploaded to the cloud is then configured with a REST API to expose an endpoint. Stakeholders are notified of the API endpoint and how to access it.

[1126] In addition, this system incorporates an emotion engine that recognizes the user's emotions. The terminal (user) collects emotion data through cameras, microphones, and other sensor devices. For example, it uses the smartphone's camera and microphone to capture facial expressions and voice. The server analyzes the collected emotion data and uses machine learning algorithms to recognize the user's emotions in real time. Specifically, by analyzing the user's facial expressions and voice, it can identify emotional states such as whether the user is excited or unhappy.

[1127] The recognized emotion data is reflected as feedback in the system interface. For example, if the user shows a surprised expression, the system can adjust the displayed content to emphasize information that is of interest to the user. Also, if the user shows a confused expression, the system will make dynamic changes such as displaying an operation guide.

[1128] Finally, the terminal (user) accesses the provided URL, enters authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map, for example, to monitor and control the movements of robots working in a specific area of ​​a factory. In addition, the emotion engine provides feedback on the user's emotional data, offering a more intuitive and personalized user experience.

[1129] This system allows for real-time monitoring of robot movements within the factory, analysis of user emotional states, and provision of optimal feedback, thereby improving both work efficiency and the user experience.

[1130] Example of a prompt

[1131] "Develop an application that generates a 3D digital twin model of a factory robot and displays and controls it in real time on a smartphone. Additionally, add a function to analyze the user's facial expressions and voice to provide feedback."

[1132] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1133] Step 1:

[1134] The server collects geographical information data within the factory. Input data includes robot position data, operating status, and 3D scan data of the environment, obtained from sensor devices and cameras. The server receives this data and stores it in a database.

[1135] Step 2:

[1136] The server preprocesses the collected geographic information data. Specifically, it denoises the input data, checks for data consistency, and converts it to GeoJSON format. This process removes noise and outputs a dataset in a consistent format.

[1137] Step 3:

[1138] The server trains a generative AI model. It takes manually created 3D map data as the training dataset and uses a generative adversarial network (GAN) algorithm to train the model. This training process is repeated until the model's accuracy reaches a sufficient level, and a trained model with optimal parameters is output.

[1139] Step 4:

[1140] The server generates 3D maps using a trained generative AI model. New geographic data is input into the model, which then predicts and generates 3D shapes from this data. The generated 3D map data is stored in a database.

[1141] Step 5:

[1142] The server integrates the generated 3D map with existing map data. It retrieves supplementary data from the existing map data and runs an integration script to integrate it with the 3D map. To ensure consistency, duplicates and inconsistencies are resolved, and the final digital twin data is output.

[1143] Step 6:

[1144] The server uploads the digital twin data to a cloud storage service. A REST API is configured for the uploaded data, and an endpoint is exposed. Stakeholders can access the digital twin data using the notified API endpoint.

[1145] Step 7:

[1146] The device (user) uses a camera and microphone to collect emotional data. The user's facial expressions and voice are captured as input data and sent to the server.

[1147] Step 8:

[1148] The server analyzes the received emotional data. Using machine learning algorithms, it recognizes the user's emotional state from the input data. As a result of the analysis, it outputs the emotional state, such as whether the user is excited or unhappy.

[1149] Step 9:

[1150] The server adjusts the system interface based on the analyzed emotion data. For example, if the user has a surprised expression, the system displays information to highlight information that will interest the user. If the user appears confused, dynamic changes are made, such as displaying an operation guide.

[1151] Step 10:

[1152] The user accesses the provided URL, enters their authentication information, and accesses the 3D map from the interface. The user can interactively manipulate the 3D map and monitor and control the movements of robots within the factory in real time. Furthermore, they can enjoy a more intuitive and personalized operating experience by utilizing emotional data fed back by the emotion engine.

[1153] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1154] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1155] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1156] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1157] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1158] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1159] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1160] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1161] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1162] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1163] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1164] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1165] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1166] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1167] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1168] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1169] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1170] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1171] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1172] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1173] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1174] The following is further disclosed regarding the embodiments described above.

[1175] (Claim 1)

[1176] Means for collecting geographic information data,

[1177] A means for preprocessing the collected geographic information data,

[1178] Means for training generative models,

[1179] A means for generating a 3D map using a trained generative model,

[1180] Means for integrating with existing map data,

[1181] A method for uploading the generated 3D map to a cloud storage service and distributing it to users,

[1182] A means for users to access and view and manipulate a 3D map that has been generated,

[1183] A system that includes this.

[1184] (Claim 2)

[1185] The system according to claim 1, wherein the generative model uses a machine learning algorithm.

[1186] (Claim 3)

[1187] The system according to claim 2, wherein the machine learning algorithm includes a generated adversarial network.

[1188] "Example 1"

[1189] (Claim 1)

[1190] Means for collecting geographic information data,

[1191] A means for preprocessing the collected geographic information data,

[1192] Means for training generative models,

[1193] A means for generating a three-dimensional map using a trained generative model,

[1194] Means for integrating with existing map data,

[1195] A method for uploading the generated 3D map to a cloud storage service and distributing it to users,

[1196] A means for users to access and view and manipulate a three-dimensional map that has been generated,

[1197] A system that includes this.

[1198] (Claim 2)

[1199] The system according to claim 1, wherein the geographic information data collected by the collection means is obtained using the OpenStreetMap API and includes NASA satellite imagery.

[1200] (Claim 3)

[1201] The system according to claim 1, wherein the preprocessing means includes a process of using the Python Geopandas library to transform the coordinate system of geographic information data, filter the data, and convert it to GeoJSON format.

[1202] (Claim 4)

[1203] The system according to claim 1, wherein the generative model uses a machine learning algorithm.

[1204] (Claim 5)

[1205] The system according to claim 4, wherein the machine learning algorithm includes a generative adversarial network.

[1206] (Claim 6)

[1207] The system according to claim 1, wherein the means for uploading to cloud storage includes a process of uploading data using an Amazon S3 bucket and configuring a REST API.

[1208] "Application Example 1"

[1209] (Claim 1)

[1210] Means for collecting geographic information data,

[1211] A means for preprocessing the collected geographic information data,

[1212] Means for training generative models,

[1213] A means for generating a 3D map using a trained generative model,

[1214] Means for integrating with existing map data,

[1215] A method for uploading the generated 3D map to a cloud storage service and distributing it to users,

[1216] A means for users to access and view and manipulate a 3D map that has been generated,

[1217] A means of calculating the optimal route for autonomous vehicles and managing their operation,

[1218] A system that includes this.

[1219] (Claim 2)

[1220] The system according to claim 1, wherein the generative model uses a machine learning algorithm.

[1221] (Claim 3)

[1222] The system according to claim 2, wherein the machine learning algorithm includes a generated adversarial network.

[1223] "Example 2 of combining an emotion engine"

[1224] (Claim 1)

[1225] Means for collecting geographic information data,

[1226] A means for preprocessing the collected geographic information data,

[1227] Means for training generative models,

[1228] A means for generating a 3D map using a trained generative model,

[1229] Means of integrating with existing map data,

[1230] A method for uploading the generated 3D map to a cloud storage service and distributing it to users,

[1231] A means for users to access and view and manipulate a 3D map that has been generated,

[1232] Means of collecting emotional data,

[1233] A means of analyzing collected emotional data and recognizing users' emotions in real time,

[1234] A means of dynamically adjusting the interface based on recognized emotion data,

[1235] A system that includes this.

[1236] (Claim 2)

[1237] The system according to claim 1, wherein the generative model uses a machine learning algorithm.

[1238] (Claim 3)

[1239] The system according to claim 2, wherein the machine learning algorithm includes a generated adversarial network.

[1240] "Application example 2 of combining emotional engines"

[1241] (Claim 1)

[1242] Means for collecting geographic information data,

[1243] A means for preprocessing the collected geographic information data,

[1244] Means for training generative models,

[1245] A means for generating a 3D map using a trained generative model,

[1246] Means for integrating with existing map data,

[1247] A method for uploading the generated 3D map to a cloud storage service and distributing it to users,

[1248] A means for users to access and view and manipulate a 3D map that has been generated,

[1249] A means of collecting and analyzing user sentiment data in real time,

[1250] A means of adjusting the user experience based on analyzed emotional data,

[1251] A system that includes this.

[1252] (Claim 2)

[1253] The system according to claim 1, wherein the generative model uses a machine learning algorithm.

[1254] (Claim 3)

[1255] The system according to claim 2, wherein the machine learning algorithm includes a generated adversarial network. [Explanation of Symbols]

[1256] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means for collecting geographic information data, A means for preprocessing the collected geographic information data, Means for training generative models, A means for generating a 3D map using a trained generative model, Means for integrating with existing map data, A method for uploading the generated 3D map to a cloud storage service and distributing it to users, A means for users to access and view and manipulate a 3D map that has been generated, A system that includes this.

2. The system according to claim 1, wherein the generative model uses a machine learning algorithm.

3. The system according to claim 2, wherein the machine learning algorithm includes a generative adversarial network.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A