System

A system that generates and distributes multimedia content based on regional information addresses the decline in rural tourism by promoting local attractions and culture, enhancing tourism and social vitality.

JP2026035470APending Publication Date: 2026-03-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024138313
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Rural areas face depopulation leading to decreased tourism, which results in the decline of local attractions and social issues, lacking effective means to promote their appeal.

Method used

A system that inputs a region's name, collects information, generates a scenario, creates video and audio content, and distributes it through a URL or download link, effectively promoting local tourist attractions and culture.

Benefits of technology

The system revitalizes rural areas by widely publicizing tourist spots, increasing tourism, and conveying unique regional charms at a low cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035470000001_ABST
    Figure 2026035470000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a means for inputting a name of a local area, a means for collecting information based on the name of the local area, a means for generating a scenario based on the collected information, a means for generating video based on the generated scenario, a means for generating audio corresponding to the video, a means for integrating the scenario, the video and the audio to create content, and a means for providing the created content.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] As depopulation progresses in rural areas, tourist attractions are becoming less well-known, resulting in a frequent decline in tourists. As a result, serious social problems have arisen, such as the closure of local shopping districts, a deterioration in public safety due to population decline, and the isolation of transportation networks. Another contributing factor is the lack of effective means to promote the appeal of rural areas. The present invention aims to solve these problems and revitalize rural areas. [Means for solving the problem]

[0005] The present invention provides a system including a means for inputting the name of a region, a means for collecting information based on the name of the region, a means for generating a scenario based on the collected information, a means for generating video based on the generated scenario, a means for generating audio corresponding to the video, a means for creating content by integrating the scenario, video, and audio, and a means for providing the created content. This system can promote local tourist attractions, increase tourism, and revitalize the region. Furthermore, since tourist information, historical information, and cultural information are collected based on the name of the region and content is generated based on this information, the unique charms of the region can be effectively conveyed. Furthermore, by providing a means for saving the created content to a distribution destination and generating a URL or download link, the content can be easily provided to users.

[0006] A "regional name" is a place name that indicates a specific region or area.

[0007] "Means for collecting information" means a mechanism for obtaining tourist, historical, and cultural information related to a specified region using external data sources and APIs.

[0008] A "means for generating a scenario" is a system or algorithm for automatically generating a story or script based on collected information.

[0009] The "means for generating images" is a system that automatically creates images including scenery and characters from a specified region based on the generated scenario.

[0010] The "means for generating audio" is a system for generating audio data such as narration, character voices, background music, etc., corresponding to the scenario and video.

[0011] The "means for creating content" is a system for integrating the generated scenario, video and audio to create a single piece of animation content.

[0012] "Means for providing content" refers to a mechanism for providing created content to users in a format that allows them to view or download it.

[0013] "Tourist information" refers to information related to local tourist attractions, access methods, scenic spots, etc.

[0014] "Historical information" is information relating to local history, important events, historic buildings, etc.

[0015] "Cultural information" refers to information related to local traditions, customs, festivals, local specialties, etc.

[0016] "Means for saving at the distribution destination" refers to a mechanism for saving the created content on a cloud server or storage service.

[0017] "Destination URL or download link" means the internet address where users can access, view, or download the created content. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The system of the present invention is designed to automatically generate animation content related to a region by inputting the name of the region, thereby widely publicizing the tourist spots. Specific embodiments of the system and the processing of the program therefor are described below.

[0040] System Configuration

[0041] The system is composed of multiple components, including a user terminal, a server, a scenario generation AI, a video generation AI, an audio generation AI, and an external information acquisition API. The role of each component is explained below.

[0042] User terminal

[0043] The user terminal provides an interface for inputting the name of a region. This interface generally runs on a web browser and is designed to be easy for users to operate.

[0044] server

[0045] The server is the core of the entire system. It receives the name of the region sent from the user's device and processes requests to the scenario generation AI, video generation AI, and audio generation AI in sequence. It also calls an external information acquisition API to collect tourist information, historical information, and cultural information about the region.

[0046] Program processing

[0047] The processing of the program will be explained in natural language below.

[0048] 1. User Input

[0049] The user inputs the name of a region (e.g., "Kochi Prefecture") using the terminal.

[0050] 2. Send to the server

[0051] The terminal transmits the entered name of the region to the server.

[0052] 3. Information gathering

[0053] The server calls an external information acquisition API based on the region name and collects tourist information, historical information, and cultural information.

[0054] 4. Scenario Generation

[0055] The server passes the collected information to a scenario generation AI, which generates scenarios relevant to the region.

[0056] 5. Image Generation

[0057] The server passes the generated scenario to the video generation AI, which then generates a video based on the scenario.

[0058] 6. Speech Generation

[0059] The server passes the generated scenario and video to an audio generation AI, which generates narration, character voices, and background music that correspond to the scenario.

[0060] 7. Content Integration

[0061] The server integrates the scenario, video, and audio to create the final animation content.

[0062] 8. Content Provision

[0063] The server stores the generated content in cloud storage and generates and provides a URL or download link for the content.

[0064] Specific examples

[0065] As a specific example of use, consider the case where "Kochi Prefecture" is entered.

[0066] 1. The user enters "Kochi Prefecture" into the terminal.

[0067] 2. The device sends the input information to the server.

[0068] 3. The server collects tourist information (e.g., Katsurahama Beach, Shimanto River), historical information (e.g., Sakamoto Ryoma), and cultural information (e.g., Yosakoi Festival) related to Kochi Prefecture from external APIs.

[0069] 4. The server sends a request to the scenario generation AI based on the collected information and generates an anime scenario.

[0070] 5. The server requests the video generation AI to generate a video based on the generated scenario, creating a video that includes scenery and tourist attractions in Kochi Prefecture.

[0071] 6. The server then uses the voice generation AI to create narration, character voices, and background music that correspond to the generated scenario and video.

[0072] 7. The server integrates the scenario, video, and audio, and saves the completed animation content in cloud storage.

[0073] 8. The server generates a URL or download link for the destination and notifies the user.

[0074] In this way, the system based on the present invention effectively promotes local tourist attractions, increasing the number of tourists and revitalizing the region.

[0075] The processing flow will be explained below.

[0076] Step 1:

[0077] The user uses a terminal to input the name of a region (e.g., "Kochi Prefecture").

[0078] Step 2:

[0079] The terminal sends the inputted name of the region to the server as an HTTP request.

[0080] Step 3:

[0081] Based on the name of the region received by the server, tourist information, historical information, and cultural information are collected using an external information acquisition API.

[0082] Step 4:

[0083] The server passes the collected information to a scenario generation AI, which then generates a scenario for an anime related to the region.

[0084] Step 5:

[0085] The scenario generation AI returns the generated scenario to the server, which then passes it on to the video generation AI.

[0086] Step 6:

[0087] The video generation AI generates video based on the scenario and returns the generated video data to the server.

[0088] Step 7:

[0089] The server passes the generated video data and scenario to the voice generation AI, which then generates the corresponding narration, character voices, and background music.

[0090] Step 8:

[0091] The voice generation AI returns the generated voice data to the server, which then integrates the scenario, video, and voice to create the final animation content.

[0092] Step 9:

[0093] The server saves the completed animation content in cloud storage and generates a distribution URL or download link.

[0094] Step 10:

[0095] The server sends a notification to the user containing a distribution URL or a download link, which the user receives to view or download the animation content.

[0096] Example 1

[0097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0098] Currently, many regions lack the means to effectively promote their tourism resources, making it difficult to attract tourists' attention. Expressing a region's history and culture through video and audio is also problematic, requiring significant costs and time. To address this issue, there is a demand for a system that can easily and automatically generate content that conveys a region's tourist attractions and culture in an appealing way.

[0099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0100] In this invention, the server includes means for collecting tourist information, historical information, and cultural information based on the name of a region, means for generating a scenario using a generation AI model based on the collected information, means for sending the generated scenario to a video generation AI to generate video, means for sending the generated scenario and video to an audio generation AI to generate audio, means for creating content by integrating the scenario, video, and audio, and means for saving the created content and generating a URL or download link for it, which makes it possible to effectively promote local tourist attractions and culture and increase the number of tourists.

[0101] A "local name" is a name used to identify a particular area, usually indicating a geographical, administrative or historical division of that area.

[0102] "Information gathering means" refers to the methods and devices used to obtain data from external sources and convert it into a form that can be used within the system.

[0103] The term "means for generating a scenario" refers to a method or device that automatically creates text data containing elements of a narrative or storytelling based on collected information.

[0104] "Means for generating video" refers to a method or device for automatically creating visual video content based on text data or a scenario.

[0105] "Audio generating means" refers to a method or device for creating audio data corresponding to a scenario or video, including narration, character voices, background music, etc.

[0106] The "means for creating content" refers to a method or device for integrating the generated scenario, video and audio to create a complete media content.

[0107] "Cloud storage" refers to an online storage service that stores data on remote servers on the Internet and allows it to be accessed and shared as needed.

[0108] "Generative AI model" refers to an artificial intelligence model that has been trained using machine learning to perform a specific task (e.g., text generation, image generation, speech generation).

[0109] A "prompt" is an instruction given to a generative AI model, and refers to text that guides the model in determining what output to generate.

[0110] The system of the present invention automatically generates content (especially animations) related to a region by inputting the name of the region, thereby widely publicizing the tourist spots. The system is composed of multiple components, including a user terminal, a server, a scenario generation AI, a video generation AI, a sound generation AI, and an external information acquisition API.

[0111] User terminal

[0112] The user terminal provides an interface for inputting the name of a region. This interface generally runs on a web browser and is designed to be easy for users to operate. The user inputs the name of the region through this interface.

[0113] server

[0114] The server is the core of the entire system and performs the following processes:

[0115] 1. Receive the name of the region sent from the user terminal.

[0116] 2. Call external information acquisition APIs to collect tourist information, historical information, and cultural information. For example, use Wikipedia API or Google Maps API.

[0117] 3. Based on the collected information, a scenario generation AI (e.g., OpenAI's GPT-4) is requested to generate a scenario. The following prompt is used as an example:

[0118] "Create a scenario related to tourist attractions, history, and culture based on the local name. For example, for "Kochi Prefecture," generate a scenario that includes "Katsurahama Beach," "Shimanto River," "Sakamoto Ryoma," and "Yosakoi Festival."

[0119] 4. The generated scenario is sent to an image generation AI (e.g., DALL-E, Disco Diffusion), which generates an image based on the scenario. The following prompts are used:

[0120] "Generate an animated video based on a given scenario, including the scenery and tourist attractions of the target region."

[0121] 5. The generated scenario and video are sent to a voice generation AI (e.g., Google Cloud Text-to-Speech, Amazon Polly) to generate narration, character voices, and background music corresponding to the scenario. The following prompts are used:

[0122] "Generate narration, character voices, and background music for the following scenario."

[0123] 6. The script, video, and audio are integrated to create the final animation content. This integration process is carried out using software such as Adobe Premiere Pro and FFmpeg.

[0124] 7. Save the created content to cloud storage (e.g., Amazon S3, Google Drive) and generate a URL or download link to provide to users.

[0125] Specific examples

[0126] For example, if a user inputs the name of a region, "Kochi Prefecture," the device sends this information to the server. The server collects tourist attractions (Katsurahama Beach, Shimanto River), historical information (Sakamoto Ryoma), and cultural information (Yosakoi Festival) related to Kochi Prefecture from an external information acquisition API, and generates the following scenario:

[0127] "One day, the statue of Sakamoto Ryoma, the symbol of Kochi Prefecture, begins to speak, and an adventure begins as he guides you around local tourist attractions."

[0128] Based on this scenario, the video generation AI draws beautiful scenery and tourist spots in Kochi Prefecture, and the voice generation AI generates narration and character voices. Finally, these are integrated to create the completed animation content, which is saved in cloud storage and a link is provided to the user.

[0129] In this way, the system based on this invention can effectively promote tourist attractions by simply entering the name of a region and automatically collecting related information and generating animation content. This system makes it possible to widely publicize regional tourist resources easily and at low cost.

[0130] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0131] Step 1: User enters locality name

[0132] The user uses the terminal to input the name of a region on the system's web browser. For example, they input "Kochi Prefecture." At this time, the input data is recorded on the terminal as a string including the region name.

[0133] Step 2: Send input data to the server

[0134] The terminal sends the region name entered by the user (e.g., "Kochi Prefecture") as an HTTP request to the server. This request includes the region name in JSON format. Input data: "Kochi Prefecture", Output data: JSON data as a request to the server.

[0135] Step 3: Server gathers external information

[0136] Based on the received region name, the server calls an external information acquisition API (e.g. Wikipedia API, Google Maps API) to collect tourist information, historical information, and cultural information. For example, it obtains information such as "Katsurahama Beach," "Shimanto River," "Sakamoto Ryoma," and "Yosakoi Festival." Input data: region name "Kochi Prefecture," output data: JSON data including tourist information, historical information, and cultural information.

[0137] Step 4: Request to Scenario Generation AI

[0138] Based on the collected information, the server sends the following prompt to the scenario generation AI (e.g., GPT-4):

[0139] Create an anime scenario based on Kochi Prefecture's tourist attractions (Katsurahama Beach, Shimanto River), history (Sakamoto Ryoma), and culture (Yosakoi Festival). The theme is "One day, the statue of Sakamoto Ryoma begins to talk, and an adventure begins in which he guides people around local tourist attractions."

[0140] As a result, a scenario is generated. Input data: tourist information, historical information, cultural information. Output data: generated scenario.

[0141] Step 5: Request to the video generation AI

[0142] The server sends the generated scenario to the image generation AI (e.g. DALL-E, Disco Diffusion) and uses the following prompt:

[0143] "Based on a given scenario, generate an animated video that includes scenery and tourist attractions in Kochi Prefecture (Katsurahama Beach, Shimanto River)."

[0144] This generates video data based on the scenario. Input data: scenario, output data: generated video.

[0145] Step 6: Request to speech generation AI

[0146] The server sends the generated scenario and video data to a voice generation AI (e.g., Google Cloud Text-to-Speech, Amazon Polly) and uses the following prompt:

[0147] "Generate narration, character voices, and background music for the following scenario."

[0148] This generates narration, character voices, and background music. Input data: scenario and video, output data: narration, character voices, and background music.

[0149] Step 7: Integrating animated content

[0150] The server integrates the generated scenario, video, and audio using video editing software such as Adobe Premiere Pro or FFmpeg. For example, it uses FFmpeg commands to combine video and audio to create a single animation file. Input data: scenario, video, audio. Output data: integrated animation content.

[0151] Step 8: Storing and serving content to cloud storage

[0152] The server uploads the completed animation content to cloud storage (e.g., Amazon S3, Google Drive) and generates a URL or download link. Finally, it notifies the user of this link. Input data: animation content, Output data: cloud storage URL or download link.

[0153] Through the above processing steps, users can easily create animation content that includes local tourist attractions and culture and widely disseminate it.

[0154] (Application example 1)

[0155] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0156] Visually appealing content is necessary to effectively communicate the appeal of tourist destinations. However, there is a lack of systems that can effectively collect detailed tourist, historical, and cultural information about local areas and automatically generate scenarios, video, and audio based on that information. This makes it difficult to widely communicate the appeal of a region and increase the number of tourists. Furthermore, a method for easily providing the generated content to users is also needed.

[0157] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0158] In this invention, the server includes means for inputting the name of a region, means for collecting information based on the name of the region, means for generating a scenario based on the collected information, means for generating video based on the generated scenario, means for generating audio corresponding to the video, means for creating content by integrating the scenario, video, and audio, and content distribution means including a display device for visually presenting the created content. This makes it possible to effectively collect tourist information, historical information, and cultural information about a region, automatically generate content consisting of a scenario, video, and audio created based on this information, and visually provide it to users.

[0159] The "means for inputting the name of a region" refers to an interface that a user uses to specify a particular region, and includes, for example, keyboard input and voice input.

[0160] "Means of collecting information" refers to the function of obtaining data such as tourist information, historical information, and cultural information about the designated region from external APIs and databases.

[0161] "Scenario generation means" refers to algorithms or models that create narratives or descriptions based on collected information, and in this case includes generative AI.

[0162] "Means for generating images" refers to tools or models that create visual images based on the generated scenario, including image generation AI.

[0163] "Means for generating audio" refers to tools and models for generating narration, sound effects, and background music corresponding to video, including audio generation AI.

[0164] "Means for creating content" refers to the process or system that integrates the generated scenario, video, and audio into a series of content.

[0165] "Content distribution means" refers to a system for providing created content in a form that users can enjoy visually, and includes smartphone apps, web platforms, head-mounted displays, etc.

[0166] "Tourist information" refers to information about major tourist destinations, attractions, and tourist facilities in the designated region.

[0167] "Historical information" refers to information about historical events, people, and ruins that exist in the designated region.

[0168] "Cultural information" refers to information about festivals, traditional events, cultural customs, etc. that are characteristic of the designated region.

[0169] A "display device" is a hardware device for visually presenting generated content to a user, including smartphones, tablets, visual augmentation devices, etc.

[0170] The system for implementing this invention is composed of multiple components including a user terminal, a server, a scenario generation AI, a video generation AI, a sound generation AI, and an external information acquisition API.

[0171] System Configuration

[0172] User terminal

[0173] The user terminal provides an interface for inputting the name of a region. This interface runs on a web browser or is provided as a dedicated mobile app. The system starts working when the user inputs the name of a region into the terminal.

[0174] server

[0175] The server is the core of the entire system, receiving the locality name sent from the user terminal and proceeding with the processing.

[0176] The server does the following:

[0177] 1. Information gathering

[0178] The server calls an external information retrieval API based on the region name to collect tourist information, historical information, and cultural information. This API call allows for the collection of a wide range of information about the region.

[0179] 2. Scenario Generation

[0180] The server calls a scenario generation AI based on the collected information to generate a scenario related to the relevant region. For example, ChatGPT (registered trademark) is used as the generation AI model for this scenario generation. An example of a specific prompt is, "Please generate a scenario introducing tourist attractions, historical figures, and cultural events related to this region."

[0181] 3. Image Generation

[0182] The server calls the video generation AI based on the generated scenario, and generates a video based on the scenario. An example of a prompt in this case would be, "Please generate an animated video introducing tourist attractions based on the following scenario."

[0183] 4. Speech Generation

[0184] The server then uses the AI ​​to generate audio for the generated video, such as narration and background music. An example of a prompt for voice generation is, "Please generate narration and background music that are appropriate for the video below."

[0185] 5. Content Integration and Delivery

[0186] The server integrates the scenario, video, and audio to create the final animation content. This integrated content is stored in cloud storage and a URL or download link is generated. The link is provided to the user's device, allowing the user to visually enjoy the generated content.

[0187] This system makes it possible to effectively collect tourist, historical, and cultural information about a specific region, automatically generate it as attractive multimedia content, and provide it visually, which is expected to greatly promote tourism and revitalize the region.

[0188] Specific examples

[0189] For example, if a user inputs "a certain region" as "a region they would like to visit," the system will operate as follows:

[0190] 1. User Input

[0191] The user enters "a certain region" into the terminal.

[0192] 2. Information gathering

[0193] The server collects tourist information, historical information, and cultural information about a "certain region" from an external API.

[0194] 3. Scenario Generation

[0195] The server sends the collected information to a scenario generation AI, which generates a scenario related to a "certain region."

[0196] 4. Image Generation

[0197] The server sends the scenario to the video generation AI, which generates a video including tourist attractions in a certain region.

[0198] 5. Speech Generation

[0199] The server uses a voice generation AI to generate narration, character voices, and background music that correspond to the generated video.

[0200] 6. Content Integration and Delivery

[0201] The server integrates the scenario, video, and audio, and saves the completed animation content in cloud storage. Users can view this content via the provided URL.

[0202] Through this process, users can enjoy tourist attractions and culture in a visually rich way, and the appeal of local areas can be effectively promoted.

[0203] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0204] Step 1:

[0205] The user inputs the name of a region into the terminal, either via keyboard or voice input, which is received by the user interface and sent to the server for further processing.

[0206] Input: Name of region

[0207] Output: Locality name data sent to the server

[0208] Step 2:

[0209] The server receives the name of the region sent from the user's device and sends a request to the external information acquisition API based on that name. The server collects tourist information, historical information, and cultural information obtained from the external information acquisition API.

[0210] Input: Name of region

[0211] Output: Tourist information, historical information, cultural information

[0212] Step 3:

[0213] The server passes the collected information to the scenario generation AI and sends a request to generate a scenario for the relevant region. Here, ChatGPT is used as the generation AI model. When generating this scenario, the following prompt is used: "Please generate a scenario introducing tourist attractions, historical figures, and cultural events related to this region."

[0214] Input: Tourist information, historical information, cultural information

[0215] Output: Scenario

[0216] Step 4:

[0217] The server passes the generated scenario to the video generation AI and sends a request to generate a video based on the scenario. When generating the video, the following prompt is used: "Please generate an animated video introducing tourist attractions based on the following scenario."

[0218] Input: Scenario

[0219] Output: Video data

[0220] Step 5:

[0221] The server then requests the AI ​​to generate audio such as narration and background music that corresponds to the generated video. When generating the audio, the following prompt is used: "Please generate narration and background music that are appropriate for the video below."

[0222] Input: Video data

[0223] Output: Audio data

[0224] Step 6:

[0225] The server integrates the scenario, video, and audio to create the completed animation content. This integration process generates a series of multimedia content that users can enjoy visually.

[0226] Input: Scenario, video data, audio data

[0227] Output: Integrated animation content

[0228] Step 7:

[0229] The server stores the generated content in cloud storage and generates a URL or download link for it, which is provided to the user's device so that the user can view the generated content.

[0230] Input: Integrated animation content

[0231] Output: Cloud storage URL or download link

[0232] Through the above processing steps, the system can collect information on local tourist attractions, historical figures, and cultural events, automatically generate and integrate scenarios, images, and audio, and provide them visually to users.

[0233] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0234] The system of the present invention automatically generates animated content related to a region by inputting the name of the region, thereby widely publicizing the region's tourist attractions. The system also includes an emotion engine that recognizes the user's emotions and has the function of personalizing content based on the user's emotions. A specific embodiment of the system and the processing of its program are described below.

[0235] System Configuration

[0236] The system consists of multiple components, including a user device, a server, a scenario generation AI, a video generation AI, a voice generation AI, an emotion engine, and an external information acquisition API. The role of each component is explained below.

[0237] User terminal

[0238] The user terminal provides an interface for inputting the names of regions. This interface generally runs on a web browser and is designed to be easy for users to operate. It also has the function of sending the user's input and reactions to the emotion engine.

[0239] server

[0240] The server is the core of the entire system. It receives the name of the region sent from the user's device and processes requests to the scenario generation AI, video generation AI, and audio generation AI in sequence. It also calls an external information acquisition API to collect tourist information, historical information, and cultural information about the region.

[0241] Emotion Engine

[0242] The emotion engine analyzes the user's input and reactions while watching to recognize the user's emotions. Based on the recognized emotions, the emotion engine adjusts scenario generation, video generation, or audio generation to provide personalized content according to the user's emotions.

[0243] Program processing

[0244] The processing of the program will be explained in natural language below.

[0245] 1. User Input

[0246] The user inputs the name of a region (e.g., "Kochi Prefecture") using the terminal.

[0247] 2. Send to the server

[0248] The terminal sends the entered name of the region to the server as an HTTP request.

[0249] 3. Information gathering

[0250] The server uses an external information acquisition API based on the local name to collect tourist information, historical information, and cultural information.

[0251] 4. Emotion recognition

[0252] The device sends the user's input and reactions to the emotion engine.

[0253] The emotion engine recognizes the user's emotion and notifies the server.

[0254] 5. Scenario Generation

[0255] The server passes the collected information to a scenario generation AI, which then generates a scenario for an anime related to the region.

[0256] The scenario generation AI personalizes the scenario based on the user's emotions.

[0257] 6. Image Generation

[0258] The server passes the generated scenario to the video generation AI, which then generates a video based on the scenario.

[0259] The video generation AI adjusts the video based on the user's emotions.

[0260] 7. Speech Generation

[0261] The server passes the generated scenario and video to an audio generation AI, which generates narration, character voices, and background music.

[0262] The voice generation AI adjusts the voice based on the user's emotions.

[0263] 8. Content Integration

[0264] The server integrates the scenario, video, and audio to create the final animation content.

[0265] 9. Content Provision

[0266] The server stores the generated content in cloud storage and generates and provides a URL or download link for the content.

[0267] Specific examples

[0268] As a specific example of use, consider the case where "Kochi Prefecture" is entered.

[0269] 1. The user enters "Kochi Prefecture" into the terminal.

[0270] 2. The device sends the input information to the server.

[0271] 3. The server collects tourist information (e.g., Katsurahama Beach, Shimanto River), historical information (e.g., Sakamoto Ryoma), and cultural information (e.g., Yosakoi Festival) related to Kochi Prefecture from external APIs.

[0272] 4. The device sends the user's input and reactions to the emotion engine.

[0273] 5. The emotion engine recognizes the user's emotion and notifies the server.

[0274] 6. The server sends a request to the scenario generation AI based on the collected information and generates an anime scenario.

[0275] The scenario generation AI personalizes the scenario based on the user's emotions.

[0276] 7. The server requests the video generation AI to generate a video based on the generated scenario, creating a video that includes scenery and tourist attractions in Kochi Prefecture.

[0277] The video generation AI adjusts the video based on the user's emotions.

[0278] 8. The server then uses the voice generation AI to create narration, character voices, and background music that correspond to the generated scenario and video.

[0279] The voice generation AI adjusts the voice based on the user's emotions.

[0280] 9. The server integrates the scenario, video, and audio, and saves the completed animation content in cloud storage.

[0281] 10. The server generates a destination URL or download link and notifies the user.

[0282] In this way, the system based on the present invention can effectively promote local tourist attractions, increase the number of tourists, and promote regional revitalization. It can also provide more attractive content by recognizing users' emotions and personalizing content based on those emotions.

[0283] The processing flow will be explained below.

[0284] Step 1:

[0285] The user uses a terminal to input the name of a region (e.g., "Kochi Prefecture").

[0286] Step 2:

[0287] The terminal sends the inputted name of the region to the server as an HTTP request.

[0288] Step 3:

[0289] Based on the name of the region received by the server, tourist information, historical information, and cultural information are collected using an external information acquisition API.

[0290] Step 4:

[0291] The server passes the collected information to a scenario generation AI, which then generates a scenario for an anime related to the region.

[0292] Step 5:

[0293] The device transmits the user's input and reactions (emotions) to the emotion engine in real time.

[0294] Step 6:

[0295] The emotion engine recognizes the user's emotions and sends feedback to the scenario generation AI, which then adjusts the content of the scenario.

[0296] Step 7:

[0297] The server receives the adjusted scenario from the scenario generation AI and passes it to the video generation AI.

[0298] Step 8:

[0299] The video generation AI generates video based on the scenario and returns the generated video data to the server.

[0300] Step 9:

[0301] The server passes the generated video data and scenario to the voice generation AI, which then generates the corresponding narration, character voices, and background music.

[0302] Step 10:

[0303] The emotion engine sends feedback to the voice generation AI based on the user's emotions, allowing it to adjust the content and tone of the voice.

[0304] Step 11:

[0305] The voice generation AI returns the adjusted voice data to the server, which then integrates the scenario, video, and voice to create the final animation content.

[0306] Step 12:

[0307] The server saves the completed animation content in cloud storage and generates a distribution URL or download link.

[0308] Step 13:

[0309] The server sends a notification to the user containing a distribution URL or a download link, which the user receives to view or download the animation content.

[0310] As a concrete example, in the case of "Kochi Prefecture," the user inputs "Kochi Prefecture," and tourist information (Katsurahama, Shimanto River), historical information (Sakamoto Ryoma), and cultural information (Yosakoi Festival) are collected via an external API. The emotion engine recognizes emotions based on the user's input and reactions, and feeds this back to the scenario generation AI, video generation AI, and audio generation AI, ultimately creating and providing animation content optimized to the user's emotions.

[0311] Example 2

[0312] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0313] Conventional tourist information presentation methods have difficulty personalizing content based on the user's emotions and interests. Furthermore, they lack a mechanism for automatically collecting related tourist, historical, and cultural information by simply entering the name of a region, and generating visually and auditorily appealing content based on that information.

[0314] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting the name of a region, means for collecting information based on the name of the region, means for generating a scenario based on the collected information, means for generating video based on the generated scenario, means for generating audio corresponding to the video, means for creating content by integrating the scenario, video, and audio, means for providing the created content, means for analyzing the user's input and operation to recognize emotions, and means for adjusting the scenario, video, and audio based on the emotions. This makes it possible to automatically generate and provide attractive tourist information content personalized according to the user's emotions.

[0315] "Means for inputting locality names" refers to an interface through which a user can input a particular locality name into the system.

[0316] "Means for collecting information based on the name of the region" refers to means for obtaining tourist information, historical information, and cultural information related to the input region name from external databases or APIs.

[0317] "Means for generating a scenario based on collected information" refers to the part of the system that automatically creates a story or narration structure based on the information obtained.

[0318] "Means for generating video based on a generated scenario" refers to a function that automatically produces animation or video content according to a created scenario.

[0319] The "means for generating audio corresponding to the video" refers to a function for generating audio data such as narration, character voices, background music, etc., in accordance with the content of the video.

[0320] The "means for creating content by integrating the scenario, video and audio" refers to a process for creating integrated multimedia content by integrating the generated scenario, video and audio.

[0321] "Means for providing the created content" refers to means for providing the created content to users in an accessible form, such as storing it in cloud storage and creating a URL.

[0322] The "means for analyzing the user's inputs and operations to recognize emotions" refers to the part of the system that analyzes the user's interface operations and visual reactions, etc., to determine the user's emotional state.

[0323] "Means for adjusting the scenario, video and audio based on the emotion" refers to a function for adjusting the content and tone of the generated scenario, video and audio based on the data obtained by emotion recognition.

[0324] The system of the present invention automatically generates animation content related to a region by inputting the name of the region, thereby widely publicizing tourist attractions. This system is composed of multiple components, including a user terminal, a server, a scenario generation AI, a video generation AI, a voice generation AI, an emotion engine, and an external information acquisition API. The role of each component and specific implementation methods are explained below.

[0325] User terminal

[0326] The user terminal provides an interface for inputting the name of a region. This interface runs on a web browser and is designed to be easy for users to operate. It also has a function for sending the user's input and reactions to the emotion engine. Specifically, the user accesses the system's website using a web browser (e.g., GOOGLE CHROME (registered trademark)) and inputs the name of a region (e.g., "Kochi Prefecture") into the input form.

[0327] server

[0328] The server is the core of the entire system. The server (e.g., AWS (registered trademark) EC2) receives the name of the region sent from the user's device and sequentially processes requests to the scenario generation AI, video generation AI, and audio generation AI. It also calls external information acquisition APIs to collect tourist information, historical information, and cultural information about the region. For example, it acquires information using the Google Maps API or Wikipedia API. The collected information is passed to the scenario generation AI and emotion engine.

[0329] Emotion Engine

[0330] The emotion engine analyzes the user's input and reactions while watching to recognize the user's emotions. Based on the recognized emotions, the emotion engine (e.g., Microsoft® Azure® Emotion API) adjusts scenario generation, video generation, or audio generation to provide personalized content according to the user's emotions. For example, if the user's reaction is "surprise," it adds a corresponding surprise element to the scenario.

[0331] Scenario Generation

[0332] The server passes the collected information to a scenario generation AI (e.g., OpenAI GPT-4), which then generates an anime scenario based on the collected local information. The scenario generation AI then personalizes the scenario based on the user's emotional data. An example of a prompt could be, "Based on tourist, historical, and cultural information about Kochi Prefecture, please generate a scenario that will make the user feel 'surprised.'"

[0333] Image Generation

[0334] The server passes the generated scenario to a video generation AI (e.g., DALL-E 2), which generates a video based on the scenario. The video generation AI adjusts the video based on the user's emotional data and creates a video that includes scenery and tourist attractions in Kochi Prefecture.

[0335] Voice generation

[0336] The server passes the generated scenario and video to a voice generation AI (e.g., Amazon Polly), which generates narration, character voices, and background music. The voice generation AI adjusts the voice based on the user's emotional data and generates voice data that creates an appropriate atmosphere.

[0337] Content Integration

[0338] The server integrates the script, video, and audio to create the final animation content. This integration process uses open-source video editing software (e.g., FFmpeg), ensuring that each element is properly positioned and timed.

[0339] Content provider

[0340] The server stores the generated animation content in cloud storage (e.g., Amazon S3) and generates an access URL or download link, through which users can access the generated content.

[0341] As a concrete example, if the theme is "Kochi Prefecture," the server uses the Google Maps API and Wikipedia API to collect information on Kochi Prefecture's tourist attractions (Katsurahama Beach, Shimanto River), history (Sakamoto Ryoma), and culture (Yosakoi Festival). The emotion engine then identifies the user's "surprise," and based on that, OpenAI GPT-4 generates a scenario that includes elements of surprise. DALL-E 2 generates video and Amazon Polly generates audio. Finally, FFmpeg is used to integrate the content, and the generated link is provided to the user.

[0342] In this way, the system of the present invention can effectively promote local tourist attractions, increase the number of tourists, and promote regional revitalization. Furthermore, by recognizing users' emotions and personalizing content based on those emotions, the system can provide more engaging and interactive content.

[0343] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0344] Step 1: User Input

[0345] The user opens the system's website using a web browser (e.g., Google Chrome), enters the name of a region (e.g., "Kochi Prefecture") in the input form, and clicks the "Submit" button.

[0346] Input: Name of region (e.g. "Kochi Prefecture")

[0347] Output: HTTP request (including locality name)

[0348] Step 2: Send to the server

[0349] The device sends the name of the region entered by the user to a server (e.g., AWS EC2) as an HTTP request, which also includes the user's identification information.

[0350] Input: HTTP request (region name and user information)

[0351] Output: Send data to the server

[0352] Step 3: Gather information

[0353] Based on the received name of the region, the server uses an external information acquisition API (e.g., Google Maps API, Wikipedia API) to collect tourist information, historical information, and cultural information related to that region.

[0354] The server formats the temporarily stored information and stores it in a database.

[0355] Input: Name of region

[0356] Output: A dataset of collected tourist, historical, and cultural information

[0357] Step 4: Emotion Recognition

[0358] The device monitors user input and actions (e.g., clicks, eye movements) in real time and sends them to an emotion engine (e.g., Microsoft Azure Emotion API).

[0359] The emotion engine analyzes the received data, recognizes the user's emotion, and notifies the server.

[0360] Input: User operation data

[0361] Output: Emotion data (e.g., "surprise")

[0362] Step 5: Scenario generation

[0363] The server passes the collected local information and emotion data to a scenario generation AI (e.g., OpenAI GPT-4), which then generates a scenario based on that data.

[0364] Example prompt: "Generate a scenario that will surprise the user based on tourist, historical, and cultural information about Kochi Prefecture."

[0365] Input: local information, emotion data, prompt sentence

[0366] Output: Generated scenario

[0367] Step 6: Image generation

[0368] The server passes the generated scenario to a video generation AI (e.g., DALL-E 2), which generates a video based on the scenario.

[0369] The video generation AI adjusts the video based on the user's emotional data and returns the appropriate video to the server.

[0370] Input: Generated scenario, emotion data

[0371] Output: Generated video data

[0372] Step 7: Speech generation

[0373] The server passes the generated scenario and video to an audio generation AI (e.g., a voice synthesis API), which generates narration, character voices, and background music.

[0374] The voice generation AI adjusts the tone and pitch of the voice based on the user's emotional data to generate appropriate voice data.

[0375] Input: Generated scenario, video data, emotion data

[0376] Output: Generated audio data

[0377] Step 8: Content Integration

[0378] The server integrates the generated scenario, video, and audio to create the final animation content, using video editing software (e.g., FFmpeg).

[0379] Input: Scenario, video data, audio data

[0380] Output: Integrated animation content

[0381] Step 9: Provide content

[0382] The server uploads the completed animation content to cloud storage (e.g., cloud storage service) and generates an access URL or download link for it. The generated link is notified to the user.

[0383] Input: Integrated animation content

[0384] Output: Access URL or download link

[0385] (Application example 2)

[0386] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0387] Current tourism promotion methods are limited to providing general information, making it difficult to provide personalized experiences that reflect the interests and emotions of individual users. They also have limitations as a means of effectively publicizing local tourist attractions and culture. Therefore, there is a need for a system that can recognize users' emotions and provide personalized content based on them.

[0388] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting the name of a region, means for collecting information based on the name of the region, means for generating a scenario based on the collected information, means for generating video based on the generated scenario, means for generating audio corresponding to the video, means for creating content by integrating the scenario, video, and audio, means for providing the created content, means for recognizing the user's emotions, and means for personalizing the content based on the recognized emotions. This makes it possible to provide personalized tourist content that reflects the user's emotions.

[0389] The "means for inputting the name of a region" refers to a device or program that provides an interface for the user to input the name of a region.

[0390] "Means for collecting information" refers to devices or programs that obtain tourist information, historical information, and cultural information from external APIs and databases.

[0391] A "means for generating a scenario" is a device or program that generates an animation or video story based on collected information.

[0392] "Video generation means" refers to devices or programs that create animations or visual content based on the generated scenario.

[0393] "Audio generating means" refers to a device or program that generates narration or character voices corresponding to the video.

[0394] The "content creating means" refers to a device or program that integrates the generated scenario, video, and audio to create the final multimedia content.

[0395] "Means for providing content" refers to devices or programs that distribute or provide created content to users.

[0396] "Means for recognizing emotions" refers to devices or programs that analyze emotions from the user's facial expressions and voice.

[0397] A "means for personalizing content based on recognized emotions" is a device or program that adjusts the content or expression of content according to the user's emotions.

[0398] The system of the present invention includes a means for inputting the name of a region, a means for collecting information, a means for generating a scenario, a means for generating video, a means for generating audio, a means for integrating and creating content, a means for providing the created content, a means for recognizing emotions, and a means for personalizing content based on the recognized emotions. Each component of the system functions as follows.

[0399] Hardware and software used

[0400] Hardware:

[0401] User devices: smartphones, tablets, etc.

[0402] Server: Cloud server

[0403] software:

[0404] Scenario generation AI: OpenAI GPT-4 as an example

[0405] Image generation AI: DALL-E as an example

[0406] Speech generation AI: Google Cloud Text-to-Speech as an example

[0407] Emotion engine: Microsoft Azure Emotion API as an example

[0408] External information acquisition API: Examples include the Google Places API and the TriPad (registered trademark) visor API

[0409] Cloud storage: AWS S3 as an example

[0410] Program processing

[0411] 1. How to enter the name of a region:

[0412] A user inputs the name of a region through an application interface on a smartphone or tablet. For example, the user inputs "Kyoto Prefecture."

[0413] 2. How we collect information:

[0414] The server uses an external information acquisition API to collect tourist, historical, and cultural information based on the local name, such as information about "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Gion Festival."

[0415] 3. How to recognize emotions:

[0416] The user's facial expressions and voice are recorded using the smartphone's camera and microphone, and analyzed by the emotion engine, allowing the system to recognize emotions in real time while the user waits for the animation to be generated.

[0417] 4. How to generate scenarios:

[0418] The server inputs the collected information into a scenario generation AI to generate an animated scenario. At this time, the scenario content is adjusted based on the user's emotions. For example, if the user is surprised, a scenario that reflects that emotion is generated.

[0419] 5. Means of generating images:

[0420] The server then passes the generated scenario to the image generation AI, which then creates an animation based on it. The image expression can be adjusted according to the user's emotions.

[0421] 6. Means of generating sound:

[0422] The server uses a voice generation AI to create narration and character voices that correspond to the scenario and video. For example, a lively sound is generated for a festival scene.

[0423] 7. Means of integrating and creating content:

[0424] The server integrates the scenario, video and audio to create the final animation content, which is personalized according to the user's emotions.

[0425] 8. Means of providing created content:

[0426] The created animation content is stored in cloud storage and the URL or download link is provided to the user.

[0427] Examples of concrete examples and prompts

[0428] As a usage example, consider the case where a user types "Kyoto Prefecture." The system behaves as follows:

[0429] The user enters "Kyoto Prefecture."

[0430] The server collects tourist information related to Kochi Prefecture (e.g., Kinkaku-ji Temple, Kiyomizu-dera Temple, Gion Festival).

[0431] The user's emotion is analyzed as "surprised."

[0432] Scenario generation AI generates scenarios that reflect surprise and excitement.

[0433] Video generation AI generates videos of festival scenes and other events.

[0434] Voice generation AI creates lively voices.

[0435] These are integrated to create animated content and provide a URL that is saved in cloud storage.

[0436] An example prompt might look like this:

[0437] "The destination is Kyoto Prefecture. Please generate a scenario that includes information about the following tourist attractions: Kinkaku-ji Temple, Kiyomizu-dera Temple, and Gion Festival. Also, since the user will be judged as surprised, please create content that elicits surprise and excitement."

[0438] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0439] Step 1:

[0440] The user inputs the name of a region through the terminal. The user inputs the region name (e.g., "Kyoto Prefecture") in the application interface of the terminal and presses the send button. This input data becomes the information to be sent directly to the server.

[0441] Step 2:

[0442] The server receives the input data and calls an external information acquisition API to collect information. Specifically, it uses the Google Places API and TripAdvisor API to collect tourist information, historical information, and cultural information related to "Kyoto Prefecture." The collected information is returned to the server in JSON format.

[0443] Step 3:

[0444] The server uses an emotion engine to recognize the user's emotions. The user's facial expressions and voice, captured by the camera and microphone on the user's device, are analyzed in real time and sent to the emotion engine (Microsoft Azure Emotion API). The emotion engine returns the analysis results to the server, which stores them in an internal database.

[0445] Step 4:

[0446] The server creates a prompt for the scenario generation AI based on the collected local information and user emotion data. For example, it creates a prompt such as, "The destination is Kyoto Prefecture. Please generate a scenario that includes information about the following tourist attractions: Kinkaku-ji Temple, Kiyomizu-dera Temple, and Gion Festival. Also, since the user is judged to be surprised, please create content that elicits surprise and excitement." and sends it to the scenario generation AI (OpenAI GPT-4). The scenario generation AI generates a scenario and returns it to the server.

[0447] Step 5:

[0448] The server passes the generated scenario to the video generation AI and requests it to generate an animation video. The video generation AI (DALL-E) generates visual content based on the scenario and returns it to the server. At this time, the video expression is adjusted based on the user's emotional data.

[0449] Step 6:

[0450] The server requests the voice generation AI to generate voice based on the scenario and video data. The voice generation AI (Google Cloud Text-to-Speech) generates the corresponding voice, narration, and character voices based on the scenario and returns them to the server.

[0451] Step 7:

[0452] The server integrates the generated scenario, video, and audio data to create the final animation content, and stores the integrated content in its internal storage.

[0453] Step 8:

[0454] The server uploads the final content to cloud storage (AWS S3) and generates a URL or download link for it, which is then sent to the user's device so that the user can view or download the content.

[0455] Through this series of steps, personalized tourism content is generated and provided based on the user's emotions.

[0456] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0457] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0458] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0459] [Second embodiment]

[0460] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0461] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0462] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0463] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0464] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0465] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0466] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0467] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0468] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0469] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0470] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0471] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0472] The system of the present invention is designed to automatically generate animation content related to a region by inputting the name of the region, thereby widely publicizing the tourist spots. Specific embodiments of the system and the processing of the program therefor are described below.

[0473] System Configuration

[0474] The system is composed of multiple components, including a user terminal, a server, a scenario generation AI, a video generation AI, an audio generation AI, and an external information acquisition API. The role of each component is explained below.

[0475] User terminal

[0476] The user terminal provides an interface for inputting the name of a region. This interface generally runs on a web browser and is designed to be easy for users to operate.

[0477] server

[0478] The server is the core of the entire system. It receives the name of the region sent from the user's device and processes requests to the scenario generation AI, video generation AI, and audio generation AI in sequence. It also calls an external information acquisition API to collect tourist information, historical information, and cultural information about the region.

[0479] Program processing

[0480] The processing of the program will be explained in natural language below.

[0481] 1. User Input

[0482] The user inputs the name of a region (e.g., "Kochi Prefecture") using the terminal.

[0483] 2. Send to the server

[0484] The terminal transmits the entered name of the region to the server.

[0485] 3. Information gathering

[0486] The server calls an external information acquisition API based on the region name and collects tourist information, historical information, and cultural information.

[0487] 4. Scenario Generation

[0488] The server passes the collected information to a scenario generation AI, which generates scenarios relevant to the region.

[0489] 5. Image Generation

[0490] The server passes the generated scenario to the video generation AI, which then generates a video based on the scenario.

[0491] 6. Speech Generation

[0492] The server passes the generated scenario and video to an audio generation AI, which generates narration, character voices, and background music that correspond to the scenario.

[0493] 7. Content Integration

[0494] The server integrates the scenario, video, and audio to create the final animation content.

[0495] 8. Content Provision

[0496] The server stores the generated content in cloud storage and generates and provides a URL or download link for the content.

[0497] Specific examples

[0498] As a specific example of use, consider the case where "Kochi Prefecture" is entered.

[0499] 1. The user enters "Kochi Prefecture" into the terminal.

[0500] 2. The device sends the input information to the server.

[0501] 3. The server collects tourist information (e.g., Katsurahama Beach, Shimanto River), historical information (e.g., Sakamoto Ryoma), and cultural information (e.g., Yosakoi Festival) related to Kochi Prefecture from external APIs.

[0502] 4. The server sends a request to the scenario generation AI based on the collected information and generates an anime scenario.

[0503] 5. The server requests the video generation AI to generate a video based on the generated scenario, creating a video that includes scenery and tourist attractions in Kochi Prefecture.

[0504] 6. The server then uses the voice generation AI to create narration, character voices, and background music that correspond to the generated scenario and video.

[0505] 7. The server integrates the scenario, video, and audio, and saves the completed animation content in cloud storage.

[0506] 8. The server generates a URL or download link for the destination and notifies the user.

[0507] In this way, the system based on the present invention effectively promotes local tourist attractions, increasing the number of tourists and revitalizing the region.

[0508] The processing flow will be explained below.

[0509] Step 1:

[0510] The user uses a terminal to input the name of a region (e.g., "Kochi Prefecture").

[0511] Step 2:

[0512] The terminal sends the inputted name of the region to the server as an HTTP request.

[0513] Step 3:

[0514] Based on the name of the region received by the server, tourist information, historical information, and cultural information are collected using an external information acquisition API.

[0515] Step 4:

[0516] The server passes the collected information to a scenario generation AI, which then generates a scenario for an anime related to the region.

[0517] Step 5:

[0518] The scenario generation AI returns the generated scenario to the server, which then passes it on to the video generation AI.

[0519] Step 6:

[0520] The video generation AI generates video based on the scenario and returns the generated video data to the server.

[0521] Step 7:

[0522] The server passes the generated video data and scenario to the voice generation AI, which then generates the corresponding narration, character voices, and background music.

[0523] Step 8:

[0524] The voice generation AI returns the generated voice data to the server, which then integrates the scenario, video, and voice to create the final animation content.

[0525] Step 9:

[0526] The server saves the completed animation content in cloud storage and generates a distribution URL or download link.

[0527] Step 10:

[0528] The server sends a notification to the user containing a distribution URL or a download link, which the user receives to view or download the animation content.

[0529] Example 1

[0530] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0531] Currently, many regions lack the means to effectively promote their tourism resources, making it difficult to attract tourists' attention. Expressing a region's history and culture through video and audio is also problematic, requiring significant costs and time. To address this issue, there is a demand for a system that can easily and automatically generate content that conveys a region's tourist attractions and culture in an appealing way.

[0532] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0533] In this invention, the server includes means for collecting tourist information, historical information, and cultural information based on the name of a region, means for generating a scenario using a generation AI model based on the collected information, means for sending the generated scenario to a video generation AI to generate video, means for sending the generated scenario and video to an audio generation AI to generate audio, means for creating content by integrating the scenario, video, and audio, and means for saving the created content and generating a URL or download link for it, which makes it possible to effectively promote local tourist attractions and culture and increase the number of tourists.

[0534] A "local name" is a name used to identify a particular area, usually indicating a geographical, administrative or historical division of that area.

[0535] "Information gathering means" refers to the methods and devices used to obtain data from external sources and convert it into a form that can be used within the system.

[0536] The term "means for generating a scenario" refers to a method or device that automatically creates text data containing elements of a narrative or storytelling based on collected information.

[0537] "Means for generating video" refers to a method or device for automatically creating visual video content based on text data or a scenario.

[0538] "Audio generating means" refers to a method or device for creating audio data corresponding to a scenario or video, including narration, character voices, background music, etc.

[0539] The "means for creating content" refers to a method or device for integrating the generated scenario, video and audio to create a complete media content.

[0540] "Cloud storage" refers to an online storage service that stores data on remote servers on the Internet and allows it to be accessed and shared as needed.

[0541] "Generative AI model" refers to an artificial intelligence model that has been trained using machine learning to perform a specific task (e.g., text generation, image generation, speech generation).

[0542] A "prompt" is an instruction given to a generative AI model, and refers to text that guides the model in determining what output to generate.

[0543] The system of the present invention automatically generates content (especially animations) related to a region by inputting the name of the region, thereby widely publicizing the tourist spots. The system is composed of multiple components, including a user terminal, a server, a scenario generation AI, a video generation AI, a sound generation AI, and an external information acquisition API.

[0544] User terminal

[0545] The user terminal provides an interface for inputting the name of a region. This interface generally runs on a web browser and is designed to be easy for users to operate. The user inputs the name of the region through this interface.

[0546] server

[0547] The server is the core of the entire system and performs the following processes:

[0548] 1. Receive the name of the region sent from the user terminal.

[0549] 2. Call external information acquisition APIs to collect tourist information, historical information, and cultural information. For example, use the Wikipedia API or Google Maps API.

[0550] 3. Based on the collected information, a scenario generation AI (e.g., OpenAI's GPT-4) is requested to generate a scenario. An example of a prompt sentence for this is as follows:

[0551] "Create a scenario related to tourist attractions, history, and culture based on the local name. For example, for "Kochi Prefecture," generate a scenario that includes "Katsurahama Beach," "Shimanto River," "Sakamoto Ryoma," and "Yosakoi Festival."

[0552] 4. The generated scenario is sent to an image generation AI (e.g., DALL-E, Disco Diffusion), which generates an image based on the scenario. The following prompts are used:

[0553] "Generate an animated video based on a given scenario, including the scenery and tourist attractions of the target region."

[0554] 5. The generated scenario and video are sent to a voice generation AI (e.g., Google Cloud Text-to-Speech, Amazon Polly) to generate narration, character voices, and background music corresponding to the scenario. The following prompts are used:

[0555] "Generate narration, character voices, and background music for the following scenario."

[0556] 6. The script, video, and audio are integrated to create the final animation content. This integration process is carried out using software such as Adobe Premiere Pro and FFmpeg.

[0557] 7. Save the created content to cloud storage (e.g., Amazon S3, Google Drive) and generate a URL or download link to provide to users.

[0558] Specific examples

[0559] For example, if a user inputs the name of a region, "Kochi Prefecture," the device sends this information to the server. The server collects tourist attractions (Katsurahama Beach, Shimanto River), historical information (Sakamoto Ryoma), and cultural information (Yosakoi Festival) related to Kochi Prefecture from an external information acquisition API, and generates the following scenario:

[0560] "One day, the statue of Sakamoto Ryoma, the symbol of Kochi Prefecture, begins to speak, and an adventure begins as he guides you around local tourist attractions."

[0561] Based on this scenario, the video generation AI draws beautiful scenery and tourist spots in Kochi Prefecture, and the voice generation AI generates narration and character voices. Finally, these are integrated to create the completed animation content, which is saved in cloud storage and a link is provided to the user.

[0562] In this way, the system based on this invention can effectively promote tourist attractions by simply entering the name of a region and automatically collecting related information and generating animation content. This system makes it possible to widely publicize regional tourist resources easily and at low cost.

[0563] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0564] Step 1: User enters locality name

[0565] The user uses the terminal to input the name of a region on the system's web browser. For example, they input "Kochi Prefecture." At this time, the input data is recorded on the terminal as a string including the region name.

[0566] Step 2: Send input data to the server

[0567] The terminal sends the region name entered by the user (e.g., "Kochi Prefecture") as an HTTP request to the server. This request includes the region name in JSON format. Input data: "Kochi Prefecture", Output data: JSON data as a request to the server.

[0568] Step 3: Server gathers external information

[0569] Based on the received region name, the server calls an external information acquisition API (e.g. Wikipedia API, Google Maps API) to collect tourist information, historical information, and cultural information. For example, it obtains information such as "Katsurahama Beach," "Shimanto River," "Sakamoto Ryoma," and "Yosakoi Festival." Input data: region name "Kochi Prefecture," output data: JSON data including tourist information, historical information, and cultural information.

[0570] Step 4: Request to Scenario Generation AI

[0571] Based on the collected information, the server sends the following prompt to the scenario generation AI (e.g., GPT-4):

[0572] Create an anime scenario based on Kochi Prefecture's tourist attractions (Katsurahama Beach, Shimanto River), history (Sakamoto Ryoma), and culture (Yosakoi Festival). The theme is "One day, the statue of Sakamoto Ryoma begins to talk, and an adventure begins in which he guides people around local tourist attractions."

[0573] As a result, a scenario is generated. Input data: tourist information, historical information, cultural information. Output data: generated scenario.

[0574] Step 5: Request to the video generation AI

[0575] The server sends the generated scenario to the image generation AI (e.g. DALL-E, Disco Diffusion) and uses the following prompt:

[0576] "Based on a given scenario, generate an animated video that includes scenery and tourist attractions in Kochi Prefecture (Katsurahama Beach, Shimanto River)."

[0577] This generates video data based on the scenario. Input data: scenario, output data: generated video.

[0578] Step 6: Request to speech generation AI

[0579] The server sends the generated scenario and video data to a voice generation AI (e.g., Google Cloud Text-to-Speech, Amazon Polly) and uses the following prompt:

[0580] "Generate narration, character voices, and background music for the following scenario."

[0581] This generates narration, character voices, and background music. Input data: scenario and video, output data: narration, character voices, and background music.

[0582] Step 7: Integrating animated content

[0583] The server integrates the generated scenario, video, and audio using video editing software such as Adobe Premiere Pro or FFmpeg. For example, it uses FFmpeg commands to combine video and audio to create a single animation file. Input data: scenario, video, audio. Output data: integrated animation content.

[0584] Step 8: Storing and serving content to cloud storage

[0585] The server uploads the completed animation content to cloud storage (e.g., Amazon S3, Google Drive) and generates a URL or download link. Finally, it notifies the user of this link. Input data: animation content, Output data: cloud storage URL or download link.

[0586] Through the above processing steps, users can easily create animation content that includes local tourist attractions and culture and widely disseminate it.

[0587] (Application example 1)

[0588] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0589] Visually appealing content is necessary to effectively communicate the appeal of tourist destinations. However, there is a lack of systems that can effectively collect detailed tourist, historical, and cultural information about local areas and automatically generate scenarios, video, and audio based on that information. This makes it difficult to widely communicate the appeal of a region and increase the number of tourists. Furthermore, a method for easily providing the generated content to users is also needed.

[0590] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0591] In this invention, the server includes means for inputting the name of a region, means for collecting information based on the name of the region, means for generating a scenario based on the collected information, means for generating video based on the generated scenario, means for generating audio corresponding to the video, means for creating content by integrating the scenario, video, and audio, and content distribution means including a display device for visually presenting the created content. This makes it possible to effectively collect tourist information, historical information, and cultural information about a region, automatically generate content consisting of a scenario, video, and audio created based on this information, and visually provide it to users.

[0592] The "means for inputting the name of a region" refers to an interface that a user uses to specify a particular region, and includes, for example, keyboard input and voice input.

[0593] "Means of collecting information" refers to the function of obtaining data such as tourist information, historical information, and cultural information about the designated region from external APIs and databases.

[0594] "Scenario generation means" refers to algorithms or models that create narratives or descriptions based on collected information, and in this case includes generative AI.

[0595] "Means for generating images" refers to tools or models that create visual images based on the generated scenario, including image generation AI.

[0596] "Means for generating audio" refers to tools and models for generating narration, sound effects, and background music corresponding to video, including audio generation AI.

[0597] "Means for creating content" refers to the process or system that integrates the generated scenario, video, and audio into a series of content.

[0598] "Content distribution means" refers to a system for providing created content in a form that users can enjoy visually, and includes smartphone apps, web platforms, head-mounted displays, etc.

[0599] "Tourist information" refers to information about major tourist destinations, attractions, and tourist facilities in the designated region.

[0600] "Historical information" refers to information about historical events, people, and ruins that exist in the designated region.

[0601] "Cultural information" refers to information about festivals, traditional events, cultural customs, etc. that are characteristic of the designated region.

[0602] A "display device" is a hardware device for visually presenting generated content to a user, including smartphones, tablets, visual augmentation devices, etc.

[0603] The system for implementing this invention is composed of multiple components including a user terminal, a server, a scenario generation AI, a video generation AI, a sound generation AI, and an external information acquisition API.

[0604] System Configuration

[0605] User terminal

[0606] The user terminal provides an interface for inputting the name of a region. This interface runs on a web browser or is provided as a dedicated mobile app. The system starts working when the user inputs the name of a region into the terminal.

[0607] server

[0608] The server is the core of the entire system, receiving the locality name sent from the user terminal and proceeding with the processing.

[0609] The server does the following:

[0610] 1. Information gathering

[0611] The server calls an external information retrieval API based on the region name to collect tourist information, historical information, and cultural information. This API call allows for the collection of a wide range of information about the region.

[0612] 2. Scenario Generation

[0613] The server calls a scenario generation AI based on the collected information to generate a scenario related to the relevant region. For example, ChatGPT is used as the generation AI model for this scenario generation. An example of a specific prompt is, "Please generate a scenario introducing tourist attractions, historical figures, and cultural events related to this region."

[0614] 3. Image Generation

[0615] The server calls the video generation AI based on the generated scenario, and generates a video based on the scenario. An example of a prompt in this case would be, "Please generate an animated video introducing tourist attractions based on the following scenario."

[0616] 4. Speech Generation

[0617] The server then uses the AI ​​to generate audio for the generated video, such as narration and background music. An example of a prompt for voice generation is, "Please generate narration and background music that are appropriate for the video below."

[0618] 5. Content Integration and Delivery

[0619] The server integrates the scenario, video, and audio to create the final animation content. This integrated content is stored in cloud storage and a URL or download link is generated. The link is provided to the user's device, allowing the user to visually enjoy the generated content.

[0620] This system makes it possible to effectively collect tourist, historical, and cultural information about a specific region, automatically generate it as attractive multimedia content, and provide it visually, which is expected to greatly promote tourism and revitalize the region.

[0621] Specific examples

[0622] For example, if a user inputs "a certain region" as "a region they would like to visit," the system will operate as follows:

[0623] 1. User Input

[0624] The user enters "a certain region" into the terminal.

[0625] 2. Information gathering

[0626] The server collects tourist information, historical information, and cultural information about a "certain region" from an external API.

[0627] 3. Scenario Generation

[0628] The server sends the collected information to a scenario generation AI, which generates a scenario related to a "certain region."

[0629] 4. Image Generation

[0630] The server sends the scenario to the video generation AI, which generates a video including tourist attractions in a certain region.

[0631] 5. Speech Generation

[0632] The server uses a voice generation AI to generate narration, character voices, and background music that correspond to the generated video.

[0633] 6. Content Integration and Delivery

[0634] The server integrates the scenario, video, and audio, and saves the completed animation content in cloud storage. Users can view this content via the provided URL.

[0635] Through this process, users can enjoy tourist attractions and culture in a visually rich way, and the appeal of local areas can be effectively promoted.

[0636] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0637] Step 1:

[0638] The user inputs the name of a region into the terminal, either via keyboard or voice input, which is received by the user interface and sent to the server for further processing.

[0639] Input: Name of region

[0640] Output: Locality name data sent to the server

[0641] Step 2:

[0642] The server receives the name of the region sent from the user's device and sends a request to the external information acquisition API based on that name. The server collects tourist information, historical information, and cultural information obtained from the external information acquisition API.

[0643] Input: Name of region

[0644] Output: Tourist information, historical information, cultural information

[0645] Step 3:

[0646] The server passes the collected information to the scenario generation AI and sends a request to generate a scenario for the relevant region. Here, ChatGPT is used as the generation AI model. When generating this scenario, the following prompt is used: "Please generate a scenario introducing tourist attractions, historical figures, and cultural events related to this region."

[0647] Input: Tourist information, historical information, cultural information

[0648] Output: Scenario

[0649] Step 4:

[0650] The server passes the generated scenario to the video generation AI and sends a request to generate a video based on the scenario. When generating the video, the following prompt is used: "Please generate an animated video introducing tourist attractions based on the following scenario."

[0651] Input: Scenario

[0652] Output: Video data

[0653] Step 5:

[0654] The server then requests the AI ​​to generate audio such as narration and background music that corresponds to the generated video. When generating the audio, the following prompt is used: "Please generate narration and background music that are appropriate for the video below."

[0655] Input: Video data

[0656] Output: Audio data

[0657] Step 6:

[0658] The server integrates the scenario, video, and audio to create the completed animation content. This integration process generates a series of multimedia content that users can enjoy visually.

[0659] Input: Scenario, video data, audio data

[0660] Output: Integrated animation content

[0661] Step 7:

[0662] The server stores the generated content in cloud storage and generates a URL or download link for it, which is provided to the user's device so that the user can view the generated content.

[0663] Input: Integrated animation content

[0664] Output: Cloud storage URL or download link

[0665] Through the above processing steps, the system can collect information on local tourist attractions, historical figures, and cultural events, automatically generate and integrate scenarios, images, and audio, and provide them visually to users.

[0666] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0667] The system of the present invention automatically generates animated content related to a region by inputting the name of the region, thereby widely publicizing the region's tourist attractions. The system also includes an emotion engine that recognizes the user's emotions and has the function of personalizing content based on the user's emotions. A specific embodiment of the system and the processing of its program are described below.

[0668] System Configuration

[0669] The system consists of multiple components, including a user device, a server, a scenario generation AI, a video generation AI, a voice generation AI, an emotion engine, and an external information acquisition API. The role of each component is explained below.

[0670] User terminal

[0671] The user terminal provides an interface for inputting the names of regions. This interface generally runs on a web browser and is designed to be easy for users to operate. It also has the function of sending the user's input and reactions to the emotion engine.

[0672] server

[0673] The server is the core of the entire system. It receives the name of the region sent from the user's device and processes requests to the scenario generation AI, video generation AI, and audio generation AI in sequence. It also calls an external information acquisition API to collect tourist information, historical information, and cultural information about the region.

[0674] Emotion Engine

[0675] The emotion engine analyzes the user's input and reactions while watching to recognize the user's emotions. Based on the recognized emotions, the emotion engine adjusts scenario generation, video generation, or audio generation to provide personalized content according to the user's emotions.

[0676] Program processing

[0677] The processing of the program will be explained in natural language below.

[0678] 1. User Input

[0679] The user inputs the name of a region (e.g., "Kochi Prefecture") using the terminal.

[0680] 2. Send to the server

[0681] The terminal sends the entered name of the region to the server as an HTTP request.

[0682] 3. Information gathering

[0683] The server uses an external information acquisition API based on the local name to collect tourist information, historical information, and cultural information.

[0684] 4. Emotion recognition

[0685] The device sends the user's input and reactions to the emotion engine.

[0686] The emotion engine recognizes the user's emotion and notifies the server.

[0687] 5. Scenario Generation

[0688] The server passes the collected information to a scenario generation AI, which then generates a scenario for an anime related to the region.

[0689] The scenario generation AI personalizes the scenario based on the user's emotions.

[0690] 6. Image Generation

[0691] The server passes the generated scenario to the video generation AI, which then generates a video based on the scenario.

[0692] The video generation AI adjusts the video based on the user's emotions.

[0693] 7. Speech Generation

[0694] The server passes the generated scenario and video to an audio generation AI, which generates narration, character voices, and background music.

[0695] The voice generation AI adjusts the voice based on the user's emotions.

[0696] 8. Content Integration

[0697] The server integrates the scenario, video, and audio to create the final animation content.

[0698] 9. Content Provision

[0699] The server stores the generated content in cloud storage and generates and provides a URL or download link for the content.

[0700] Specific examples

[0701] As a specific example of use, consider the case where "Kochi Prefecture" is entered.

[0702] 1. The user enters "Kochi Prefecture" into the terminal.

[0703] 2. The device sends the input information to the server.

[0704] 3. The server collects tourist information (e.g., Katsurahama Beach, Shimanto River), historical information (e.g., Sakamoto Ryoma), and cultural information (e.g., Yosakoi Festival) related to Kochi Prefecture from external APIs.

[0705] 4. The device sends the user's input and reactions to the emotion engine.

[0706] 5. The emotion engine recognizes the user's emotion and notifies the server.

[0707] 6. The server sends a request to the scenario generation AI based on the collected information and generates an anime scenario.

[0708] The scenario generation AI personalizes the scenario based on the user's emotions.

[0709] 7. The server requests the video generation AI to generate a video based on the generated scenario, creating a video that includes scenery and tourist attractions in Kochi Prefecture.

[0710] The video generation AI adjusts the video based on the user's emotions.

[0711] 8. The server then uses the voice generation AI to create narration, character voices, and background music that correspond to the generated scenario and video.

[0712] The voice generation AI adjusts the voice based on the user's emotions.

[0713] 9. The server integrates the scenario, video, and audio, and saves the completed animation content in cloud storage.

[0714] 10. The server generates a destination URL or download link and notifies the user.

[0715] In this way, the system based on the present invention can effectively promote local tourist attractions, increase the number of tourists, and promote regional revitalization. It can also provide more attractive content by recognizing users' emotions and personalizing content based on those emotions.

[0716] The processing flow will be explained below.

[0717] Step 1:

[0718] The user uses a terminal to input the name of a region (e.g., "Kochi Prefecture").

[0719] Step 2:

[0720] The terminal sends the inputted name of the region to the server as an HTTP request.

[0721] Step 3:

[0722] Based on the name of the region received by the server, tourist information, historical information, and cultural information are collected using an external information acquisition API.

[0723] Step 4:

[0724] The server passes the collected information to a scenario generation AI, which then generates a scenario for an anime related to the region.

[0725] Step 5:

[0726] The device transmits the user's input and reactions (emotions) to the emotion engine in real time.

[0727] Step 6:

[0728] The emotion engine recognizes the user's emotions and sends feedback to the scenario generation AI, which then adjusts the content of the scenario.

[0729] Step 7:

[0730] The server receives the adjusted scenario from the scenario generation AI and passes it to the video generation AI.

[0731] Step 8:

[0732] The video generation AI generates video based on the scenario and returns the generated video data to the server.

[0733] Step 9:

[0734] The server passes the generated video data and scenario to the voice generation AI, which then generates the corresponding narration, character voices, and background music.

[0735] Step 10:

[0736] The emotion engine sends feedback to the voice generation AI based on the user's emotions, allowing it to adjust the content and tone of the voice.

[0737] Step 11:

[0738] The voice generation AI returns the adjusted voice data to the server, which then integrates the scenario, video, and voice to create the final animation content.

[0739] Step 12:

[0740] The server saves the completed animation content in cloud storage and generates a distribution URL or download link.

[0741] Step 13:

[0742] The server sends a notification to the user containing a distribution URL or a download link, which the user receives to view or download the animation content.

[0743] As a concrete example, in the case of "Kochi Prefecture," the user inputs "Kochi Prefecture," and tourist information (Katsurahama, Shimanto River), historical information (Sakamoto Ryoma), and cultural information (Yosakoi Festival) are collected via an external API. The emotion engine recognizes emotions based on the user's input and reactions, and feeds this back to the scenario generation AI, video generation AI, and audio generation AI, ultimately creating and providing animation content optimized to the user's emotions.

[0744] Example 2

[0745] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0746] Conventional tourist information presentation methods have difficulty personalizing content based on the user's emotions and interests. Furthermore, they lack a mechanism for automatically collecting related tourist, historical, and cultural information by simply entering the name of a region, and generating visually and auditorily appealing content based on that information.

[0747] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting the name of a region, means for collecting information based on the name of the region, means for generating a scenario based on the collected information, means for generating video based on the generated scenario, means for generating audio corresponding to the video, means for creating content by integrating the scenario, video, and audio, means for providing the created content, means for analyzing the user's input and operation to recognize emotions, and means for adjusting the scenario, video, and audio based on the emotions. This makes it possible to automatically generate and provide attractive tourist information content personalized according to the user's emotions.

[0748] "Means for inputting locality names" refers to an interface through which a user can input a particular locality name into the system.

[0749] "Means for collecting information based on the name of the region" refers to means for obtaining tourist information, historical information, and cultural information related to the input region name from external databases or APIs.

[0750] "Means for generating a scenario based on collected information" refers to the part of the system that automatically creates a story or narration structure based on the information obtained.

[0751] "Means for generating video based on a generated scenario" refers to a function that automatically produces animation or video content according to a created scenario.

[0752] The "means for generating audio corresponding to the video" refers to a function for generating audio data such as narration, character voices, background music, etc., in accordance with the content of the video.

[0753] The "means for creating content by integrating the scenario, video and audio" refers to a process for creating integrated multimedia content by integrating the generated scenario, video and audio.

[0754] "Means for providing the created content" refers to means for providing the created content to users in an accessible form, such as storing it in cloud storage and creating a URL.

[0755] The "means for analyzing the user's inputs and operations to recognize emotions" refers to the part of the system that analyzes the user's interface operations and visual reactions, etc., to determine the user's emotional state.

[0756] "Means for adjusting the scenario, video and audio based on the emotion" refers to a function for adjusting the content and tone of the generated scenario, video and audio based on the data obtained by emotion recognition.

[0757] The system of the present invention automatically generates animation content related to a region by inputting the name of the region, thereby widely publicizing tourist attractions. This system is composed of multiple components, including a user terminal, a server, a scenario generation AI, a video generation AI, a voice generation AI, an emotion engine, and an external information acquisition API. The role of each component and specific implementation methods are explained below.

[0758] User terminal

[0759] The user device provides an interface for inputting the name of a region. This interface runs on a web browser and is designed to be easy for users to operate. It also has a function for sending the user's input and reactions to the emotion engine. Specifically, the user accesses the system's website using a web browser (e.g., Google Chrome) and inputs the name of a region (e.g., "Kochi Prefecture") into the input form.

[0760] server

[0761] The server is the core of the entire system. The server (e.g., AWS EC2) receives the name of the region sent from the user's device and sequentially processes requests to the scenario generation AI, video generation AI, and audio generation AI. It also calls external information acquisition APIs to collect tourist, historical, and cultural information about the region. For example, it acquires information using the Google Maps API or Wikipedia API. The collected information is passed to the scenario generation AI and emotion engine.

[0762] Emotion Engine

[0763] The emotion engine analyzes the user's input and reactions while watching to recognize the user's emotions. Based on the recognized emotions, the emotion engine (e.g., Microsoft Azure Emotion API) adjusts scenario generation, video generation, or audio generation to provide personalized content according to the user's emotions. For example, if the user's reaction is "surprise," it adds a corresponding surprise element to the scenario.

[0764] Scenario Generation

[0765] The server passes the collected information to a scenario generation AI (e.g., OpenAI GPT-4), which then generates an anime scenario based on the collected local information. The scenario generation AI then personalizes the scenario based on the user's emotional data. An example of a prompt could be, "Based on tourist, historical, and cultural information about Kochi Prefecture, please generate a scenario that will make the user feel 'surprised.'"

[0766] Image Generation

[0767] The server passes the generated scenario to a video generation AI (e.g., DALL-E 2), which generates a video based on the scenario. The video generation AI adjusts the video based on the user's emotional data and creates a video that includes scenery and tourist attractions in Kochi Prefecture.

[0768] Voice generation

[0769] The server passes the generated scenario and video to a voice generation AI (e.g., Amazon Polly), which generates narration, character voices, and background music. The voice generation AI adjusts the voice based on the user's emotional data and generates voice data that creates an appropriate atmosphere.

[0770] Content Integration

[0771] The server integrates the script, video, and audio to create the final animation content. This integration process uses open-source video editing software (e.g., FFmpeg), ensuring that each element is properly positioned and timed.

[0772] Content provider

[0773] The server stores the generated animation content in cloud storage (e.g., Amazon S3) and generates an access URL or download link, through which users can access the generated content.

[0774] As a concrete example, if the theme is "Kochi Prefecture," the server uses the Google Maps API and Wikipedia API to collect information on Kochi Prefecture's tourist attractions (Katsurahama Beach, Shimanto River), history (Sakamoto Ryoma), and culture (Yosakoi Festival). The emotion engine then identifies the user's "surprise," and based on that, OpenAI GPT-4 generates a scenario that includes elements of surprise. DALL-E 2 generates video and Amazon Polly generates audio. Finally, FFmpeg is used to integrate the content, and the generated link is provided to the user.

[0775] In this way, the system of the present invention can effectively promote local tourist attractions, increase the number of tourists, and promote regional revitalization. Furthermore, by recognizing users' emotions and personalizing content based on those emotions, the system can provide more engaging and interactive content.

[0776] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0777] Step 1: User Input

[0778] The user opens the system's website using a web browser (e.g., Google Chrome), enters the name of a region (e.g., "Kochi Prefecture") in the input form, and clicks the "Submit" button.

[0779] Input: Name of region (e.g. "Kochi Prefecture")

[0780] Output: HTTP request (including locality name)

[0781] Step 2: Send to the server

[0782] The device sends the name of the region entered by the user to a server (e.g., AWS EC2) as an HTTP request, which also includes the user's identification information.

[0783] Input: HTTP request (region name and user information)

[0784] Output: Send data to the server

[0785] Step 3: Gather information

[0786] Based on the received name of the region, the server uses an external information acquisition API (e.g., Google Maps API, Wikipedia API) to collect tourist information, historical information, and cultural information related to that region.

[0787] The server formats the temporarily stored information and stores it in a database.

[0788] Input: Name of region

[0789] Output: A dataset of collected tourist, historical, and cultural information

[0790] Step 4: Emotion Recognition

[0791] The device monitors user input and actions (e.g., clicks, eye movements) in real time and sends them to an emotion engine (e.g., Microsoft Azure Emotion API).

[0792] The emotion engine analyzes the received data, recognizes the user's emotion, and notifies the server.

[0793] Input: User operation data

[0794] Output: Emotion data (e.g., "surprise")

[0795] Step 5: Scenario generation

[0796] The server passes the collected local information and emotion data to a scenario generation AI (e.g., OpenAI GPT-4), which then generates a scenario based on that data.

[0797] Example prompt: "Generate a scenario that will surprise the user based on tourist, historical, and cultural information about Kochi Prefecture."

[0798] Input: local information, emotion data, prompt sentence

[0799] Output: Generated scenario

[0800] Step 6: Image generation

[0801] The server passes the generated scenario to a video generation AI (e.g., DALL-E 2), which generates a video based on the scenario.

[0802] The video generation AI adjusts the video based on the user's emotional data and returns the appropriate video to the server.

[0803] Input: Generated scenario, emotion data

[0804] Output: Generated video data

[0805] Step 7: Speech generation

[0806] The server passes the generated scenario and video to an audio generation AI (e.g., a voice synthesis API), which generates narration, character voices, and background music.

[0807] The voice generation AI adjusts the tone and pitch of the voice based on the user's emotional data to generate appropriate voice data.

[0808] Input: Generated scenario, video data, emotion data

[0809] Output: Generated audio data

[0810] Step 8: Content Integration

[0811] The server integrates the generated scenario, video, and audio to create the final animation content, using video editing software (e.g., FFmpeg).

[0812] Input: Scenario, video data, audio data

[0813] Output: Integrated animation content

[0814] Step 9: Provide content

[0815] The server uploads the completed animation content to cloud storage (e.g., cloud storage service) and generates an access URL or download link for it. The generated link is notified to the user.

[0816] Input: Integrated animation content

[0817] Output: Access URL or download link

[0818] (Application example 2)

[0819] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0820] Current tourism promotion methods are limited to providing general information, making it difficult to provide personalized experiences that reflect the interests and emotions of individual users. They also have limitations as a means of effectively publicizing local tourist attractions and culture. Therefore, there is a need for a system that can recognize users' emotions and provide personalized content based on them.

[0821] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting the name of a region, means for collecting information based on the name of the region, means for generating a scenario based on the collected information, means for generating video based on the generated scenario, means for generating audio corresponding to the video, means for creating content by integrating the scenario, video, and audio, means for providing the created content, means for recognizing the user's emotions, and means for personalizing the content based on the recognized emotions. This makes it possible to provide personalized tourist content that reflects the user's emotions.

[0822] The "means for inputting the name of a region" refers to a device or program that provides an interface for the user to input the name of a region.

[0823] "Means for collecting information" refers to devices or programs that obtain tourist information, historical information, and cultural information from external APIs and databases.

[0824] A "means for generating a scenario" is a device or program that generates an animation or video story based on collected information.

[0825] "Video generation means" refers to devices or programs that create animations or visual content based on the generated scenario.

[0826] "Audio generating means" refers to a device or program that generates narration or character voices corresponding to the video.

[0827] The "content creating means" refers to a device or program that integrates the generated scenario, video, and audio to create the final multimedia content.

[0828] "Means for providing content" refers to devices or programs that distribute or provide created content to users.

[0829] "Means for recognizing emotions" refers to devices or programs that analyze emotions from the user's facial expressions and voice.

[0830] A "means for personalizing content based on recognized emotions" is a device or program that adjusts the content or expression of content according to the user's emotions.

[0831] The system of the present invention includes a means for inputting the name of a region, a means for collecting information, a means for generating a scenario, a means for generating video, a means for generating audio, a means for integrating and creating content, a means for providing the created content, a means for recognizing emotions, and a means for personalizing content based on the recognized emotions. Each component of the system functions as follows.

[0832] Hardware and software used

[0833] Hardware:

[0834] User devices: smartphones, tablets, etc.

[0835] Server: Cloud server

[0836] software:

[0837] Scenario generation AI: OpenAI GPT-4 as an example

[0838] Image generation AI: DALL-E as an example

[0839] Speech generation AI: Google Cloud Text-to-Speech as an example

[0840] Emotion engine: Microsoft Azure Emotion API as an example

[0841] External information retrieval APIs: Examples include Google Places API and TripAdvisor API

[0842] Cloud storage: AWS S3 as an example

[0843] Program processing

[0844] 1. How to enter the name of a region:

[0845] A user inputs the name of a region through an application interface on a smartphone or tablet. For example, the user inputs "Kyoto Prefecture."

[0846] 2. How we collect information:

[0847] The server uses an external information acquisition API to collect tourist, historical, and cultural information based on the local name, such as information about "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Gion Festival."

[0848] 3. How to recognize emotions:

[0849] The user's facial expressions and voice are recorded using the smartphone's camera and microphone, and analyzed by the emotion engine, allowing the system to recognize emotions in real time while the user waits for the animation to be generated.

[0850] 4. How to generate scenarios:

[0851] The server inputs the collected information into a scenario generation AI to generate an animated scenario. At this time, the scenario content is adjusted based on the user's emotions. For example, if the user is surprised, a scenario that reflects that emotion is generated.

[0852] 5. Means of generating images:

[0853] The server then passes the generated scenario to the image generation AI, which then creates an animation based on it. The image expression can be adjusted according to the user's emotions.

[0854] 6. Means of generating sound:

[0855] The server uses a voice generation AI to create narration and character voices that correspond to the scenario and video. For example, a lively sound is generated for a festival scene.

[0856] 7. Means of integrating and creating content:

[0857] The server integrates the scenario, video and audio to create the final animation content, which is personalized according to the user's emotions.

[0858] 8. Means of providing created content:

[0859] The created animation content is stored in cloud storage and the URL or download link is provided to the user.

[0860] Examples of concrete examples and prompts

[0861] As a usage example, consider the case where a user types "Kyoto Prefecture." The system behaves as follows:

[0862] The user enters "Kyoto Prefecture."

[0863] The server collects tourist information related to Kochi Prefecture (e.g., Kinkaku-ji Temple, Kiyomizu-dera Temple, Gion Festival).

[0864] The user's emotion is analyzed as "surprised."

[0865] Scenario generation AI generates scenarios that reflect surprise and excitement.

[0866] Video generation AI generates videos of festival scenes and other events.

[0867] Voice generation AI creates lively voices.

[0868] These are integrated to create animated content and provide a URL that is saved in cloud storage.

[0869] An example prompt might look like this:

[0870] "The destination is Kyoto Prefecture. Please generate a scenario that includes information about the following tourist attractions: Kinkaku-ji Temple, Kiyomizu-dera Temple, and Gion Festival. Also, since the user will be judged as surprised, please create content that elicits surprise and excitement."

[0871] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0872] Step 1:

[0873] The user inputs the name of a region through the terminal. The user inputs the region name (e.g., "Kyoto Prefecture") in the application interface of the terminal and presses the send button. This input data becomes the information to be sent directly to the server.

[0874] Step 2:

[0875] The server receives the input data and calls an external information acquisition API to collect information. Specifically, it uses the Google Places API and TripAdvisor API to collect tourist information, historical information, and cultural information related to "Kyoto Prefecture." The collected information is returned to the server in JSON format.

[0876] Step 3:

[0877] The server uses an emotion engine to recognize the user's emotions. The user's facial expressions and voice, captured by the camera and microphone on the user's device, are analyzed in real time and sent to the emotion engine (Microsoft Azure Emotion API). The emotion engine returns the analysis results to the server, which stores them in an internal database.

[0878] Step 4:

[0879] The server creates a prompt for the scenario generation AI based on the collected local information and user emotion data. For example, it creates a prompt such as, "The destination is Kyoto Prefecture. Please generate a scenario that includes information about the following tourist attractions: Kinkaku-ji Temple, Kiyomizu-dera Temple, and Gion Festival. Also, since the user is judged to be surprised, please create content that elicits surprise and excitement." and sends it to the scenario generation AI (OpenAI GPT-4). The scenario generation AI generates a scenario and returns it to the server.

[0880] Step 5:

[0881] The server passes the generated scenario to the video generation AI and requests it to generate an animation video. The video generation AI (DALL-E) generates visual content based on the scenario and returns it to the server. At this time, the video expression is adjusted based on the user's emotional data.

[0882] Step 6:

[0883] The server requests the voice generation AI to generate voice based on the scenario and video data. The voice generation AI (Google Cloud Text-to-Speech) generates the corresponding voice, narration, and character voices based on the scenario and returns them to the server.

[0884] Step 7:

[0885] The server integrates the generated scenario, video, and audio data to create the final animation content, and stores the integrated content in its internal storage.

[0886] Step 8:

[0887] The server uploads the final content to cloud storage (AWS S3) and generates a URL or download link for it, which is then sent to the user's device so that the user can view or download the content.

[0888] Through this series of steps, personalized tourism content is generated and provided based on the user's emotions.

[0889] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0890] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0891] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0892] [Third embodiment]

[0893] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0894] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0895] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0896] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0897] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0898] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0899] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0900] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0901] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0902] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0903] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0904] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0905] The system of the present invention is designed to automatically generate animation content related to a region by inputting the name of the region, thereby widely publicizing the tourist spots. Specific embodiments of the system and the processing of the program therefor are described below.

[0906] System Configuration

[0907] The system is composed of multiple components, including a user terminal, a server, a scenario generation AI, a video generation AI, an audio generation AI, and an external information acquisition API. The role of each component is explained below.

[0908] User terminal

[0909] The user terminal provides an interface for inputting the name of a region. This interface generally runs on a web browser and is designed to be easy for users to operate.

[0910] server

[0911] The server is the core of the entire system. It receives the name of the region sent from the user's device and processes requests to the scenario generation AI, video generation AI, and audio generation AI in sequence. It also calls an external information acquisition API to collect tourist information, historical information, and cultural information about the region.

[0912] Program processing

[0913] The processing of the program will be explained in natural language below.

[0914] 1. User Input

[0915] The user inputs the name of a region (e.g., "Kochi Prefecture") using the terminal.

[0916] 2. Send to the server

[0917] The terminal transmits the entered name of the region to the server.

[0918] 3. Information gathering

[0919] The server calls an external information acquisition API based on the region name and collects tourist information, historical information, and cultural information.

[0920] 4. Scenario Generation

[0921] The server passes the collected information to a scenario generation AI, which generates scenarios relevant to the region.

[0922] 5. Image Generation

[0923] The server passes the generated scenario to the video generation AI, which then generates a video based on the scenario.

[0924] 6. Speech Generation

[0925] The server passes the generated scenario and video to an audio generation AI, which generates narration, character voices, and background music that correspond to the scenario.

[0926] 7. Content Integration

[0927] The server integrates the scenario, video, and audio to create the final animation content.

[0928] 8. Content Provision

[0929] The server stores the generated content in cloud storage and generates and provides a URL or download link for the content.

[0930] Specific examples

[0931] As a specific example of use, consider the case where "Kochi Prefecture" is entered.

[0932] 1. The user enters "Kochi Prefecture" into the terminal.

[0933] 2. The device sends the input information to the server.

[0934] 3. The server collects tourist information (e.g., Katsurahama Beach, Shimanto River), historical information (e.g., Sakamoto Ryoma), and cultural information (e.g., Yosakoi Festival) related to Kochi Prefecture from external APIs.

[0935] 4. The server sends a request to the scenario generation AI based on the collected information and generates an anime scenario.

[0936] 5. The server requests the video generation AI to generate a video based on the generated scenario, creating a video that includes scenery and tourist attractions in Kochi Prefecture.

[0937] 6. The server then uses the voice generation AI to create narration, character voices, and background music that correspond to the generated scenario and video.

[0938] 7. The server integrates the scenario, video, and audio, and saves the completed animation content in cloud storage.

[0939] 8. The server generates a URL or download link for the destination and notifies the user.

[0940] In this way, the system based on the present invention effectively promotes local tourist attractions, increasing the number of tourists and revitalizing the region.

[0941] The processing flow will be explained below.

[0942] Step 1:

[0943] The user uses a terminal to input the name of a region (e.g., "Kochi Prefecture").

[0944] Step 2:

[0945] The terminal sends the inputted name of the region to the server as an HTTP request.

[0946] Step 3:

[0947] Based on the name of the region received by the server, tourist information, historical information, and cultural information are collected using an external information acquisition API.

[0948] Step 4:

[0949] The server passes the collected information to a scenario generation AI, which then generates a scenario for an anime related to the region.

[0950] Step 5:

[0951] The scenario generation AI returns the generated scenario to the server, which then passes it on to the video generation AI.

[0952] Step 6:

[0953] The video generation AI generates video based on the scenario and returns the generated video data to the server.

[0954] Step 7:

[0955] The server passes the generated video data and scenario to the voice generation AI, which then generates the corresponding narration, character voices, and background music.

[0956] Step 8:

[0957] The voice generation AI returns the generated voice data to the server, which then integrates the scenario, video, and voice to create the final animation content.

[0958] Step 9:

[0959] The server saves the completed animation content in cloud storage and generates a distribution URL or download link.

[0960] Step 10:

[0961] The server sends a notification to the user containing a distribution URL or a download link, which the user receives to view or download the animation content.

[0962] Example 1

[0963] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0964] Currently, many regions lack the means to effectively promote their tourism resources, making it difficult to attract tourists' attention. Expressing a region's history and culture through video and audio is also problematic, requiring significant costs and time. To address this issue, there is a demand for a system that can easily and automatically generate content that conveys a region's tourist attractions and culture in an appealing way.

[0965] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0966] In this invention, the server includes means for collecting tourist information, historical information, and cultural information based on the name of a region, means for generating a scenario using a generation AI model based on the collected information, means for sending the generated scenario to a video generation AI to generate video, means for sending the generated scenario and video to an audio generation AI to generate audio, means for creating content by integrating the scenario, video, and audio, and means for saving the created content and generating a URL or download link for it, which makes it possible to effectively promote local tourist attractions and culture and increase the number of tourists.

[0967] A "local name" is a name used to identify a particular area, usually indicating a geographical, administrative or historical division of that area.

[0968] "Information gathering means" refers to the methods and devices used to obtain data from external sources and convert it into a form that can be used within the system.

[0969] The term "means for generating a scenario" refers to a method or device that automatically creates text data containing elements of a narrative or storytelling based on collected information.

[0970] "Means for generating video" refers to a method or device for automatically creating visual video content based on text data or a scenario.

[0971] "Audio generating means" refers to a method or device for creating audio data corresponding to a scenario or video, including narration, character voices, background music, etc.

[0972] The "means for creating content" refers to a method or device for integrating the generated scenario, video and audio to create a complete media content.

[0973] "Cloud storage" refers to an online storage service that stores data on remote servers on the Internet and allows it to be accessed and shared as needed.

[0974] "Generative AI model" refers to an artificial intelligence model that has been trained using machine learning to perform a specific task (e.g., text generation, image generation, speech generation).

[0975] A "prompt" is an instruction given to a generative AI model, and refers to text that guides the model in determining what output to generate.

[0976] The system of the present invention automatically generates content (especially animations) related to a region by inputting the name of the region, thereby widely publicizing the tourist spots. The system is composed of multiple components, including a user terminal, a server, a scenario generation AI, a video generation AI, a sound generation AI, and an external information acquisition API.

[0977] User terminal

[0978] The user terminal provides an interface for inputting the name of a region. This interface generally runs on a web browser and is designed to be easy for users to operate. The user inputs the name of the region through this interface.

[0979] server

[0980] The server is the core of the entire system and performs the following processes:

[0981] 1. Receive the name of the region sent from the user terminal.

[0982] 2. Call external information acquisition APIs to collect tourist information, historical information, and cultural information. For example, use the Wikipedia API or Google Maps API.

[0983] 3. Based on the collected information, a scenario generation AI (e.g., OpenAI's GPT-4) is requested to generate a scenario. An example of a prompt sentence for this is as follows:

[0984] "Create a scenario related to tourist attractions, history, and culture based on the local name. For example, for "Kochi Prefecture," generate a scenario that includes "Katsurahama Beach," "Shimanto River," "Sakamoto Ryoma," and "Yosakoi Festival."

[0985] 4. The generated scenario is sent to an image generation AI (e.g., DALL-E, Disco Diffusion), which generates an image based on the scenario. The following prompts are used:

[0986] "Generate an animated video based on a given scenario, including the scenery and tourist attractions of the target region."

[0987] 5. The generated scenario and video are sent to a voice generation AI (e.g., Google Cloud Text-to-Speech, Amazon Polly) to generate narration, character voices, and background music corresponding to the scenario. The following prompts are used:

[0988] "Generate narration, character voices, and background music for the following scenario."

[0989] 6. The script, video, and audio are integrated to create the final animation content. This integration process is carried out using software such as Adobe Premiere Pro and FFmpeg.

[0990] 7. Save the created content to cloud storage (e.g., Amazon S3, Google Drive) and generate a URL or download link to provide to users.

[0991] Specific examples

[0992] For example, if a user inputs the name of a region, "Kochi Prefecture," the device sends this information to the server. The server collects tourist attractions (Katsurahama Beach, Shimanto River), historical information (Sakamoto Ryoma), and cultural information (Yosakoi Festival) related to Kochi Prefecture from an external information acquisition API, and generates the following scenario:

[0993] "One day, the statue of Sakamoto Ryoma, the symbol of Kochi Prefecture, begins to speak, and an adventure begins as he guides you around local tourist attractions."

[0994] Based on this scenario, the video generation AI draws beautiful scenery and tourist spots in Kochi Prefecture, and the voice generation AI generates narration and character voices. Finally, these are integrated to create the completed animation content, which is saved in cloud storage and a link is provided to the user.

[0995] In this way, the system based on this invention can effectively promote tourist attractions by simply entering the name of a region and automatically collecting related information and generating animation content. This system makes it possible to widely publicize regional tourist resources easily and at low cost.

[0996] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0997] Step 1: User enters locality name

[0998] The user uses the terminal to input the name of a region on the system's web browser. For example, they input "Kochi Prefecture." At this time, the input data is recorded on the terminal as a string including the region name.

[0999] Step 2: Send input data to the server

[1000] The terminal sends the region name entered by the user (e.g., "Kochi Prefecture") as an HTTP request to the server. This request includes the region name in JSON format. Input data: "Kochi Prefecture", Output data: JSON data as a request to the server.

[1001] Step 3: Server gathers external information

[1002] Based on the received region name, the server calls an external information acquisition API (e.g. Wikipedia API, Google Maps API) to collect tourist information, historical information, and cultural information. For example, it obtains information such as "Katsurahama Beach," "Shimanto River," "Sakamoto Ryoma," and "Yosakoi Festival." Input data: region name "Kochi Prefecture," output data: JSON data including tourist information, historical information, and cultural information.

[1003] Step 4: Request to Scenario Generation AI

[1004] Based on the collected information, the server sends the following prompt to the scenario generation AI (e.g., GPT-4):

[1005] Create an anime scenario based on Kochi Prefecture's tourist attractions (Katsurahama Beach, Shimanto River), history (Sakamoto Ryoma), and culture (Yosakoi Festival). The theme is "One day, the statue of Sakamoto Ryoma begins to talk, and an adventure begins in which he guides people around local tourist attractions."

[1006] As a result, a scenario is generated. Input data: tourist information, historical information, cultural information. Output data: generated scenario.

[1007] Step 5: Request to the video generation AI

[1008] The server sends the generated scenario to the image generation AI (e.g. DALL-E, Disco Diffusion) and uses the following prompt:

[1009] "Based on a given scenario, generate an animated video that includes scenery and tourist attractions in Kochi Prefecture (Katsurahama Beach, Shimanto River)."

[1010] This generates video data based on the scenario. Input data: scenario, output data: generated video.

[1011] Step 6: Request to speech generation AI

[1012] The server sends the generated scenario and video data to a voice generation AI (e.g., Google Cloud Text-to-Speech, Amazon Polly) and uses the following prompt:

[1013] "Generate narration, character voices, and background music for the following scenario."

[1014] This generates narration, character voices, and background music. Input data: scenario and video, output data: narration, character voices, and background music.

[1015] Step 7: Integrating animated content

[1016] The server integrates the generated scenario, video, and audio using video editing software such as Adobe Premiere Pro or FFmpeg. For example, it uses FFmpeg commands to combine video and audio to create a single animation file. Input data: scenario, video, audio. Output data: integrated animation content.

[1017] Step 8: Storing and serving content to cloud storage

[1018] The server uploads the completed animation content to cloud storage (e.g., Amazon S3, Google Drive) and generates a URL or download link. Finally, it notifies the user of this link. Input data: animation content, Output data: cloud storage URL or download link.

[1019] Through the above processing steps, users can easily create animation content that includes local tourist attractions and culture and widely disseminate it.

[1020] (Application example 1)

[1021] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1022] Visually appealing content is necessary to effectively communicate the appeal of tourist destinations. However, there is a lack of systems that can effectively collect detailed tourist, historical, and cultural information about local areas and automatically generate scenarios, video, and audio based on that information. This makes it difficult to widely communicate the appeal of a region and increase the number of tourists. Furthermore, a method for easily providing the generated content to users is also needed.

[1023] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1024] In this invention, the server includes means for inputting the name of a region, means for collecting information based on the name of the region, means for generating a scenario based on the collected information, means for generating video based on the generated scenario, means for generating audio corresponding to the video, means for creating content by integrating the scenario, video, and audio, and content distribution means including a display device for visually presenting the created content. This makes it possible to effectively collect tourist information, historical information, and cultural information about a region, automatically generate content consisting of a scenario, video, and audio created based on this information, and visually provide it to users.

[1025] The "means for inputting the name of a region" refers to an interface that a user uses to specify a particular region, and includes, for example, keyboard input and voice input.

[1026] "Means of collecting information" refers to the function of obtaining data such as tourist information, historical information, and cultural information about the designated region from external APIs and databases.

[1027] "Scenario generation means" refers to algorithms or models that create narratives or descriptions based on collected information, and in this case includes generative AI.

[1028] "Means for generating images" refers to tools or models that create visual images based on the generated scenario, including image generation AI.

[1029] "Means for generating audio" refers to tools and models for generating narration, sound effects, and background music corresponding to video, including audio generation AI.

[1030] "Means for creating content" refers to the process or system that integrates the generated scenario, video, and audio into a series of content.

[1031] "Content distribution means" refers to a system for providing created content in a form that users can enjoy visually, and includes smartphone apps, web platforms, head-mounted displays, etc.

[1032] "Tourist information" refers to information about major tourist destinations, attractions, and tourist facilities in the designated region.

[1033] "Historical information" refers to information about historical events, people, and ruins that exist in the designated region.

[1034] "Cultural information" refers to information about festivals, traditional events, cultural customs, etc. that are characteristic of the designated region.

[1035] A "display device" is a hardware device for visually presenting generated content to a user, including smartphones, tablets, visual augmentation devices, etc.

[1036] The system for implementing this invention is composed of multiple components including a user terminal, a server, a scenario generation AI, a video generation AI, a sound generation AI, and an external information acquisition API.

[1037] System Configuration

[1038] User terminal

[1039] The user terminal provides an interface for inputting the name of a region. This interface runs on a web browser or is provided as a dedicated mobile app. The system starts working when the user inputs the name of a region into the terminal.

[1040] server

[1041] The server is the core of the entire system, receiving the locality name sent from the user terminal and proceeding with the processing.

[1042] The server does the following:

[1043] 1. Information gathering

[1044] The server calls an external information retrieval API based on the region name to collect tourist information, historical information, and cultural information. This API call allows for the collection of a wide range of information about the region.

[1045] 2. Scenario Generation

[1046] The server calls a scenario generation AI based on the collected information to generate a scenario related to the relevant region. For example, ChatGPT is used as the generation AI model for this scenario generation. An example of a specific prompt is, "Please generate a scenario introducing tourist attractions, historical figures, and cultural events related to this region."

[1047] 3. Image Generation

[1048] The server calls the video generation AI based on the generated scenario, and generates a video based on the scenario. An example of a prompt in this case would be, "Please generate an animated video introducing tourist attractions based on the following scenario."

[1049] 4. Speech Generation

[1050] The server then uses the AI ​​to generate audio for the generated video, such as narration and background music. An example of a prompt for voice generation is, "Please generate narration and background music that are appropriate for the video below."

[1051] 5. Content Integration and Delivery

[1052] The server integrates the scenario, video, and audio to create the final animation content. This integrated content is stored in cloud storage and a URL or download link is generated. The link is provided to the user's device, allowing the user to visually enjoy the generated content.

[1053] This system makes it possible to effectively collect tourist, historical, and cultural information about a specific region, automatically generate it as attractive multimedia content, and provide it visually, which is expected to greatly promote tourism and revitalize the region.

[1054] Specific examples

[1055] For example, if a user inputs "a certain region" as "a region they would like to visit," the system will operate as follows:

[1056] 1. User Input

[1057] The user enters "a certain region" into the terminal.

[1058] 2. Information gathering

[1059] The server collects tourist information, historical information, and cultural information about a "certain region" from an external API.

[1060] 3. Scenario Generation

[1061] The server sends the collected information to a scenario generation AI, which generates a scenario related to a "certain region."

[1062] 4. Image Generation

[1063] The server sends the scenario to the video generation AI, which generates a video including tourist attractions in a certain region.

[1064] 5. Speech Generation

[1065] The server uses a voice generation AI to generate narration, character voices, and background music that correspond to the generated video.

[1066] 6. Content Integration and Delivery

[1067] The server integrates the scenario, video, and audio, and saves the completed animation content in cloud storage. Users can view this content via the provided URL.

[1068] Through this process, users can enjoy tourist attractions and culture in a visually rich way, and the appeal of local areas can be effectively promoted.

[1069] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1070] Step 1:

[1071] The user inputs the name of a region into the terminal, either via keyboard or voice input, which is received by the user interface and sent to the server for further processing.

[1072] Input: Name of region

[1073] Output: Locality name data sent to the server

[1074] Step 2:

[1075] The server receives the name of the region sent from the user's device and sends a request to the external information acquisition API based on that name. The server collects tourist information, historical information, and cultural information obtained from the external information acquisition API.

[1076] Input: Name of region

[1077] Output: Tourist information, historical information, cultural information

[1078] Step 3:

[1079] The server passes the collected information to the scenario generation AI and sends a request to generate a scenario for the relevant region. Here, ChatGPT is used as the generation AI model. When generating this scenario, the following prompt is used: "Please generate a scenario introducing tourist attractions, historical figures, and cultural events related to this region."

[1080] Input: Tourist information, historical information, cultural information

[1081] Output: Scenario

[1082] Step 4:

[1083] The server passes the generated scenario to the video generation AI and sends a request to generate a video based on the scenario. When generating the video, the following prompt is used: "Please generate an animated video introducing tourist attractions based on the following scenario."

[1084] Input: Scenario

[1085] Output: Video data

[1086] Step 5:

[1087] The server then requests the AI ​​to generate audio such as narration and background music that corresponds to the generated video. When generating the audio, the following prompt is used: "Please generate narration and background music that are appropriate for the video below."

[1088] Input: Video data

[1089] Output: Audio data

[1090] Step 6:

[1091] The server integrates the scenario, video, and audio to create the completed animation content. This integration process generates a series of multimedia content that users can enjoy visually.

[1092] Input: Scenario, video data, audio data

[1093] Output: Integrated animation content

[1094] Step 7:

[1095] The server stores the generated content in cloud storage and generates a URL or download link for it, which is provided to the user's device so that the user can view the generated content.

[1096] Input: Integrated animation content

[1097] Output: Cloud storage URL or download link

[1098] Through the above processing steps, the system can collect information on local tourist attractions, historical figures, and cultural events, automatically generate and integrate scenarios, images, and audio, and provide them visually to users.

[1099] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1100] The system of the present invention automatically generates animated content related to a region by inputting the name of the region, thereby widely publicizing the region's tourist attractions. The system also includes an emotion engine that recognizes the user's emotions and has the function of personalizing content based on the user's emotions. A specific embodiment of the system and the processing of its program are described below.

[1101] System Configuration

[1102] The system consists of multiple components, including a user device, a server, a scenario generation AI, a video generation AI, a voice generation AI, an emotion engine, and an external information acquisition API. The role of each component is explained below.

[1103] User terminal

[1104] The user terminal provides an interface for inputting the names of regions. This interface generally runs on a web browser and is designed to be easy for users to operate. It also has the function of sending the user's input and reactions to the emotion engine.

[1105] server

[1106] The server is the core of the entire system. It receives the name of the region sent from the user's device and processes requests to the scenario generation AI, video generation AI, and audio generation AI in sequence. It also calls an external information acquisition API to collect tourist information, historical information, and cultural information about the region.

[1107] Emotion Engine

[1108] The emotion engine analyzes the user's input and reactions while watching to recognize the user's emotions. Based on the recognized emotions, the emotion engine adjusts scenario generation, video generation, or audio generation to provide personalized content according to the user's emotions.

[1109] Program processing

[1110] The processing of the program will be explained in natural language below.

[1111] 1. User Input

[1112] The user inputs the name of a region (e.g., "Kochi Prefecture") using the terminal.

[1113] 2. Send to the server

[1114] The terminal sends the entered name of the region to the server as an HTTP request.

[1115] 3. Information gathering

[1116] The server uses an external information acquisition API based on the local name to collect tourist information, historical information, and cultural information.

[1117] 4. Emotion recognition

[1118] The device sends the user's input and reactions to the emotion engine.

[1119] The emotion engine recognizes the user's emotion and notifies the server.

[1120] 5. Scenario Generation

[1121] The server passes the collected information to a scenario generation AI, which then generates a scenario for an anime related to the region.

[1122] The scenario generation AI personalizes the scenario based on the user's emotions.

[1123] 6. Image Generation

[1124] The server passes the generated scenario to the video generation AI, which then generates a video based on the scenario.

[1125] The video generation AI adjusts the video based on the user's emotions.

[1126] 7. Speech Generation

[1127] The server passes the generated scenario and video to an audio generation AI, which generates narration, character voices, and background music.

[1128] The voice generation AI adjusts the voice based on the user's emotions.

[1129] 8. Content Integration

[1130] The server integrates the scenario, video, and audio to create the final animation content.

[1131] 9. Content Provision

[1132] The server stores the generated content in cloud storage and generates and provides a URL or download link for the content.

[1133] Specific examples

[1134] As a specific example of use, consider the case where "Kochi Prefecture" is entered.

[1135] 1. The user enters "Kochi Prefecture" into the terminal.

[1136] 2. The device sends the input information to the server.

[1137] 3. The server collects tourist information (e.g., Katsurahama Beach, Shimanto River), historical information (e.g., Sakamoto Ryoma), and cultural information (e.g., Yosakoi Festival) related to Kochi Prefecture from external APIs.

[1138] 4. The device sends the user's input and reactions to the emotion engine.

[1139] 5. The emotion engine recognizes the user's emotion and notifies the server.

[1140] 6. The server sends a request to the scenario generation AI based on the collected information and generates an anime scenario.

[1141] The scenario generation AI personalizes the scenario based on the user's emotions.

[1142] 7. The server requests the video generation AI to generate a video based on the generated scenario, creating a video that includes scenery and tourist attractions in Kochi Prefecture.

[1143] The video generation AI adjusts the video based on the user's emotions.

[1144] 8. The server then uses the voice generation AI to create narration, character voices, and background music that correspond to the generated scenario and video.

[1145] The voice generation AI adjusts the voice based on the user's emotions.

[1146] 9. The server integrates the scenario, video, and audio, and saves the completed animation content in cloud storage.

[1147] 10. The server generates a destination URL or download link and notifies the user.

[1148] In this way, the system based on the present invention can effectively promote local tourist attractions, increase the number of tourists, and promote regional revitalization. It can also provide more attractive content by recognizing users' emotions and personalizing content based on those emotions.

[1149] The processing flow will be explained below.

[1150] Step 1:

[1151] The user uses a terminal to input the name of a region (e.g., "Kochi Prefecture").

[1152] Step 2:

[1153] The terminal sends the inputted name of the region to the server as an HTTP request.

[1154] Step 3:

[1155] Based on the name of the region received by the server, tourist information, historical information, and cultural information are collected using an external information acquisition API.

[1156] Step 4:

[1157] The server passes the collected information to a scenario generation AI, which then generates a scenario for an anime related to the region.

[1158] Step 5:

[1159] The device transmits the user's input and reactions (emotions) to the emotion engine in real time.

[1160] Step 6:

[1161] The emotion engine recognizes the user's emotions and sends feedback to the scenario generation AI, which then adjusts the content of the scenario.

[1162] Step 7:

[1163] The server receives the adjusted scenario from the scenario generation AI and passes it to the video generation AI.

[1164] Step 8:

[1165] The video generation AI generates video based on the scenario and returns the generated video data to the server.

[1166] Step 9:

[1167] The server passes the generated video data and scenario to the voice generation AI, which then generates the corresponding narration, character voices, and background music.

[1168] Step 10:

[1169] The emotion engine sends feedback to the voice generation AI based on the user's emotions, allowing it to adjust the content and tone of the voice.

[1170] Step 11:

[1171] The voice generation AI returns the adjusted voice data to the server, which then integrates the scenario, video, and voice to create the final animation content.

[1172] Step 12:

[1173] The server saves the completed animation content in cloud storage and generates a distribution URL or download link.

[1174] Step 13:

[1175] The server sends a notification to the user containing a distribution URL or a download link, which the user receives to view or download the animation content.

[1176] As a concrete example, in the case of "Kochi Prefecture," the user inputs "Kochi Prefecture," and tourist information (Katsurahama, Shimanto River), historical information (Sakamoto Ryoma), and cultural information (Yosakoi Festival) are collected via an external API. The emotion engine recognizes emotions based on the user's input and reactions, and feeds this back to the scenario generation AI, video generation AI, and audio generation AI, ultimately creating and providing animation content optimized to the user's emotions.

[1177] Example 2

[1178] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1179] Conventional tourist information presentation methods have difficulty personalizing content based on the user's emotions and interests. Furthermore, they lack a mechanism for automatically collecting related tourist, historical, and cultural information by simply entering the name of a region, and generating visually and auditorily appealing content based on that information.

[1180] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting the name of a region, means for collecting information based on the name of the region, means for generating a scenario based on the collected information, means for generating video based on the generated scenario, means for generating audio corresponding to the video, means for creating content by integrating the scenario, video, and audio, means for providing the created content, means for analyzing the user's input and operation to recognize emotions, and means for adjusting the scenario, video, and audio based on the emotions. This makes it possible to automatically generate and provide attractive tourist information content personalized according to the user's emotions.

[1181] "Means for inputting locality names" refers to an interface through which a user can input a particular locality name into the system.

[1182] "Means for collecting information based on the name of the region" refers to means for obtaining tourist information, historical information, and cultural information related to the input region name from external databases or APIs.

[1183] "Means for generating a scenario based on collected information" refers to the part of the system that automatically creates a story or narration structure based on the information obtained.

[1184] "Means for generating video based on a generated scenario" refers to a function that automatically produces animation or video content according to a created scenario.

[1185] The "means for generating audio corresponding to the video" refers to a function for generating audio data such as narration, character voices, background music, etc., in accordance with the content of the video.

[1186] The "means for creating content by integrating the scenario, video and audio" refers to a process for creating integrated multimedia content by integrating the generated scenario, video and audio.

[1187] "Means for providing the created content" refers to means for providing the created content to users in an accessible form, such as storing it in cloud storage and creating a URL.

[1188] The "means for analyzing the user's inputs and operations to recognize emotions" refers to the part of the system that analyzes the user's interface operations and visual reactions, etc., to determine the user's emotional state.

[1189] "Means for adjusting the scenario, video and audio based on the emotion" refers to a function for adjusting the content and tone of the generated scenario, video and audio based on the data obtained by emotion recognition.

[1190] The system of the present invention automatically generates animation content related to a region by inputting the name of the region, thereby widely publicizing tourist attractions. This system is composed of multiple components, including a user terminal, a server, a scenario generation AI, a video generation AI, a voice generation AI, an emotion engine, and an external information acquisition API. The role of each component and specific implementation methods are explained below.

[1191] User terminal

[1192] The user device provides an interface for inputting the name of a region. This interface runs on a web browser and is designed to be easy for users to operate. It also has a function for sending the user's input and reactions to the emotion engine. Specifically, the user accesses the system's website using a web browser (e.g., Google Chrome) and inputs the name of a region (e.g., "Kochi Prefecture") into the input form.

[1193] server

[1194] The server is the core of the entire system. The server (e.g., AWS EC2) receives the name of the region sent from the user's device and sequentially processes requests to the scenario generation AI, video generation AI, and audio generation AI. It also calls external information acquisition APIs to collect tourist, historical, and cultural information about the region. For example, it acquires information using the Google Maps API or Wikipedia API. The collected information is passed to the scenario generation AI and emotion engine.

[1195] Emotion Engine

[1196] The emotion engine analyzes the user's input and reactions while watching to recognize the user's emotions. Based on the recognized emotions, the emotion engine (e.g., Microsoft Azure Emotion API) adjusts scenario generation, video generation, or audio generation to provide personalized content according to the user's emotions. For example, if the user's reaction is "surprise," it adds a corresponding surprise element to the scenario.

[1197] Scenario Generation

[1198] The server passes the collected information to a scenario generation AI (e.g., OpenAI GPT-4), which then generates an anime scenario based on the collected local information. The scenario generation AI then personalizes the scenario based on the user's emotional data. An example of a prompt could be, "Based on tourist, historical, and cultural information about Kochi Prefecture, please generate a scenario that will make the user feel 'surprised.'"

[1199] Image Generation

[1200] The server passes the generated scenario to a video generation AI (e.g., DALL-E 2), which generates a video based on the scenario. The video generation AI adjusts the video based on the user's emotional data and creates a video that includes scenery and tourist attractions in Kochi Prefecture.

[1201] Voice generation

[1202] The server passes the generated scenario and video to a voice generation AI (e.g., Amazon Polly), which generates narration, character voices, and background music. The voice generation AI adjusts the voice based on the user's emotional data and generates voice data that creates an appropriate atmosphere.

[1203] Content Integration

[1204] The server integrates the script, video, and audio to create the final animation content. This integration process uses open-source video editing software (e.g., FFmpeg), ensuring that each element is properly positioned and timed.

[1205] Content provider

[1206] The server stores the generated animation content in cloud storage (e.g., Amazon S3) and generates an access URL or download link, through which users can access the generated content.

[1207] As a concrete example, if the theme is "Kochi Prefecture," the server uses the Google Maps API and Wikipedia API to collect information on Kochi Prefecture's tourist attractions (Katsurahama Beach, Shimanto River), history (Sakamoto Ryoma), and culture (Yosakoi Festival). The emotion engine then identifies the user's "surprise," and based on that, OpenAI GPT-4 generates a scenario that includes elements of surprise. DALL-E 2 generates video and Amazon Polly generates audio. Finally, FFmpeg is used to integrate the content, and the generated link is provided to the user.

[1208] In this way, the system of the present invention can effectively promote local tourist attractions, increase the number of tourists, and promote regional revitalization. Furthermore, by recognizing users' emotions and personalizing content based on those emotions, the system can provide more engaging and interactive content.

[1209] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1210] Step 1: User Input

[1211] The user opens the system's website using a web browser (e.g., Google Chrome), enters the name of a region (e.g., "Kochi Prefecture") in the input form, and clicks the "Submit" button.

[1212] Input: Name of region (e.g. "Kochi Prefecture")

[1213] Output: HTTP request (including locality name)

[1214] Step 2: Send to the server

[1215] The device sends the name of the region entered by the user to a server (e.g., AWS EC2) as an HTTP request, which also includes the user's identification information.

[1216] Input: HTTP request (region name and user information)

[1217] Output: Send data to the server

[1218] Step 3: Gather information

[1219] Based on the received name of the region, the server uses an external information acquisition API (e.g., Google Maps API, Wikipedia API) to collect tourist information, historical information, and cultural information related to that region.

[1220] The server formats the temporarily stored information and stores it in a database.

[1221] Input: Name of region

[1222] Output: A dataset of collected tourist, historical, and cultural information

[1223] Step 4: Emotion Recognition

[1224] The device monitors user input and actions (e.g., clicks, eye movements) in real time and sends them to an emotion engine (e.g., Microsoft Azure Emotion API).

[1225] The emotion engine analyzes the received data, recognizes the user's emotion, and notifies the server.

[1226] Input: User operation data

[1227] Output: Emotion data (e.g., "surprise")

[1228] Step 5: Scenario generation

[1229] The server passes the collected local information and emotion data to a scenario generation AI (e.g., OpenAI GPT-4), which then generates a scenario based on that data.

[1230] Example prompt: "Generate a scenario that will surprise the user based on tourist, historical, and cultural information about Kochi Prefecture."

[1231] Input: local information, emotion data, prompt sentence

[1232] Output: Generated scenario

[1233] Step 6: Image generation

[1234] The server passes the generated scenario to a video generation AI (e.g., DALL-E 2), which generates a video based on the scenario.

[1235] The video generation AI adjusts the video based on the user's emotional data and returns the appropriate video to the server.

[1236] Input: Generated scenario, emotion data

[1237] Output: Generated video data

[1238] Step 7: Speech generation

[1239] The server passes the generated scenario and video to an audio generation AI (e.g., a voice synthesis API), which generates narration, character voices, and background music.

[1240] The voice generation AI adjusts the tone and pitch of the voice based on the user's emotional data to generate appropriate voice data.

[1241] Input: Generated scenario, video data, emotion data

[1242] Output: Generated audio data

[1243] Step 8: Content Integration

[1244] The server integrates the generated scenario, video, and audio to create the final animation content, using video editing software (e.g., FFmpeg).

[1245] Input: Scenario, video data, audio data

[1246] Output: Integrated animation content

[1247] Step 9: Provide content

[1248] The server uploads the completed animation content to cloud storage (e.g., cloud storage service) and generates an access URL or download link for it. The generated link is notified to the user.

[1249] Input: Integrated animation content

[1250] Output: Access URL or download link

[1251] (Application example 2)

[1252] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1253] Current tourism promotion methods are limited to providing general information, making it difficult to provide personalized experiences that reflect the interests and emotions of individual users. They also have limitations as a means of effectively publicizing local tourist attractions and culture. Therefore, there is a need for a system that can recognize users' emotions and provide personalized content based on them.

[1254] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting the name of a region, means for collecting information based on the name of the region, means for generating a scenario based on the collected information, means for generating video based on the generated scenario, means for generating audio corresponding to the video, means for creating content by integrating the scenario, video, and audio, means for providing the created content, means for recognizing the user's emotions, and means for personalizing the content based on the recognized emotions. This makes it possible to provide personalized tourist content that reflects the user's emotions.

[1255] The "means for inputting the name of a region" refers to a device or program that provides an interface for the user to input the name of a region.

[1256] "Means for collecting information" refers to devices or programs that obtain tourist information, historical information, and cultural information from external APIs and databases.

[1257] A "means for generating a scenario" is a device or program that generates an animation or video story based on collected information.

[1258] "Video generation means" refers to devices or programs that create animations or visual content based on the generated scenario.

[1259] "Audio generating means" refers to a device or program that generates narration or character voices corresponding to the video.

[1260] The "content creating means" refers to a device or program that integrates the generated scenario, video, and audio to create the final multimedia content.

[1261] "Means for providing content" refers to devices or programs that distribute or provide created content to users.

[1262] "Means for recognizing emotions" refers to devices or programs that analyze emotions from the user's facial expressions and voice.

[1263] A "means for personalizing content based on recognized emotions" is a device or program that adjusts the content or expression of content according to the user's emotions.

[1264] The system of the present invention includes a means for inputting the name of a region, a means for collecting information, a means for generating a scenario, a means for generating video, a means for generating audio, a means for integrating and creating content, a means for providing the created content, a means for recognizing emotions, and a means for personalizing content based on the recognized emotions. Each component of the system functions as follows.

[1265] Hardware and software used

[1266] Hardware:

[1267] User devices: smartphones, tablets, etc.

[1268] Server: Cloud server

[1269] software:

[1270] Scenario generation AI: OpenAI GPT-4 as an example

[1271] Image generation AI: DALL-E as an example

[1272] Speech generation AI: Google Cloud Text-to-Speech as an example

[1273] Emotion engine: Microsoft Azure Emotion API as an example

[1274] External information retrieval APIs: Examples include Google Places API and TripAdvisor API

[1275] Cloud storage: AWS S3 as an example

[1276] Program processing

[1277] 1. How to enter the name of a region:

[1278] A user inputs the name of a region through an application interface on a smartphone or tablet. For example, the user inputs "Kyoto Prefecture."

[1279] 2. How we collect information:

[1280] The server uses an external information acquisition API to collect tourist, historical, and cultural information based on the local name, such as information about "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Gion Festival."

[1281] 3. How to recognize emotions:

[1282] The user's facial expressions and voice are recorded using the smartphone's camera and microphone, and analyzed by the emotion engine, allowing the system to recognize emotions in real time while the user waits for the animation to be generated.

[1283] 4. How to generate scenarios:

[1284] The server inputs the collected information into a scenario generation AI to generate an animated scenario. At this time, the scenario content is adjusted based on the user's emotions. For example, if the user is surprised, a scenario that reflects that emotion is generated.

[1285] 5. Means of generating images:

[1286] The server then passes the generated scenario to the image generation AI, which then creates an animation based on it. The image expression can be adjusted according to the user's emotions.

[1287] 6. Means of generating sound:

[1288] The server uses a voice generation AI to create narration and character voices that correspond to the scenario and video. For example, a lively sound is generated for a festival scene.

[1289] 7. Means of integrating and creating content:

[1290] The server integrates the scenario, video and audio to create the final animation content, which is personalized according to the user's emotions.

[1291] 8. Means of providing created content:

[1292] The created animation content is stored in cloud storage and the URL or download link is provided to the user.

[1293] Examples of concrete examples and prompts

[1294] As a usage example, consider the case where a user types "Kyoto Prefecture." The system behaves as follows:

[1295] The user enters "Kyoto Prefecture."

[1296] The server collects tourist information related to Kochi Prefecture (e.g., Kinkaku-ji Temple, Kiyomizu-dera Temple, Gion Festival).

[1297] The user's emotion is analyzed as "surprised."

[1298] Scenario generation AI generates scenarios that reflect surprise and excitement.

[1299] Video generation AI generates videos of festival scenes and other events.

[1300] Voice generation AI creates lively voices.

[1301] These are integrated to create animated content and provide a URL that is saved in cloud storage.

[1302] An example prompt might look like this:

[1303] "The destination is Kyoto Prefecture. Please generate a scenario that includes information about the following tourist attractions: Kinkaku-ji Temple, Kiyomizu-dera Temple, and Gion Festival. Also, since the user will be judged as surprised, please create content that elicits surprise and excitement."

[1304] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1305] Step 1:

[1306] The user inputs the name of a region through the terminal. The user inputs the region name (e.g., "Kyoto Prefecture") in the application interface of the terminal and presses the send button. This input data becomes the information to be sent directly to the server.

[1307] Step 2:

[1308] The server receives the input data and calls an external information acquisition API to collect information. Specifically, it uses the Google Places API and TripAdvisor API to collect tourist information, historical information, and cultural information related to "Kyoto Prefecture." The collected information is returned to the server in JSON format.

[1309] Step 3:

[1310] The server uses an emotion engine to recognize the user's emotions. The user's facial expressions and voice, captured by the camera and microphone on the user's device, are analyzed in real time and sent to the emotion engine (Microsoft Azure Emotion API). The emotion engine returns the analysis results to the server, which stores them in an internal database.

[1311] Step 4:

[1312] The server creates a prompt for the scenario generation AI based on the collected local information and user emotion data. For example, it creates a prompt such as, "The destination is Kyoto Prefecture. Please generate a scenario that includes information about the following tourist attractions: Kinkaku-ji Temple, Kiyomizu-dera Temple, and Gion Festival. Also, since the user is judged to be surprised, please create content that elicits surprise and excitement." and sends it to the scenario generation AI (OpenAI GPT-4). The scenario generation AI generates a scenario and returns it to the server.

[1313] Step 5:

[1314] The server passes the generated scenario to the video generation AI and requests it to generate an animation video. The video generation AI (DALL-E) generates visual content based on the scenario and returns it to the server. At this time, the video expression is adjusted based on the user's emotional data.

[1315] Step 6:

[1316] The server requests the voice generation AI to generate voice based on the scenario and video data. The voice generation AI (Google Cloud Text-to-Speech) generates the corresponding voice, narration, and character voices based on the scenario and returns them to the server.

[1317] Step 7:

[1318] The server integrates the generated scenario, video, and audio data to create the final animation content, and stores the integrated content in its internal storage.

[1319] Step 8:

[1320] The server uploads the final content to cloud storage (AWS S3) and generates a URL or download link for it, which is then sent to the user's device so that the user can view or download the content.

[1321] Through this series of steps, personalized tourism content is generated and provided based on the user's emotions.

[1322] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1323] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1324] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1325] [Fourth embodiment]

[1326] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1327] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1328] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1329] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1330] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1331] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1332] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1333] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1334] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1335] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1336] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1337] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1338] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1339] The system of the present invention is designed to automatically generate animation content related to a region by inputting the name of the region, thereby widely publicizing the tourist spots. Specific embodiments of the system and the processing of the program therefor are described below.

[1340] System Configuration

[1341] The system is composed of multiple components, including a user terminal, a server, a scenario generation AI, a video generation AI, an audio generation AI, and an external information acquisition API. The role of each component is explained below.

[1342] User terminal

[1343] The user terminal provides an interface for inputting the name of a region. This interface generally runs on a web browser and is designed to be easy for users to operate.

[1344] server

[1345] The server is the core of the entire system. It receives the name of the region sent from the user's device and processes requests to the scenario generation AI, video generation AI, and audio generation AI in sequence. It also calls an external information acquisition API to collect tourist information, historical information, and cultural information about the region.

[1346] Program processing

[1347] The processing of the program will be explained in natural language below.

[1348] 1. User Input

[1349] The user inputs the name of a region (e.g., "Kochi Prefecture") using the terminal.

[1350] 2. Send to the server

[1351] The terminal transmits the entered name of the region to the server.

[1352] 3. Information gathering

[1353] The server calls an external information acquisition API based on the region name and collects tourist information, historical information, and cultural information.

[1354] 4. Scenario Generation

[1355] The server passes the collected information to a scenario generation AI, which generates scenarios relevant to the region.

[1356] 5. Image Generation

[1357] The server passes the generated scenario to the video generation AI, which then generates a video based on the scenario.

[1358] 6. Speech Generation

[1359] The server passes the generated scenario and video to an audio generation AI, which generates narration, character voices, and background music that correspond to the scenario.

[1360] 7. Content Integration

[1361] The server integrates the scenario, video, and audio to create the final animation content.

[1362] 8. Content Provision

[1363] The server stores the generated content in cloud storage and generates and provides a URL or download link for the content.

[1364] Specific examples

[1365] As a specific example of use, consider the case where "Kochi Prefecture" is entered.

[1366] 1. The user enters "Kochi Prefecture" into the terminal.

[1367] 2. The device sends the input information to the server.

[1368] 3. The server collects tourist information (e.g., Katsurahama Beach, Shimanto River), historical information (e.g., Sakamoto Ryoma), and cultural information (e.g., Yosakoi Festival) related to Kochi Prefecture from external APIs.

[1369] 4. The server sends a request to the scenario generation AI based on the collected information and generates an anime scenario.

[1370] 5. The server requests the video generation AI to generate a video based on the generated scenario, creating a video that includes scenery and tourist attractions in Kochi Prefecture.

[1371] 6. The server then uses the voice generation AI to create narration, character voices, and background music that correspond to the generated scenario and video.

[1372] 7. The server integrates the scenario, video, and audio, and saves the completed animation content in cloud storage.

[1373] 8. The server generates a URL or download link for the destination and notifies the user.

[1374] In this way, the system based on the present invention effectively promotes local tourist attractions, increasing the number of tourists and revitalizing the region.

[1375] The processing flow will be explained below.

[1376] Step 1:

[1377] The user uses a terminal to input the name of a region (e.g., "Kochi Prefecture").

[1378] Step 2:

[1379] The terminal sends the inputted name of the region to the server as an HTTP request.

[1380] Step 3:

[1381] Based on the name of the region received by the server, tourist information, historical information, and cultural information are collected using an external information acquisition API.

[1382] Step 4:

[1383] The server passes the collected information to a scenario generation AI, which then generates a scenario for an anime related to the region.

[1384] Step 5:

[1385] The scenario generation AI returns the generated scenario to the server, which then passes it on to the video generation AI.

[1386] Step 6:

[1387] The video generation AI generates video based on the scenario and returns the generated video data to the server.

[1388] Step 7:

[1389] The server passes the generated video data and scenario to the voice generation AI, which then generates the corresponding narration, character voices, and background music.

[1390] Step 8:

[1391] The voice generation AI returns the generated voice data to the server, which then integrates the scenario, video, and voice to create the final animation content.

[1392] Step 9:

[1393] The server saves the completed animation content in cloud storage and generates a distribution URL or download link.

[1394] Step 10:

[1395] The server sends a notification to the user containing a distribution URL or a download link, which the user receives to view or download the animation content.

[1396] Example 1

[1397] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1398] Currently, many regions lack the means to effectively promote their tourism resources, making it difficult to attract tourists' attention. Expressing a region's history and culture through video and audio is also problematic, requiring significant costs and time. To address this issue, there is a demand for a system that can easily and automatically generate content that conveys a region's tourist attractions and culture in an appealing way.

[1399] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1400] In this invention, the server includes means for collecting tourist information, historical information, and cultural information based on the name of a region, means for generating a scenario using a generation AI model based on the collected information, means for sending the generated scenario to a video generation AI to generate video, means for sending the generated scenario and video to an audio generation AI to generate audio, means for creating content by integrating the scenario, video, and audio, and means for saving the created content and generating a URL or download link for it, which makes it possible to effectively promote local tourist attractions and culture and increase the number of tourists.

[1401] A "local name" is a name used to identify a particular area, usually indicating a geographical, administrative or historical division of that area.

[1402] "Information gathering means" refers to the methods and devices used to obtain data from external sources and convert it into a form that can be used within the system.

[1403] The term "means for generating a scenario" refers to a method or device that automatically creates text data containing elements of a narrative or storytelling based on collected information.

[1404] "Means for generating video" refers to a method or device for automatically creating visual video content based on text data or a scenario.

[1405] "Audio generating means" refers to a method or device for creating audio data corresponding to a scenario or video, including narration, character voices, background music, etc.

[1406] The "means for creating content" refers to a method or device for integrating the generated scenario, video and audio to create a complete media content.

[1407] "Cloud storage" refers to an online storage service that stores data on remote servers on the Internet and allows it to be accessed and shared as needed.

[1408] "Generative AI model" refers to an artificial intelligence model that has been trained using machine learning to perform a specific task (e.g., text generation, image generation, speech generation).

[1409] A "prompt" is an instruction given to a generative AI model, and refers to text that guides the model in determining what output to generate.

[1410] The system of the present invention automatically generates content (especially animations) related to a region by inputting the name of the region, thereby widely publicizing the tourist spots. The system is composed of multiple components, including a user terminal, a server, a scenario generation AI, a video generation AI, a sound generation AI, and an external information acquisition API.

[1411] User terminal

[1412] The user terminal provides an interface for inputting the name of a region. This interface generally runs on a web browser and is designed to be easy for users to operate. The user inputs the name of the region through this interface.

[1413] server

[1414] The server is the core of the entire system and performs the following processes:

[1415] 1. Receive the name of the region sent from the user terminal.

[1416] 2. Call external information acquisition APIs to collect tourist information, historical information, and cultural information. For example, use the Wikipedia API or Google Maps API.

[1417] 3. Based on the collected information, a scenario generation AI (e.g., OpenAI's GPT-4) is requested to generate a scenario. An example of a prompt sentence for this is as follows:

[1418] "Create a scenario related to tourist attractions, history, and culture based on the local name. For example, for "Kochi Prefecture," generate a scenario that includes "Katsurahama Beach," "Shimanto River," "Sakamoto Ryoma," and "Yosakoi Festival."

[1419] 4. The generated scenario is sent to an image generation AI (e.g., DALL-E, Disco Diffusion), which generates an image based on the scenario. The following prompts are used:

[1420] "Generate an animated video based on a given scenario, including the scenery and tourist attractions of the target region."

[1421] 5. The generated scenario and video are sent to a voice generation AI (e.g., Google Cloud Text-to-Speech, Amazon Polly) to generate narration, character voices, and background music corresponding to the scenario. The following prompts are used:

[1422] "Generate narration, character voices, and background music for the following scenario."

[1423] 6. The script, video, and audio are integrated to create the final animation content. This integration process is carried out using software such as Adobe Premiere Pro and FFmpeg.

[1424] 7. Save the created content to cloud storage (e.g., Amazon S3, Google Drive) and generate a URL or download link to provide to users.

[1425] Specific examples

[1426] For example, if a user inputs the name of a region, "Kochi Prefecture," the device sends this information to the server. The server collects tourist attractions (Katsurahama Beach, Shimanto River), historical information (Sakamoto Ryoma), and cultural information (Yosakoi Festival) related to Kochi Prefecture from an external information acquisition API, and generates the following scenario:

[1427] "One day, the statue of Sakamoto Ryoma, the symbol of Kochi Prefecture, begins to speak, and an adventure begins as he guides you around local tourist attractions."

[1428] Based on this scenario, the video generation AI draws beautiful scenery and tourist spots in Kochi Prefecture, and the voice generation AI generates narration and character voices. Finally, these are integrated to create the completed animation content, which is saved in cloud storage and a link is provided to the user.

[1429] In this way, the system based on this invention can effectively promote tourist attractions by simply entering the name of a region and automatically collecting related information and generating animation content. This system makes it possible to widely publicize regional tourist resources easily and at low cost.

[1430] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1431] Step 1: User enters locality name

[1432] The user uses the terminal to input the name of a region on the system's web browser. For example, they input "Kochi Prefecture." At this time, the input data is recorded on the terminal as a string including the region name.

[1433] Step 2: Send input data to the server

[1434] The terminal sends the region name entered by the user (e.g., "Kochi Prefecture") as an HTTP request to the server. This request includes the region name in JSON format. Input data: "Kochi Prefecture", Output data: JSON data as a request to the server.

[1435] Step 3: Server gathers external information

[1436] Based on the received region name, the server calls an external information acquisition API (e.g. Wikipedia API, Google Maps API) to collect tourist information, historical information, and cultural information. For example, it obtains information such as "Katsurahama Beach," "Shimanto River," "Sakamoto Ryoma," and "Yosakoi Festival." Input data: region name "Kochi Prefecture," output data: JSON data including tourist information, historical information, and cultural information.

[1437] Step 4: Request to Scenario Generation AI

[1438] Based on the collected information, the server sends the following prompt to the scenario generation AI (e.g., GPT-4):

[1439] Create an anime scenario based on Kochi Prefecture's tourist attractions (Katsurahama Beach, Shimanto River), history (Sakamoto Ryoma), and culture (Yosakoi Festival). The theme is "One day, the statue of Sakamoto Ryoma begins to talk, and an adventure begins in which he guides people around local tourist attractions."

[1440] As a result, a scenario is generated. Input data: tourist information, historical information, cultural information. Output data: generated scenario.

[1441] Step 5: Request to the video generation AI

[1442] The server sends the generated scenario to the image generation AI (e.g. DALL-E, Disco Diffusion) and uses the following prompt:

[1443] "Based on a given scenario, generate an animated video that includes scenery and tourist attractions in Kochi Prefecture (Katsurahama Beach, Shimanto River)."

[1444] This generates video data based on the scenario. Input data: scenario, output data: generated video.

[1445] Step 6: Request to speech generation AI

[1446] The server sends the generated scenario and video data to a voice generation AI (e.g., Google Cloud Text-to-Speech, Amazon Polly) and uses the following prompt:

[1447] "Generate narration, character voices, and background music for the following scenario."

[1448] This generates narration, character voices, and background music. Input data: scenario and video, output data: narration, character voices, and background music.

[1449] Step 7: Integrating animated content

[1450] The server integrates the generated scenario, video, and audio using video editing software such as Adobe Premiere Pro or FFmpeg. For example, it uses FFmpeg commands to combine video and audio to create a single animation file. Input data: scenario, video, audio. Output data: integrated animation content.

[1451] Step 8: Storing and serving content to cloud storage

[1452] The server uploads the completed animation content to cloud storage (e.g., Amazon S3, Google Drive) and generates a URL or download link. Finally, it notifies the user of this link. Input data: animation content, Output data: cloud storage URL or download link.

[1453] Through the above processing steps, users can easily create animation content that includes local tourist attractions and culture and widely disseminate it.

[1454] (Application example 1)

[1455] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1456] Visually appealing content is necessary to effectively communicate the appeal of tourist destinations. However, there is a lack of systems that can effectively collect detailed tourist, historical, and cultural information about local areas and automatically generate scenarios, video, and audio based on that information. This makes it difficult to widely communicate the appeal of a region and increase the number of tourists. Furthermore, a method for easily providing the generated content to users is also needed.

[1457] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1458] In this invention, the server includes means for inputting the name of a region, means for collecting information based on the name of the region, means for generating a scenario based on the collected information, means for generating video based on the generated scenario, means for generating audio corresponding to the video, means for creating content by integrating the scenario, video, and audio, and content distribution means including a display device for visually presenting the created content. This makes it possible to effectively collect tourist information, historical information, and cultural information about a region, automatically generate content consisting of a scenario, video, and audio created based on this information, and visually provide it to users.

[1459] The "means for inputting the name of a region" refers to an interface that a user uses to specify a particular region, and includes, for example, keyboard input and voice input.

[1460] "Means of collecting information" refers to the function of obtaining data such as tourist information, historical information, and cultural information about the designated region from external APIs and databases.

[1461] "Scenario generation means" refers to algorithms or models that create narratives or descriptions based on collected information, and in this case includes generative AI.

[1462] "Means for generating images" refers to tools or models that create visual images based on the generated scenario, including image generation AI.

[1463] "Means for generating audio" refers to tools and models for generating narration, sound effects, and background music corresponding to video, including audio generation AI.

[1464] "Means for creating content" refers to the process or system that integrates the generated scenario, video, and audio into a series of content.

[1465] "Content distribution means" refers to a system for providing created content in a form that users can enjoy visually, and includes smartphone apps, web platforms, head-mounted displays, etc.

[1466] "Tourist information" refers to information about major tourist destinations, attractions, and tourist facilities in the designated region.

[1467] "Historical information" refers to information about historical events, people, and ruins that exist in the designated region.

[1468] "Cultural information" refers to information about festivals, traditional events, cultural customs, etc. that are characteristic of the designated region.

[1469] A "display device" is a hardware device for visually presenting generated content to a user, including smartphones, tablets, visual augmentation devices, etc.

[1470] The system for implementing this invention is composed of multiple components including a user terminal, a server, a scenario generation AI, a video generation AI, a sound generation AI, and an external information acquisition API.

[1471] System Configuration

[1472] User terminal

[1473] The user terminal provides an interface for inputting the name of a region. This interface runs on a web browser or is provided as a dedicated mobile app. The system starts working when the user inputs the name of a region into the terminal.

[1474] server

[1475] The server is the core of the entire system, receiving the locality name sent from the user terminal and proceeding with the processing.

[1476] The server does the following:

[1477] 1. Information gathering

[1478] The server calls an external information retrieval API based on the region name to collect tourist information, historical information, and cultural information. This API call allows for the collection of a wide range of information about the region.

[1479] 2. Scenario Generation

[1480] The server calls a scenario generation AI based on the collected information to generate a scenario related to the relevant region. For example, ChatGPT is used as the generation AI model for this scenario generation. An example of a specific prompt is, "Please generate a scenario introducing tourist attractions, historical figures, and cultural events related to this region."

[1481] 3. Image Generation

[1482] The server calls the video generation AI based on the generated scenario, and generates a video based on the scenario. An example of a prompt in this case would be, "Please generate an animated video introducing tourist attractions based on the following scenario."

[1483] 4. Speech Generation

[1484] The server then uses the AI ​​to generate audio for the generated video, such as narration and background music. An example of a prompt for voice generation is, "Please generate narration and background music that are appropriate for the video below."

[1485] 5. Content Integration and Delivery

[1486] The server integrates the scenario, video, and audio to create the final animation content. This integrated content is stored in cloud storage and a URL or download link is generated. The link is provided to the user's device, allowing the user to visually enjoy the generated content.

[1487] This system makes it possible to effectively collect tourist, historical, and cultural information about a specific region, automatically generate it as attractive multimedia content, and provide it visually, which is expected to greatly promote tourism and revitalize the region.

[1488] Specific examples

[1489] For example, if a user inputs "a certain region" as "a region they would like to visit," the system will operate as follows:

[1490] 1. User Input

[1491] The user enters "a certain region" into the terminal.

[1492] 2. Information gathering

[1493] The server collects tourist information, historical information, and cultural information about a "certain region" from an external API.

[1494] 3. Scenario Generation

[1495] The server sends the collected information to a scenario generation AI, which generates a scenario related to a "certain region."

[1496] 4. Image Generation

[1497] The server sends the scenario to the video generation AI, which generates a video including tourist attractions in a certain region.

[1498] 5. Speech Generation

[1499] The server uses a voice generation AI to generate narration, character voices, and background music that correspond to the generated video.

[1500] 6. Content Integration and Delivery

[1501] The server integrates the scenario, video, and audio, and saves the completed animation content in cloud storage. Users can view this content via the provided URL.

[1502] Through this process, users can enjoy tourist attractions and culture in a visually rich way, and the appeal of local areas can be effectively promoted.

[1503] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1504] Step 1:

[1505] The user inputs the name of a region into the terminal, either via keyboard or voice input, which is received by the user interface and sent to the server for further processing.

[1506] Input: Name of region

[1507] Output: Locality name data sent to the server

[1508] Step 2:

[1509] The server receives the name of the region sent from the user's device and sends a request to the external information acquisition API based on that name. The server collects tourist information, historical information, and cultural information obtained from the external information acquisition API.

[1510] Input: Name of region

[1511] Output: Tourist information, historical information, cultural information

[1512] Step 3:

[1513] The server passes the collected information to the scenario generation AI and sends a request to generate a scenario for the relevant region. Here, ChatGPT is used as the generation AI model. When generating this scenario, the following prompt is used: "Please generate a scenario introducing tourist attractions, historical figures, and cultural events related to this region."

[1514] Input: Tourist information, historical information, cultural information

[1515] Output: Scenario

[1516] Step 4:

[1517] The server passes the generated scenario to the video generation AI and sends a request to generate a video based on the scenario. When generating the video, the following prompt is used: "Please generate an animated video introducing tourist attractions based on the following scenario."

[1518] Input: Scenario

[1519] Output: Video data

[1520] Step 5:

[1521] The server then requests the AI ​​to generate audio such as narration and background music that corresponds to the generated video. When generating the audio, the following prompt is used: "Please generate narration and background music that are appropriate for the video below."

[1522] Input: Video data

[1523] Output: Audio data

[1524] Step 6:

[1525] The server integrates the scenario, video, and audio to create the completed animation content. This integration process generates a series of multimedia content that users can enjoy visually.

[1526] Input: Scenario, video data, audio data

[1527] Output: Integrated animation content

[1528] Step 7:

[1529] The server stores the generated content in cloud storage and generates a URL or download link for it, which is provided to the user's device so that the user can view the generated content.

[1530] Input: Integrated animation content

[1531] Output: Cloud storage URL or download link

[1532] Through the above processing steps, the system can collect information on local tourist attractions, historical figures, and cultural events, automatically generate and integrate scenarios, images, and audio, and provide them visually to users.

[1533] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1534] The system of the present invention automatically generates animated content related to a region by inputting the name of the region, thereby widely publicizing the region's tourist attractions. The system also includes an emotion engine that recognizes the user's emotions and has the function of personalizing content based on the user's emotions. A specific embodiment of the system and the processing of its program are described below.

[1535] System Configuration

[1536] The system consists of multiple components, including a user device, a server, a scenario generation AI, a video generation AI, a voice generation AI, an emotion engine, and an external information acquisition API. The role of each component is explained below.

[1537] User terminal

[1538] The user terminal provides an interface for inputting the names of regions. This interface generally runs on a web browser and is designed to be easy for users to operate. It also has the function of sending the user's input and reactions to the emotion engine.

[1539] server

[1540] The server is the core of the entire system. It receives the name of the region sent from the user's device and processes requests to the scenario generation AI, video generation AI, and audio generation AI in sequence. It also calls an external information acquisition API to collect tourist information, historical information, and cultural information about the region.

[1541] Emotion Engine

[1542] The emotion engine analyzes the user's input and reactions while watching to recognize the user's emotions. Based on the recognized emotions, the emotion engine adjusts scenario generation, video generation, or audio generation to provide personalized content according to the user's emotions.

[1543] Program processing

[1544] The processing of the program will be explained in natural language below.

[1545] 1. User Input

[1546] The user inputs the name of a region (e.g., "Kochi Prefecture") using the terminal.

[1547] 2. Send to the server

[1548] The terminal sends the entered name of the region to the server as an HTTP request.

[1549] 3. Information gathering

[1550] The server uses an external information acquisition API based on the local name to collect tourist information, historical information, and cultural information.

[1551] 4. Emotion recognition

[1552] The device sends the user's input and reactions to the emotion engine.

[1553] The emotion engine recognizes the user's emotion and notifies the server.

[1554] 5. Scenario Generation

[1555] The server passes the collected information to a scenario generation AI, which then generates a scenario for an anime related to the region.

[1556] The scenario generation AI personalizes the scenario based on the user's emotions.

[1557] 6. Image Generation

[1558] The server passes the generated scenario to the video generation AI, which then generates a video based on the scenario.

[1559] The video generation AI adjusts the video based on the user's emotions.

[1560] 7. Speech Generation

[1561] The server passes the generated scenario and video to an audio generation AI, which generates narration, character voices, and background music.

[1562] The voice generation AI adjusts the voice based on the user's emotions.

[1563] 8. Content Integration

[1564] The server integrates the scenario, video, and audio to create the final animation content.

[1565] 9. Content Provision

[1566] The server stores the generated content in cloud storage and generates and provides a URL or download link for the content.

[1567] Specific examples

[1568] As a specific example of use, consider the case where "Kochi Prefecture" is entered.

[1569] 1. The user enters "Kochi Prefecture" into the terminal.

[1570] 2. The device sends the input information to the server.

[1571] 3. The server collects tourist information (e.g., Katsurahama Beach, Shimanto River), historical information (e.g., Sakamoto Ryoma), and cultural information (e.g., Yosakoi Festival) related to Kochi Prefecture from external APIs.

[1572] 4. The device sends the user's input and reactions to the emotion engine.

[1573] 5. The emotion engine recognizes the user's emotion and notifies the server.

[1574] 6. The server sends a request to the scenario generation AI based on the collected information and generates an anime scenario.

[1575] The scenario generation AI personalizes the scenario based on the user's emotions.

[1576] 7. The server requests the video generation AI to generate a video based on the generated scenario, creating a video that includes scenery and tourist attractions in Kochi Prefecture.

[1577] The video generation AI adjusts the video based on the user's emotions.

[1578] 8. The server then uses the voice generation AI to create narration, character voices, and background music that correspond to the generated scenario and video.

[1579] The voice generation AI adjusts the voice based on the user's emotions.

[1580] 9. The server integrates the scenario, video, and audio, and saves the completed animation content in cloud storage.

[1581] 10. The server generates a destination URL or download link and notifies the user.

[1582] In this way, the system based on the present invention can effectively promote local tourist attractions, increase the number of tourists, and promote regional revitalization. It can also provide more attractive content by recognizing users' emotions and personalizing content based on those emotions.

[1583] The processing flow will be explained below.

[1584] Step 1:

[1585] The user uses a terminal to input the name of a region (e.g., "Kochi Prefecture").

[1586] Step 2:

[1587] The terminal sends the inputted name of the region to the server as an HTTP request.

[1588] Step 3:

[1589] Based on the name of the region received by the server, tourist information, historical information, and cultural information are collected using an external information acquisition API.

[1590] Step 4:

[1591] The server passes the collected information to a scenario generation AI, which then generates a scenario for an anime related to the region.

[1592] Step 5:

[1593] The device transmits the user's input and reactions (emotions) to the emotion engine in real time.

[1594] Step 6:

[1595] The emotion engine recognizes the user's emotions and sends feedback to the scenario generation AI, which then adjusts the content of the scenario.

[1596] Step 7:

[1597] The server receives the adjusted scenario from the scenario generation AI and passes it to the video generation AI.

[1598] Step 8:

[1599] The video generation AI generates video based on the scenario and returns the generated video data to the server.

[1600] Step 9:

[1601] The server passes the generated video data and scenario to the voice generation AI, which then generates the corresponding narration, character voices, and background music.

[1602] Step 10:

[1603] The emotion engine sends feedback to the voice generation AI based on the user's emotions, allowing it to adjust the content and tone of the voice.

[1604] Step 11:

[1605] The voice generation AI returns the adjusted voice data to the server, which then integrates the scenario, video, and voice to create the final animation content.

[1606] Step 12:

[1607] The server saves the completed animation content in cloud storage and generates a distribution URL or download link.

[1608] Step 13:

[1609] The server sends a notification to the user containing a distribution URL or a download link, which the user receives to view or download the animation content.

[1610] As a concrete example, in the case of "Kochi Prefecture," the user inputs "Kochi Prefecture," and tourist information (Katsurahama, Shimanto River), historical information (Sakamoto Ryoma), and cultural information (Yosakoi Festival) are collected via an external API. The emotion engine recognizes emotions based on the user's input and reactions, and feeds this back to the scenario generation AI, video generation AI, and audio generation AI, ultimately creating and providing animation content optimized to the user's emotions.

[1611] Example 2

[1612] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1613] Conventional tourist information presentation methods have difficulty personalizing content based on the user's emotions and interests. Furthermore, they lack a mechanism for automatically collecting related tourist, historical, and cultural information by simply entering the name of a region, and generating visually and auditorily appealing content based on that information.

[1614] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting the name of a region, means for collecting information based on the name of the region, means for generating a scenario based on the collected information, means for generating video based on the generated scenario, means for generating audio corresponding to the video, means for creating content by integrating the scenario, video, and audio, means for providing the created content, means for analyzing the user's input and operation to recognize emotions, and means for adjusting the scenario, video, and audio based on the emotions. This makes it possible to automatically generate and provide attractive tourist information content personalized according to the user's emotions.

[1615] "Means for inputting locality names" refers to an interface through which a user can input a particular locality name into the system.

[1616] "Means for collecting information based on the name of the region" refers to means for obtaining tourist information, historical information, and cultural information related to the input region name from external databases or APIs.

[1617] "Means for generating a scenario based on collected information" refers to the part of the system that automatically creates a story or narration structure based on the information obtained.

[1618] "Means for generating video based on a generated scenario" refers to a function that automatically produces animation or video content according to a created scenario.

[1619] The "means for generating audio corresponding to the video" refers to a function for generating audio data such as narration, character voices, background music, etc., in accordance with the content of the video.

[1620] The "means for creating content by integrating the scenario, video and audio" refers to a process for creating integrated multimedia content by integrating the generated scenario, video and audio.

[1621] "Means for providing the created content" refers to means for providing the created content to users in an accessible form, such as storing it in cloud storage and creating a URL.

[1622] The "means for analyzing the user's inputs and operations to recognize emotions" refers to the part of the system that analyzes the user's interface operations and visual reactions, etc., to determine the user's emotional state.

[1623] "Means for adjusting the scenario, video and audio based on the emotion" refers to a function for adjusting the content and tone of the generated scenario, video and audio based on the data obtained by emotion recognition.

[1624] The system of the present invention automatically generates animation content related to a region by inputting the name of the region, thereby widely publicizing tourist attractions. This system is composed of multiple components, including a user terminal, a server, a scenario generation AI, a video generation AI, a voice generation AI, an emotion engine, and an external information acquisition API. The role of each component and specific implementation methods are explained below.

[1625] User terminal

[1626] The user device provides an interface for inputting the name of a region. This interface runs on a web browser and is designed to be easy for users to operate. It also has a function for sending the user's input and reactions to the emotion engine. Specifically, the user accesses the system's website using a web browser (e.g., Google Chrome) and inputs the name of a region (e.g., "Kochi Prefecture") into the input form.

[1627] server

[1628] The server is the core of the entire system. The server (e.g., AWS EC2) receives the name of the region sent from the user's device and sequentially processes requests to the scenario generation AI, video generation AI, and audio generation AI. It also calls external information acquisition APIs to collect tourist, historical, and cultural information about the region. For example, it acquires information using the Google Maps API or Wikipedia API. The collected information is passed to the scenario generation AI and emotion engine.

[1629] Emotion Engine

[1630] The emotion engine analyzes the user's input and reactions while watching to recognize the user's emotions. Based on the recognized emotions, the emotion engine (e.g., Microsoft Azure Emotion API) adjusts scenario generation, video generation, or audio generation to provide personalized content according to the user's emotions. For example, if the user's reaction is "surprise," it adds a corresponding surprise element to the scenario.

[1631] Scenario Generation

[1632] The server passes the collected information to a scenario generation AI (e.g., OpenAI GPT-4), which then generates an anime scenario based on the collected local information. The scenario generation AI then personalizes the scenario based on the user's emotional data. An example of a prompt could be, "Based on tourist, historical, and cultural information about Kochi Prefecture, please generate a scenario that will make the user feel 'surprised.'"

[1633] Image Generation

[1634] The server passes the generated scenario to a video generation AI (e.g., DALL-E 2), which generates a video based on the scenario. The video generation AI adjusts the video based on the user's emotional data and creates a video that includes scenery and tourist attractions in Kochi Prefecture.

[1635] Voice generation

[1636] The server passes the generated scenario and video to a voice generation AI (e.g., Amazon Polly), which generates narration, character voices, and background music. The voice generation AI adjusts the voice based on the user's emotional data and generates voice data that creates an appropriate atmosphere.

[1637] Content Integration

[1638] The server integrates the script, video, and audio to create the final animation content. This integration process uses open-source video editing software (e.g., FFmpeg), ensuring that each element is properly positioned and timed.

[1639] Content provider

[1640] The server stores the generated animation content in cloud storage (e.g., Amazon S3) and generates an access URL or download link, through which users can access the generated content.

[1641] As a concrete example, if the theme is "Kochi Prefecture," the server uses the Google Maps API and Wikipedia API to collect information on Kochi Prefecture's tourist attractions (Katsurahama Beach, Shimanto River), history (Sakamoto Ryoma), and culture (Yosakoi Festival). The emotion engine then identifies the user's "surprise," and based on that, OpenAI GPT-4 generates a scenario that includes elements of surprise. DALL-E 2 generates video and Amazon Polly generates audio. Finally, FFmpeg is used to integrate the content, and the generated link is provided to the user.

[1642] In this way, the system of the present invention can effectively promote local tourist attractions, increase the number of tourists, and promote regional revitalization. Furthermore, by recognizing users' emotions and personalizing content based on those emotions, the system can provide more engaging and interactive content.

[1643] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1644] Step 1: User Input

[1645] The user opens the system's website using a web browser (e.g., Google Chrome), enters the name of a region (e.g., "Kochi Prefecture") in the input form, and clicks the "Submit" button.

[1646] Input: Name of region (e.g. "Kochi Prefecture")

[1647] Output: HTTP request (including locality name)

[1648] Step 2: Send to the server

[1649] The device sends the name of the region entered by the user to a server (e.g., AWS EC2) as an HTTP request, which also includes the user's identification information.

[1650] Input: HTTP request (region name and user information)

[1651] Output: Send data to the server

[1652] Step 3: Gather information

[1653] Based on the received name of the region, the server uses an external information acquisition API (e.g., Google Maps API, Wikipedia API) to collect tourist information, historical information, and cultural information related to that region.

[1654] The server formats the temporarily stored information and stores it in a database.

[1655] Input: Name of region

[1656] Output: A dataset of collected tourist, historical, and cultural information

[1657] Step 4: Emotion Recognition

[1658] The device monitors user input and actions (e.g., clicks, eye movements) in real time and sends them to an emotion engine (e.g., Microsoft Azure Emotion API).

[1659] The emotion engine analyzes the received data, recognizes the user's emotion, and notifies the server.

[1660] Input: User operation data

[1661] Output: Emotion data (e.g., "surprise")

[1662] Step 5: Scenario generation

[1663] The server passes the collected local information and emotion data to a scenario generation AI (e.g., OpenAI GPT-4), which then generates a scenario based on that data.

[1664] Example prompt: "Generate a scenario that will surprise the user based on tourist, historical, and cultural information about Kochi Prefecture."

[1665] Input: local information, emotion data, prompt sentence

[1666] Output: Generated scenario

[1667] Step 6: Image generation

[1668] The server passes the generated scenario to a video generation AI (e.g., DALL-E 2), which generates a video based on the scenario.

[1669] The video generation AI adjusts the video based on the user's emotional data and returns the appropriate video to the server.

[1670] Input: Generated scenario, emotion data

[1671] Output: Generated video data

[1672] Step 7: Speech generation

[1673] The server passes the generated scenario and video to an audio generation AI (e.g., a voice synthesis API), which generates narration, character voices, and background music.

[1674] The voice generation AI adjusts the tone and pitch of the voice based on the user's emotional data to generate appropriate voice data.

[1675] Input: Generated scenario, video data, emotion data

[1676] Output: Generated audio data

[1677] Step 8: Content Integration

[1678] The server integrates the generated scenario, video, and audio to create the final animation content, using video editing software (e.g., FFmpeg).

[1679] Input: Scenario, video data, audio data

[1680] Output: Integrated animation content

[1681] Step 9: Provide content

[1682] The server uploads the completed animation content to cloud storage (e.g., cloud storage service) and generates an access URL or download link for it. The generated link is notified to the user.

[1683] Input: Integrated animation content

[1684] Output: Access URL or download link

[1685] (Application example 2)

[1686] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1687] Current tourism promotion methods are limited to providing general information, making it difficult to provide personalized experiences that reflect the interests and emotions of individual users. They also have limitations as a means of effectively publicizing local tourist attractions and culture. Therefore, there is a need for a system that can recognize users' emotions and provide personalized content based on them.

[1688] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting the name of a region, means for collecting information based on the name of the region, means for generating a scenario based on the collected information, means for generating video based on the generated scenario, means for generating audio corresponding to the video, means for creating content by integrating the scenario, video, and audio, means for providing the created content, means for recognizing the user's emotions, and means for personalizing the content based on the recognized emotions. This makes it possible to provide personalized tourist content that reflects the user's emotions.

[1689] The "means for inputting the name of a region" refers to a device or program that provides an interface for the user to input the name of a region.

[1690] "Means for collecting information" refers to devices or programs that obtain tourist information, historical information, and cultural information from external APIs and databases.

[1691] A "means for generating a scenario" is a device or program that generates an animation or video story based on collected information.

[1692] "Video generation means" refers to devices or programs that create animations or visual content based on the generated scenario.

[1693] "Audio generating means" refers to a device or program that generates narration or character voices corresponding to the video.

[1694] The "content creating means" refers to a device or program that integrates the generated scenario, video, and audio to create the final multimedia content.

[1695] "Means for providing content" refers to devices or programs that distribute or provide created content to users.

[1696] "Means for recognizing emotions" refers to devices or programs that analyze emotions from the user's facial expressions and voice.

[1697] A "means for personalizing content based on recognized emotions" is a device or program that adjusts the content or expression of content according to the user's emotions.

[1698] The system of the present invention includes a means for inputting the name of a region, a means for collecting information, a means for generating a scenario, a means for generating video, a means for generating audio, a means for integrating and creating content, a means for providing the created content, a means for recognizing emotions, and a means for personalizing content based on the recognized emotions. Each component of the system functions as follows.

[1699] Hardware and software used

[1700] Hardware:

[1701] User devices: smartphones, tablets, etc.

[1702] Server: Cloud server

[1703] software:

[1704] Scenario generation AI: OpenAI GPT-4 as an example

[1705] Image generation AI: DALL-E as an example

[1706] Speech generation AI: Google Cloud Text-to-Speech as an example

[1707] Emotion engine: Microsoft Azure Emotion API as an example

[1708] External information retrieval APIs: Examples include Google Places API and TripAdvisor API

[1709] Cloud storage: AWS S3 as an example

[1710] Program processing

[1711] 1. How to enter the name of a region:

[1712] A user inputs the name of a region through an application interface on a smartphone or tablet. For example, the user inputs "Kyoto Prefecture."

[1713] 2. How we collect information:

[1714] The server uses an external information acquisition API to collect tourist, historical, and cultural information based on the local name, such as information about "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Gion Festival."

[1715] 3. How to recognize emotions:

[1716] The user's facial expressions and voice are recorded using the smartphone's camera and microphone, and analyzed by the emotion engine, allowing the system to recognize emotions in real time while the user waits for the animation to be generated.

[1717] 4. How to generate scenarios:

[1718] The server inputs the collected information into a scenario generation AI to generate an animated scenario. At this time, the scenario content is adjusted based on the user's emotions. For example, if the user is surprised, a scenario that reflects that emotion is generated.

[1719] 5. Means of generating images:

[1720] The server then passes the generated scenario to the image generation AI, which then creates an animation based on it. The image expression can be adjusted according to the user's emotions.

[1721] 6. Means of generating sound:

[1722] The server uses a voice generation AI to create narration and character voices that correspond to the scenario and video. For example, a lively sound is generated for a festival scene.

[1723] 7. Means of integrating and creating content:

[1724] The server integrates the scenario, video and audio to create the final animation content, which is personalized according to the user's emotions.

[1725] 8. Means of providing created content:

[1726] The created animation content is stored in cloud storage and the URL or download link is provided to the user.

[1727] Examples of concrete examples and prompts

[1728] As a usage example, consider the case where a user types "Kyoto Prefecture." The system behaves as follows:

[1729] The user enters "Kyoto Prefecture."

[1730] The server collects tourist information related to Kochi Prefecture (e.g., Kinkaku-ji Temple, Kiyomizu-dera Temple, Gion Festival).

[1731] The user's emotion is analyzed as "surprised."

[1732] Scenario generation AI generates scenarios that reflect surprise and excitement.

[1733] Video generation AI generates videos of festival scenes and other events.

[1734] Voice generation AI creates lively voices.

[1735] These are integrated to create animated content and provide a URL that is saved in cloud storage.

[1736] An example prompt might look like this:

[1737] "The destination is Kyoto Prefecture. Please generate a scenario that includes information about the following tourist attractions: Kinkaku-ji Temple, Kiyomizu-dera Temple, and Gion Festival. Also, since the user will be judged as surprised, please create content that elicits surprise and excitement."

[1738] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1739] Step 1:

[1740] The user inputs the name of a region through the terminal. The user inputs the region name (e.g., "Kyoto Prefecture") in the application interface of the terminal and presses the send button. This input data becomes the information to be sent directly to the server.

[1741] Step 2:

[1742] The server receives the input data and calls an external information acquisition API to collect information. Specifically, it uses the Google Places API and TripAdvisor API to collect tourist information, historical information, and cultural information related to "Kyoto Prefecture." The collected information is returned to the server in JSON format.

[1743] Step 3:

[1744] The server uses an emotion engine to recognize the user's emotions. The user's facial expressions and voice, captured by the camera and microphone on the user's device, are analyzed in real time and sent to the emotion engine (Microsoft Azure Emotion API). The emotion engine returns the analysis results to the server, which stores them in an internal database.

[1745] Step 4:

[1746] The server creates a prompt for the scenario generation AI based on the collected local information and user emotion data. For example, it creates a prompt such as, "The destination is Kyoto Prefecture. Please generate a scenario that includes information about the following tourist attractions: Kinkaku-ji Temple, Kiyomizu-dera Temple, and Gion Festival. Also, since the user is judged to be surprised, please create content that elicits surprise and excitement." and sends it to the scenario generation AI (OpenAI GPT-4). The scenario generation AI generates a scenario and returns it to the server.

[1747] Step 5:

[1748] The server passes the generated scenario to the video generation AI and requests it to generate an animation video. The video generation AI (DALL-E) generates visual content based on the scenario and returns it to the server. At this time, the video expression is adjusted based on the user's emotional data.

[1749] Step 6:

[1750] The server requests the voice generation AI to generate voice based on the scenario and video data. The voice generation AI (Google Cloud Text-to-Speech) generates the corresponding voice, narration, and character voices based on the scenario and returns them to the server.

[1751] Step 7:

[1752] The server integrates the generated scenario, video, and audio data to create the final animation content, and stores the integrated content in its internal storage.

[1753] Step 8:

[1754] The server uploads the final content to cloud storage (AWS S3) and generates a URL or download link for it, which is then sent to the user's device so that the user can view or download the content.

[1755] Through this series of steps, personalized tourism content is generated and provided based on the user's emotions.

[1756] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1757] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1758] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1759] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1760] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1761] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1762] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1763] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1764] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1765] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1766] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1767] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1768] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1769] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1770] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1771] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1772] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1773] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1774] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1775] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1776] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1777] The following is further disclosed regarding the above embodiment.

[1778] (Claim 1)

[1779] means for inputting the name of a locality;

[1780] means for collecting information based on the name of the locality;

[1781] means for generating scenarios based on the collected information;

[1782] a means for generating a video based on the generated scenario;

[1783] means for generating audio corresponding to the video;

[1784] a means for creating content by integrating the scenario, video and audio;

[1785] A system including a means for providing created content.

[1786] (Claim 2)

[1787] 10. The system of claim 1, further comprising means for collecting tourist information, historical information and cultural information based on the name of the region.

[1788] (Claim 3)

[1789] 2. The system according to claim 1, further comprising means for saving the created content at a distribution destination and generating a distribution destination URL or a download link.

[1790] "Example 1"

[1791] (Claim 1)

[1792] means for inputting the name of a locality;

[1793] means for collecting information based on the name of the locality;

[1794] means for generating scenarios based on the collected information;

[1795] a means for generating a video based on the generated scenario;

[1796] means for generating audio corresponding to the video;

[1797] a means for creating content by integrating the scenario, video and audio;

[1798] A system that includes a means to store the created content and generate a URL or download link for it.

[1799] (Claim 2)

[1800] 10. The system of claim 1, further comprising means for collecting tourist information, historical information and cultural information based on the name of the region.

[1801] (Claim 3)

[1802] 10. The system of claim 1, further comprising means for generating a scenario using a generative AI model based on the collected information.

[1803] (Claim 4)

[1804] The system according to claim 1, further comprising means for transmitting the generated scenario to an image generation AI and generating an image.

[1805] (Claim 5)

[1806] The system according to claim 1, further comprising means for transmitting the generated scenario and video to an audio generation AI and generating audio.

[1807] (Claim 6)

[1808] 10. The system of claim 1, further comprising means for using multimedia editing software to integrate the script, video and audio.

[1809] (Claim 7)

[1810] 2. The system according to claim 1, further comprising means for saving the created content at a distribution destination and generating a distribution destination URL or a download link.

[1811] "Application Example 1"

[1812] (Claim 1)

[1813] means for inputting the name of a locality;

[1814] means for collecting information based on the name of the locality;

[1815] means for generating scenarios based on the collected information;

[1816] a means for generating a video based on the generated scenario;

[1817] means for generating audio corresponding to the video;

[1818] a means for creating content by integrating the scenario, video and audio;

[1819] a content distribution means including a display device for visually presenting the created content;

[1820] A system including:

[1821] (Claim 2)

[1822] 10. The system of claim 1, further comprising means for collecting tourist information, historical information and cultural information based on the name of the region.

[1823] (Claim 3)

[1824] 2. The system according to claim 1, further comprising means for saving the created content at a distribution destination and generating a distribution destination URL or a download link.

[1825] "Example 2: Combining Emotion Engines"

[1826] (Claim 1)

[1827] means for inputting the name of a locality;

[1828] means for collecting information based on the name of the locality;

[1829] means for generating scenarios based on the collected information;

[1830] a means for generating a video based on the generated scenario;

[1831] means for generating audio corresponding to the video;

[1832] a means for creating content by integrating the scenario, video and audio;

[1833] a means for providing the created content;

[1834] means for analyzing the input or operation of the user to recognize emotions;

[1835] means for adjusting a scenario, video and audio based on the emotion;

[1836] A system including:

[1837] (Claim 2)

[1838] 10. The system of claim 1, further comprising means for collecting tourist information, historical information and cultural information based on the name of the region.

[1839] (Claim 3)

[1840] 2. The system according to claim 1, further comprising means for saving the created content at a distribution destination and generating a distribution destination URL or a download link.

[1841] "Application example 2 when combining emotion engines"

[1842] (Claim 1)

[1843] means for inputting the name of a locality;

[1844] means for collecting information based on the name of the locality;

[1845] means for generating scenarios based on the collected information;

[1846] a means for generating a video based on the generated scenario;

[1847] means for generating audio corresponding to the video;

[1848] a means for creating content by integrating the scenario, video and audio;

[1849] a means for providing the created content;

[1850] means for recognizing a user's emotion;

[1851] a means for personalizing content based on perceived sentiment;

[1852] A system including:

[1853] (Claim 2)

[1854] 10. The system of claim 1, further comprising means for collecting tourist information, historical information and cultural information based on the name of the region.

[1855] (Claim 3)

[1856] 2. The system according to claim 1, further comprising means for saving the created content at a distribution destination and generating a distribution destination URL or a download link. [Explanation of symbols]

[1857] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for inputting the name of a locality; means for collecting information based on the name of the locality; means for generating scenarios based on the collected information; a means for generating a video based on the generated scenario; means for generating audio corresponding to the video; a means for creating content by integrating the scenario, video and audio; A system including a means for providing created content.

2. 2. The system of claim 1, further comprising means for collecting tourist information, historical information and cultural information based on the name of the locality.

3. 2. The system according to claim 1, further comprising means for saving the created content at a distribution destination and generating a distribution destination URL or a download link.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A