System

The system addresses the inefficiencies in documentary production by automating information collection, credibility assessment, and story construction to generate high-quality documentary footage efficiently and accurately.

JP2026022482APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123999
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Traditional documentary production is time-consuming and resource-intensive, and the challenges include information gathering, credibility assessment, and story construction, exacerbated by the prevalence of fake news and inaccurate information, necessitating a system for efficient generation of high-quality documentary footage.

Method used

A system that accepts input of a theme or keyword, collects and evaluates information, categorizes it for reliability, sets target audience and story genre, generates a story structure, and composes scenes to render high-quality documentary video.

Benefits of technology

Enables efficient and accurate generation of high-quality documentary video by automating information collection, credibility assessment, and story construction, while filtering out fake news.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022482000001_ABST
    Figure 2026022482000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving an input of a theme or a keyword; means for collecting information related to the theme or the keyword from the Internet and an internal database; means for evaluating reliability of the collected information and determining whether the information is true or false; means for classifying the evaluated information by category; means for generating a story structure based on the set viewing condition; and means for rendering and outputting a final documentary video.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Traditional documentary production requires a lot of time and resources, and is particularly problematic due to the complicated processes of information gathering, credibility assessment, and story construction. Furthermore, in a world flooded with fake news and inaccurate information, selecting accurate and reliable information is also a challenge. Therefore, a system that can efficiently generate high-quality documentary footage is needed. [Means for solving the problem]

[0005] The present invention provides a system that accepts input of a theme or keyword, collects information related to that theme, and evaluates and classifies its reliability. Specifically, the system includes means for accepting input of the theme or keyword, means for collecting information from the Internet and an internal database, means for evaluating the reliability of the collected information and determining its authenticity, means for categorizing the evaluated information, means for the user to set the target audience, story genre, and emotion to be conveyed, means for generating a story structure based on the settings, means for setting the intention of each cut based on the generated story structure and composing scenes, and means for rendering and outputting the final documentary video. This system enables the efficient and accurate generation of high-quality documentary video.

[0006] "Themes and keywords" refer to concepts and phrases that the user wants to focus on when producing documentary footage.

[0007] "Information" refers to data collected from the internet and internal databases, including various media formats such as text, images, video, and audio.

[0008] "Means of collection" refers to the technical capabilities for obtaining information related to themes and keywords from the Internet and internal databases.

[0009] "Measures to assess trustworthiness and determine authenticity" refers to analytical tools and algorithms used to determine whether collected information is accurate and trustworthy.

[0010] "Means for categorizing" is a function for organizing collected information by specific themes or categories (e.g., scientific data, news, policy responses, etc.).

[0011] "Target audience" refers to the person or group who is intended to view the documentary footage.

[0012] "Story genre" refers to the type or format of story the video belongs to (e.g., documentary, news, educational, etc.).

[0013] "Emotions to be conveyed" are specific emotions or feelings (e.g., tension, hope) that you want to convey to viewers through the documentary footage.

[0014] The "means for generating a story structure" is a function for constructing the narrative flow of the entire video based on the collected information and the user's settings.

[0015] "Means for setting the intention of each cut and composing a scene" refers to a technique for selecting images and narration that align with the user's intentions for each part of a documentary video and putting them together as an entire scene.

[0016] "Means for rendering and outputting documentary footage" is a function for generating edited footage as a final file and providing it in a format that can be viewed by users. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] This system provides a technical means for efficiently generating documentary footage. A specific embodiment will now be described.

[0039] The user first inputs the documentary's theme or keywords, which are the initial instructions for the system, for example, "marine pollution."

[0040] The server then collects information based on the input topic and keywords, searches the internet and internal databases for relevant information, and collects the necessary text, images, video, and audio data. The collected data is then stored in storage.

[0041] To assess the reliability of the collected information, the server uses natural language processing (NLP) techniques to calculate a credibility score for each piece of information based on its source, number of citations, related research data, etc. It also uses social media analysis tools to filter out fake news and inaccurate information.

[0042] Once the information has been assessed for trustworthiness, it is then sorted by the server into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews." Each piece of information is automatically labeled with the relevant category.

[0043] Next, the user sets the target audience, the genre of the story, the emotion they want to convey, etc. These settings are input through a user interface.

[0044] The server selects a story template based on the input settings and generates the story structure. This process uses a template matching algorithm to select the most suitable story template. The selected template is then incorporated with the collected information to generate the storyboard.

[0045] The next step is to set the intention for each cut and compose the scene. The device receives instructions from the user and selects the appropriate video clip for each cut. At this stage, video clips are dragged and dropped, narration is added, and transition effects are set.

[0046] Finally, the server renders all the scenes and outputs the final documentary video, which is generated in the specified output format (e.g. MP4, MOV) and made available to the user for download or viewing.

[0047] As a concrete example, consider a user creating a documentary on the theme of "marine pollution." First, the user enters "marine pollution," and the server collects the latest research papers, news articles, government reports, and footage from environmental organizations. After assessing the reliability and eliminating fake news, the information is classified into categories such as "scientific data," "news," and "policy responses." The user sets the target audience as general viewers, the story genre as documentary, and the emotion they want to convey as tension. The server composes the story based on a template, and the device places the footage and narration for each cut. Finally, the server integrates all the scenes and outputs the completed documentary footage.

[0048] In this way, the present invention allows for efficient and accurate generation of high quality documentary footage.

[0049] The processing flow will be explained below.

[0050] Step 1:

[0051] The user inputs the documentary's theme or keywords. For example, "marine pollution."

[0052] Step 2:

[0053] The server collects information based on themes and keywords. This includes retrieving relevant text, images, video, and audio data from the internet and internal databases. The server gathers the necessary data through web scraping and API integration.

[0054] Step 3:

[0055] The server evaluates the credibility of the collected information, using natural language processing (NLP) techniques to analyze each piece of information's source, number of citations, and related research data to calculate a credibility score, and also uses social media analysis tools to identify and remove fake news and inaccurate information.

[0056] Step 4:

[0057] The server categorizes the evaluated information into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and automatically labels and organizes the information.

[0058] Step 5:

[0059] The user sets the target audience, the genre of the story, and the emotion they want to convey. These settings are input through a user interface. For example, the target audience is set to general viewers, the genre is set to documentary, and the emotion they want to convey is set to tension.

[0060] Step 6:

[0061] The server selects a story template based on the input settings, uses a template matching algorithm to select the most suitable template, and places the collected information into the storyboard.

[0062] Step 7:

[0063] The device determines the intent of each cut and composes the scene. The user uses drag and drop to place the appropriate video clips for each cut, add narration, and set transition effects. The device then makes the final adjustments to the scene.

[0064] Step 8:

[0065] The server stitches all the scenes together and renders the final documentary video, generating the final video file in the specified output format (e.g. MP4, MOV) for users to download or watch.

[0066] Through the above steps, users can efficiently generate high-quality documentary footage.

[0067] Example 1

[0068] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0069] In today's information society, extracting reliable information from vast amounts of data and efficiently creating video content based on that information is a major challenge. While documentary video production in particular requires accurate and reliable information, it also requires a great deal of time and effort, including information gathering, reliability assessment, story construction, and video editing. Therefore, there is a demand for technological tools that allow users to easily generate high-quality documentary video.

[0070] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0071] In this invention, the server includes means for accepting input of a theme or keyword, means for collecting information related to the theme or keyword from a network and a database, means for evaluating the reliability of the collected information using natural language processing and determining its authenticity, means for categorizing the evaluated information, means for the user to set the target audience, story genre, and emotion to be conveyed, means for generating a story structure based on the set viewing conditions, means for setting the intention of each cut and composing scenes based on the generated story structure, means for rendering and outputting the final video content, means for identifying inaccurate information using a social media analysis tool, and means for selecting a story template optimal for the viewing conditions using a template matching algorithm. This allows the user to efficiently generate high-quality video content based on reliable information while eliminating cumbersome tasks.

[0072] "Theme" is a keyword that the user inputs as the central topic of the documentary video.

[0073] "Keywords" are relevant search terms that the system uses to gather information.

[0074] "Information" refers to any data, including text, images, video, and audio data, collected from networks and databases.

[0075] "Reliability" is the criterion for assessing whether the collected information is accurate and trustworthy.

[0076] "Natural language processing (NLP)" is a technology that analyzes collected text data and evaluates its reliability and categorizes it.

[0077] "Category" refers to the field or type used to classify collected information, such as "scientific data," "news articles," and "policy reports."

[0078] "Target audience" refers to the main audience of the documentary footage, and includes general viewers and experts.

[0079] "Story genre" refers to the type or style of documentary footage, including, for example, documentaries and educational films.

[0080] The "emotions to be conveyed" refer to the emotions and atmosphere that the viewer wants to convey through the video, and include, for example, tension and emotion.

[0081] "Story construction" involves organizing information and planning the flow of the footage based on set viewing conditions.

[0082] A "scene" refers to an individual scene or cut from a documentary film.

[0083] "Rendering" is the process of generating the entire image based on the collected information and the user's settings.

[0084] "Social Media Analytics Tools" are tools used to collect data from social media sites and identify inaccurate information.

[0085] The "template matching algorithm" is an algorithm for selecting the optimal story template based on the set viewing conditions.

[0086] "Video Content" refers to the final video file created based on the collected information.

[0087] This system provides a technical means for efficiently generating documentary footage. A specific embodiment will now be described.

[0088] System configuration

[0089] The system consists of the following main components: a user interface (GUI), a data collection module, a natural language processing (NLP) module, a template matching algorithm, a video editing module, and a rendering module.

[0090] User Interface

[0091] The user first inputs the documentary's theme and keywords through the GUI. This input is the initial instruction for the system, and for example, a word such as "marine pollution" is entered.

[0092] Data Collection Module

[0093] The server collects information based on the input topic or keyword, searches for related information from the network or internal database, and collects the necessary text, image, video, and audio data. For example, it uses Google Custom Search API to collect information from the Internet and pulls the necessary information from the internal database.

[0094] Natural Language Processing (NLP) Module

[0095] The server uses NLP technology to evaluate the reliability of the collected information. Specifically, it calculates a reliability score for each piece of information based on its source, number of citations, and related research data. It also uses social media analysis tools to filter out fake news and inaccurate information. NLP libraries such as spaCy and NLTK are used.

[0096] Categorizing information

[0097] Once the information has been assessed for reliability, it is sorted by the server into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and each piece of information is automatically labeled with the relevant category.

[0098] Viewing condition settings

[0099] Next, the user sets the target audience, the genre of the story, the emotions they want to convey, etc. These settings are again entered through the GUI and sent to the server. For example, possible settings include "general audience," "documentary," and "tension."

[0100] Selecting and configuring a story template

[0101] The server selects a story template based on the user's input settings and generates the story structure using a template matching algorithm. The selected template is then combined with the collected information to generate the storyboard.

[0102] Video editing

[0103] The device receives user instructions and selects the appropriate video clip for each cut. At this stage, video clips are dragged and dropped, narration is added, and transition effects are set. A preview of the video can also be viewed on the GUI.

[0104] Video rendering and output

[0105] Finally, the server renders all the scenes and outputs the final documentary video. The video is generated using a rendering tool such as FFmpeg in the specified format (e.g., MP4, MOV) and made available to users for download or viewing.

[0106] Specific examples

[0107] Consider a case where a user wants to create a documentary on the theme of "marine pollution." The user first enters "marine pollution," and the server uses the Google Custom Search API and an internal database to collect the latest research papers, news articles, government reports, and footage from environmental organizations. The collected information is evaluated for reliability using NLP, and inaccurate information is eliminated. Each piece of information is categorized as "scientific data," "news," "policy response," etc. The user sets viewing conditions such as "general audience," "documentary," and "urgency," and the server composes a story based on the template. The device arranges the footage and narration for each cut, and finally, the server integrates all the scenes and outputs the completed documentary video.

[0108] Prompt Sentence Examples

[0109] As a prompt sentence, provide the following sentence as input to the generative AI model:

[0110] Theme: Marine pollution

[0111] Keywords: plastics, climate change, marine life

[0112] Target audience: General audience

[0113] Story genre: Documentary

[0114] Emotions to convey: Urgency

[0115] Sources: networks, databases, news articles, reports, research papers

[0116] In this way, the present invention allows users to effectively and accurately generate high quality documentary footage.

[0117] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0118] Step 1: User enters topic and keywords

[0119] The user logs into the system's user interface (GUI) and is taken to a screen where they can enter a theme and keywords. The user enters the keywords "marine pollution" and "plastic" in the text field and clicks the "Search" button.

[0120] Input: topic (marine pollution) and keyword (plastic)

[0121] Output: The entered theme and keywords are sent to the server.

[0122] Step 2: Server collects information

[0123] The server receives the input topic and keywords and collects related information from the network and internal database. The server uses the Google Custom Search API to perform an internet search for the query "ocean pollution plastic" and scrapes useful information (text, images, video, audio data). It also queries the internal database to pull out related information.

[0124] Input: topic (marine pollution) and keyword (plastic)

[0125] Output: A list of collected text, image, video, and audio data

[0126] Step 3: The server evaluates trustworthiness using natural language processing (NLP)

[0127] The server applies NLP techniques to the collected information, specifically using the spaCy and NLTK libraries to calculate a credibility score for each piece of information. The server also evaluates the source, number of citations, and related research data for each piece of information, and uses social media analysis tools to filter out inaccurate information.

[0128] Input: A list of collected text, image, video, and audio data

[0129] Output: List of information with confidence scores

[0130] Step 4: The server categorizes the information

[0131] The server categorizes the information whose reliability has been evaluated. It analyzes the content of the text data and automatically categorizes it into categories such as "scientific data," "news articles," "policy reports," and "personal interviews." The server updates and saves this categorization information in a database.

[0132] Input: List of information with confidence scores

[0133] Output: A list of information sorted by category

[0134] Step 5: User sets target audience, genre, and emotion

[0135] The user again uses the GUI to set the target audience, the genre of the story, and the emotion they want to convey. Users select settings such as "general audience," "documentary," or "urgency" from drop-down menus and click the "Set" button.

[0136] Input: audience, story genre, emotion you want to convey

[0137] Output: Configuration information sent to the server

[0138] Step 6: The server selects a story template and generates a story structure.

[0139] The server runs a template matching algorithm based on the viewing conditions you set and selects the most suitable story template. The collected information is automatically incorporated into the selected template to generate a storyboard.

[0140] Input: audience, story genre, emotion to convey, list of information categorized

[0141] Output: Storyboard

[0142] Step 7: Select and edit video clips on your device

[0143] The user uses the device to select and edit video clips based on the generated storyboard. They drag and drop the appropriate video clips for each scene, add narration, and set transition effects. Editing proceeds while checking the video preview on the GUI.

[0144] Input: Storyboard, video clips

[0145] Output: Edited video sequence

[0146] Step 8: The server renders the entire scene and outputs the final video.

[0147] The server uses a rendering tool such as FFmpeg to combine all the scenes and render the video in the specified output format (e.g. MP4, MOV). The final video file is provided to the user, and a download link is displayed on the GUI.

[0148] Input: Edited video sequence

[0149] Output: Final video file (e.g. MP4, MOV)

[0150] Through the above steps, the user can efficiently generate high-quality documentary footage.

[0151] (Application example 1)

[0152] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0153] In today's world, documentary video production requires a high level of skill and time. Gathering accurate information tailored to the theme, assessing its reliability, and structuring a story that meets the needs of viewers are particularly challenging. Furthermore, amid a proliferation of fake information, there is a demand for video production that uses accurate and reliable information. Rendering and outputting the final video work is also a laborious process that requires increased efficiency. To address these challenges, there is an urgent need to provide a system that allows users to easily and efficiently generate and distribute high-quality documentary video.

[0154] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0155] In this invention, the server includes means for accepting input of a theme or keyword, means for collecting information related to the theme or keyword from a communication network and an internal database, means for evaluating the reliability of the collected information and determining its veracity, means for classifying the evaluated information by category, means for the user to set the target audience, story genre, and desired emotion, means for generating a story structure based on the set viewing conditions, means for setting the intention of each cut and composing scenes based on the generated story structure, means for rendering and outputting the final video work, and means for providing a user interface for the user to input a theme and automatically generate and distribute documentary video based on the collected information, thereby enabling users to easily and efficiently generate and distribute high-quality documentary video.

[0156] "Theme" is a concept that allows the user to indicate the central subject or interest of the documentary footage.

[0157] "Keywords" are indicators that contain important words or terms related to a topic and are used to narrow down information.

[0158] A "communications network" is a digital network for sending and receiving information, including the Internet.

[0159] An "internal database" is a source of information held within the system and is a storage device for storing and managing a portion of the collected data.

[0160] "Credibility" is a measure of the accuracy and veracity of collected information, and is calculated using NLP technology.

[0161] "Veracity" is the quality that determines whether information is based on actual, verified facts.

[0162] "User" refers to an individual or organization that performs operations to create documentary footage.

[0163] "Target audience" refers to the audience or target group that intends to watch the documentary footage.

[0164] "Story genre" is a classification that indicates the format and style of documentary footage.

[0165] "Emotions you want to convey" refers to the emotions and atmosphere you want viewers to feel through the documentary footage.

[0166] The "viewing conditions" are parameters related to the viewer set by the user, such as the viewing target, the genre of the story, and the emotion to be conveyed.

[0167] "Story structure" is the framework for designing the overall flow and development of a documentary film.

[0168] "The intention of each cut" refers to the purpose and message of each scene or cut in the video.

[0169] "Scenes" are the individual scenes or cuts that make up documentary footage.

[0170] "Video work" is a general term for the final documentary video.

[0171] "Rendering" is the process of converting an edited video clip into an output format as a single file.

[0172] "Output" refers to the act of saving or distributing the final video work in a specified format.

[0173] "User interface" refers to the screen and operating means through which a user interacts with and operates a system.

[0174] This invention relates to a system that efficiently generates and distributes documentary footage by allowing users to input themes and keywords. The system is configured as follows:

[0175] First, users input the documentary's theme and keywords through the smartphone application interface. For example, they can specify a theme such as "marine pollution." To accept this input, the system uses libraries such as React Native as its front end.

[0176] The server receives the input topic and keywords and collects related information from communication networks and internal databases. In this step, a crawler is used to collect text, images, video, and audio data from academic papers, news sites, government sites, and other sources on the Internet. The collected data is then stored in storage. The software used is Python and technologies such as Beautiful Soup and Scrapy.

[0177] Next, the reliability of the collected information is evaluated. TensorFlow or PyTorch is used to calculate a reliability score for the information using natural language processing (NLP) techniques. Furthermore, fake news and inaccurate information is filtered out using social media analysis tools. This analysis uses techniques such as sentiment analysis and named entity recognition (NER).

[0178] The evaluated information is then sorted into categories on the server, such as "scientific data," "news articles," "policy responses," and "personal interviews." The information is sorted into categories using labeling technology based on machine learning models.

[0179] Next, the user sets the target audience, story genre, and desired emotion on the application. These settings are necessary for the server to generate the optimal story structure. A template matching algorithm is used to select a story template and automatically generate the storyboard.

[0180] Based on the generated storyboard, the intention of each cut is set and the scene is composed. At this stage, the user inserts video clips, adds narration, and sets transition effects through the interface. The front-end editing function uses HTML5, CSS3, and JavaScript.

[0181] FFmpeg is used to render the final video and save it in the output format, which is then made available to users for download.

[0182] As a concrete example, consider a case where a user is creating a documentary on the theme of "marine pollution." When the user enters "marine pollution," the server collects the latest scientific data, news articles, and policy response information from the Internet. It evaluates the reliability of the collected information and categorizes it into categories such as "scientific data," "news," and "policy response." The user sets the target audience as general viewers, the story genre as documentary, and the emotion they want to convey as tension. The server composes the story based on a template and edits each cut through the user interface. Finally, the server integrates all the scenes and outputs the completed documentary video in MP4 format.

[0183] Examples of prompts include:

[0184] "A user wants to create a documentary about marine pollution. They need to gather the latest research, news, policy reports, etc., create a storyboard based on that information, and finally output it in MP4 format."

[0185] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0186] Step 1:

[0187] Users enter the documentary's theme and keywords into the smartphone application.

[0188] Input: Theme or keyword (e.g., marine pollution)

[0189] Output: Themes and keywords are sent to the server

[0190] Step 2:

[0191] The server collects related information from the communication network and internal database based on the input topic and keywords.

[0192] Input: Theme or keyword

[0193] Data processing: Using crawlers to collect information from academic papers, news sites, government sites, etc. on the Internet and store it in storage.

[0194] Output: Collected text, images, video, and audio data

[0195] Step 3:

[0196] The server uses NLP techniques to assess the reliability of the collected information.

[0197] Input: Collected information

[0198] Data Computation: Use TensorFlow or PyTorch to calculate the credibility score of information. Use social media analysis tools to filter out fake news and inaccurate information.

[0199] Output: reliable information

[0200] Step 4:

[0201] The server categorizes the information whose reliability has been evaluated.

[0202] Input: Information that has undergone a reliability assessment

[0203] Data processing: Using machine learning models, information is labeled into categories such as "scientific data," "news articles," "policy responses," and "personal interviews."

[0204] Output: Information sorted by category

[0205] Step 5:

[0206] Users set their viewing target, story genre, and the emotion they want to convey on the application.

[0207] Input: Target audience, story genre, and desired emotion (e.g., general audience, documentary, suspense)

[0208] Output: The configured viewing conditions are sent to the server.

[0209] Step 6:

[0210] The server selects the most suitable story template based on the set viewing conditions and generates a storyboard.

[0211] Input: Viewing conditions, information categorized by category

[0212] Data calculation: Using a template matching algorithm, a storyboard is automatically generated by selecting the story template that best suits the viewing conditions.

[0213] Output: Generated storyboard

[0214] Step 7:

[0215] The user sets the intention for each cut based on the generated storyboard and composes the scene.

[0216] Input: Generated storyboard

[0217] Specific operations: The user inserts a video clip, adds narration, and sets transition effects in the application interface.

[0218] Output: Edited storyboard

[0219] Step 8:

[0220] The server renders the final video production and outputs it in the specified format.

[0221] Input: Edited storyboard

[0222] Specific operation: All scenes are integrated using FFmpeg and the final video is rendered in MP4 format or similar.

[0223] Output: The finished documentary footage is made available to users in a downloadable format

[0224] As an example, the following prompt sentence is used:

[0225] "A user wants to create a documentary about marine pollution. They need to gather the latest research, news, policy reports, etc., create a storyboard based on that information, and finally output it in MP4 format."

[0226] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0227] The present invention is a system that collects related information based on specific themes and keywords, evaluates and classifies reliability, and generates high-quality documentary footage by recognizing the user's viewing preferences and emotions.

[0228] System Configuration

[0229] 1. How to input themes and keywords

[0230] The user inputs the documentary's theme or keywords. For example, "marine pollution." The user interface (UI) is designed to be intuitive.

[0231] 2. Information gathering methods

[0232] The server collects relevant information from the internet and internal databases using web scraping and API integration to obtain text, images, video, and audio data.

[0233] 3. Reliability assessment methods

[0234] The server uses natural language processing (NLP) techniques to assess the reliability of collected information, analyzing its source, number of citations, and related research data to calculate a credibility score. It also uses social media analysis tools to identify and filter out fake news.

[0235] 4. Information classification means

[0236] The server categorizes the evaluated information into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and automatically labels each piece of information with the appropriate category.

[0237] 5. Viewing setting input method

[0238] The user sets the target audience (e.g., general audience), the genre of the story (e.g., documentary), and the emotion they want to convey (e.g., tension). These settings are made through a user interface.

[0239] 6. Emotion Engine

[0240] The emotion engine analyzes the user's emotions in real time as they type, and reflects this in the story structure and scene settings. Emotion analysis is performed based on the user's facial expressions, tone of voice, and typing speed.

[0241] 7. Story Generation Methods

[0242] The server selects the optimal story template based on the viewing preferences and the analysis results of the emotion engine, and generates a story structure. Using a template matching algorithm, the collected information is arranged on the storyboard.

[0243] 8. Scene Composition Methods

[0244] The device allows the user to compose scenes by setting the intention for each cut. Using drag and drop, the user places the appropriate video clips for each cut, adds narration, and sets transition effects. Based on the emotional state recognized by the emotion engine, appropriate narration and music are automatically selected and placed.

[0245] 9. Video rendering and output methods

[0246] The server stitches all the scenes together and renders the final documentary video. The video file is generated in the specified output format (e.g. MP4, MOV) and made available to the user for download or viewing.

[0247] Specific examples

[0248] If a user wants to create a documentary on the theme of "marine pollution," they first input the theme. The server collects related information and evaluates and classifies it for reliability. The user then selects a general audience, a documentary genre, and a desire to convey a sense of tension. The emotion engine recognizes the user's sense of tension as they input, and reflects this in the story structure and scene settings. Appropriate video clips, narration, and music are selected, and finally, the server renders all the scenes, outputting the finished documentary video.

[0249] In this way, the present invention provides a system that efficiently generates high-quality documentary footage while taking into account the user's emotions.

[0250] The processing flow will be explained below.

[0251] Step 1:

[0252] The user inputs the documentary's theme or keywords. For example, "marine pollution."

[0253] Step 2:

[0254] The server collects information based on themes and keywords. The information is collected using web scraping and API integration to obtain text, images, video, and audio data from the internet and internal databases.

[0255] Step 3:

[0256] The server evaluates the credibility of the collected information, using natural language processing (NLP) techniques to analyze the information's source, number of citations, and related research data to calculate a credibility score, and also uses social media analysis tools to identify and remove fake news and inaccurate information.

[0257] Step 4:

[0258] The server categorizes the evaluated information into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and each piece of information is automatically labeled with the appropriate category.

[0259] Step 5:

[0260] The user sets the target audience, the genre of the story, and the emotion they want to convey. These settings are entered through a highly visible user interface. For example, the target audience can be set to general viewers, the genre to documentary, and the emotion they want to convey to be tension.

[0261] Step 6:

[0262] The emotion engine analyzes the user's emotions in real time, analyzing facial expressions, tone of voice, and input speed when entering settings to recognize the user's current emotional state.

[0263] Step 7:

[0264] The server selects the optimal story template based on the user's viewing preferences and the analysis results of the emotion engine. Using a template matching algorithm, it selects the template that best suits the preferences and analysis data, and places the collected information on the storyboard.

[0265] Step 8:

[0266] The device sets the intention for each cut and composes the scene. The user uses drag and drop to place the appropriate video clips for each cut, add narration, and set transition effects. Appropriate narration and music are also automatically selected and placed based on the emotional state recognized by the emotion engine.

[0267] Step 9:

[0268] The server stitches all the scenes together and renders the final documentary video. The video file is generated in the specified output format (e.g. MP4, MOV) and made available to the user for download or viewing.

[0269] By following these steps, users can efficiently generate high-quality documentary footage that reflects emotions.

[0270] Example 2

[0271] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0272] In conventional documentary video production systems, the processes of information collection, credibility assessment, information classification, viewing settings, emotion analysis, story generation, scene composition, and video rendering are all performed manually, which is extremely time-consuming and labor-intensive. Furthermore, they lack the ability to analyze user emotions in real time and reflect them in video production, resulting in a lack of emotional impact on viewers. Furthermore, with the rise of fake news, unreliable information is often mixed in, making it difficult to produce high-quality documentary videos based on accurate information.

[0273] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for accepting input of themes and keywords, means for collecting information from a network and an internal database, means for evaluating the reliability of the collected information and determining its authenticity, means for identifying fake information using social media analysis technology, means for categorizing the evaluated information, means for the user to set the viewing target, story genre, and emotion to be conveyed, means for analyzing the user's emotions in real time using an emotion analysis means and reflecting the results in scene settings, means for selecting a story template optimal for the viewing conditions using a template matching algorithm, means for setting the intention of each scene based on the generated story structure and composing scenes, and means for synthesizing and outputting the final documentary video. This enables automatic collection and reliability evaluation of information, appropriate information classification, real-time analysis of user emotions, and automatic selection of the optimal story template, thereby enabling efficient generation of high-quality documentary video with emotional impact.

[0274] The "means for accepting input of themes and keywords" is a means for providing an interface that allows a user to input themes and keywords necessary for producing a documentary video.

[0275] "Means for collecting information from networks and internal databases" refers to means for automatically collecting information related to the entered topic or keyword via the Internet or internal databases.

[0276] "Means for assessing the reliability of collected information and determining its authenticity" refers to means for assessing the reliability of collected information using natural language processing technology and source analysis, and determining the authenticity of the information.

[0277] "Methods for identifying fake information using social media analysis technology" refers to methods that use technology to analyze data on social media and identify misinformation and fake news.

[0278] A "means of categorizing evaluated information" is a means of classifying information that has been collected and evaluated for reliability into specific categories, such as "scientific data" or "news articles."

[0279] "A means for users to set the target audience, story genre, and emotions they want to convey" is a means for providing an interface that allows users to set the audience, genre, and emotions they want to convey for the documentary footage.

[0280] "Means for analyzing a user's emotions in real time using emotion analysis means and reflecting the emotions in scene settings" refers to means for analyzing a user's facial expressions and tone of voice, grasping the user's emotional state in real time, and reflecting the results in scene settings.

[0281] "Means for selecting a story template optimal for viewing conditions using a template matching algorithm" refers to means for executing an algorithm that selects the optimal template from among pre-prepared story templates based on the user's viewing settings and emotion analysis results.

[0282] "Means for setting the intention of each scene based on the generated story structure and composing scenes" refers to means for setting the purpose and intention of each scene based on the generated story and arranging corresponding images and narration.

[0283] The "means for compositing and outputting the final documentary footage" refers to a means for combining all the scenes into one, generating the final documentary footage, and outputting it in a specified format.

[0284] This invention is a system that collects related information based on specific themes and keywords, evaluates and classifies its reliability, recognizes the user's viewing preferences and emotions, and automatically generates high-quality documentary footage.

[0285] System Configuration

[0286] The system has the following main components:

[0287] 1. How to input themes and keywords

[0288] Users input the documentary's theme and keywords through a user interface (UI), which is designed for intuitive operation.

[0289] 2. Information gathering methods

[0290] The server collects relevant information from the internet and internal databases using web scraping tools (e.g., BeautifulSoup) and API integration (e.g., Twitter API), which allows it to obtain text, images, video, and audio data.

[0291] 3. Reliability assessment methods

[0292] The server evaluates the reliability of the collected information using natural language processing (NLP) techniques (e.g., BERT, GPT-3), and also identifies and eliminates fake information using social media analysis techniques.

[0293] 4. Information classification means

[0294] The server categorizes the evaluated information using a clustering algorithm (e.g., K-means) to automatically categorize it into categories such as "scientific data," "news articles," "policy reports," and "personal interviews."

[0295] 5. Viewing setting input method

[0296] Users can set their target audience, the genre of the story, and the emotion they want to convey through a user interface.

[0297] 6. Emotion analysis method

[0298] The server is equipped with an emotion analysis engine that analyzes the user's emotions in real time using software that analyzes facial expressions (e.g., OpenCV) and voice tone (e.g., Deepgram), and also combines the user's input speed for analysis.

[0299] 7. Story Generation Methods

[0300] The server runs a template matching algorithm to select the most appropriate story template based on viewing preferences and sentiment analysis results, and places the collected information into the appropriate storyboard.

[0301] 8. Scene Composition Methods

[0302] The user uses the drag-and-drop function on the device to place appropriate video clips for each cut on a storyboard, and then sets narration and transition effects. Appropriate narration and music are automatically selected based on the emotional state recognized by the emotion analysis means.

[0303] 9. Video rendering and output methods

[0304] The server combines all the scenes to generate the final documentary video and renders it as a video file in the specified format, such as MP4 or MOV.

[0305] Specific examples

[0306] Consider a scenario where a user is creating a documentary on the theme of "marine pollution." First, the user inputs the theme. The server collects related information, evaluates its reliability, and classifies it. Next, the user sets the target audience as "general audience," the genre as "documentary," and the emotion they want to convey as "tension." The emotion analysis engine recognizes the user's emotions in real time as they are input, and reflects them in the story structure and scene settings. Appropriate video clips, narration, and music are selected, and the server renders all scenes to output the completed documentary video.

[0307] Prompt Sentence Examples

[0308] "Collect information to create a documentary film on the following topic, assess its reliability, classify it, generate a story, and output the final film. The user settings are for general audiences, the genre is documentary, and the emotion you want to convey is a sense of urgency. The topic is 'Marine Pollution.'"

[0309] This system makes it possible to efficiently generate high-quality documentary footage while taking into account the user's emotions.

[0310] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0311] Step 1:

[0312] Accepts input of themes and keywords.

[0313] The user inputs themes and keywords through a user interface (UI).

[0314] Specific behavior: The user enters "Marine Pollution" in the topic field and clicks the submit button.

[0315] Input: Theme or keyword (e.g. "Marine pollution")

[0316] Output: Themes and keywords entered

[0317] Step 2:

[0318] Automatic Collection of Information.

[0319] The server collects information from the Internet and internal databases based on the entered themes and keywords.

[0320] Specific operation: The server uses a web scraping tool (e.g., BeautifulSoup) to retrieve articles from Google search results and uses an API (e.g., Twitter API) to collect related tweets.

[0321] Input: Theme or keyword (e.g. "Marine pollution")

[0322] Output: Collected information (text, images, video, audio data)

[0323] Step 3:

[0324] Evaluating the reliability of information.

[0325] The server evaluates the reliability of the collected information using natural language processing (NLP) techniques.

[0326] How it works: The server uses NLP models (e.g., BERT, GPT-3) to calculate a credibility score for each article, and also uses social media analysis techniques to identify and remove misinformation.

[0327] Input: Collected information

[0328] Output: Information with assessed reliability (with score)

[0329] Step 4:

[0330] Classification of information.

[0331] The server categorizes the evaluated information.

[0332] What it does: The server uses a clustering algorithm (e.g., K-means) to classify information into categories such as "scientific data," "news articles," "policy reports," and "personal interviews."

[0333] Input: Information with reliability assessment

[0334] Output: Categorically labeled information

[0335] Step 5:

[0336] Enter viewing settings.

[0337] The user inputs the target audience, the genre of the story, and the emotion they want to convey.

[0338] Specific operation: The user selects the target audience as "general audience," the genre as "documentary," and the emotion to be conveyed as "tension."

[0339] Input: target audience, story genre, emotion you want to convey

[0340] Output: Viewing Settings

[0341] Step 6:

[0342] Performing sentiment analysis.

[0343] The server analyzes the user's emotions in real time.

[0344] How it works: The server uses facial expression analysis software (e.g., OpenCV) and voice analysis software (e.g., Deepgram) to detect the user's emotions. The user's typing speed is also used for analysis.

[0345] Input: User's facial expression data, tone of voice, typing speed

[0346] Output: Real-time analyzed emotion data

[0347] Step 7:

[0348] Story generation.

[0349] The server selects a story template based on viewing preferences and emotion analysis results, and generates a story structure.

[0350] Specific operation: The server uses a template matching algorithm to select an appropriate story template and places the collected information on the storyboard.

[0351] Input: Viewing settings, sentiment analysis results, and information categorized by category

[0352] Output: Generated story structure

[0353] Step 8:

[0354] Scene placement and adjustment.

[0355] The terminal allows the user to set the intention for each cut and compose a scene.

[0356] How it works: Users use the drag-and-drop function to place video clips on the storyboard and set narration and transition effects. Appropriate narration and music are automatically selected based on emotional data.

[0357] Input: Generated story structure, video clips, narration, transition effects

[0358] Output: Composed Scene

[0359] Step 9:

[0360] Video rendering and output.

[0361] The server stitches all the scenes together and renders the final documentary footage.

[0362] What it does: The server uses video editing software (e.g., FFmpeg) to combine and render all the scenes into a single video file in MP4 or MOV format.

[0363] Input: Composed Scene

[0364] Output: Final documentary footage file

[0365] (Application example 2)

[0366] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0367] In today's world, users are seeking to generate and view high-quality documentary videos based on their own interests. However, conventional methods have difficulty collecting reliable information and providing it in an optimal format based on the viewer's emotions and preferences. Furthermore, there is no system that can analyze the viewer's emotional state in real time and dynamically change the video content accordingly, which hinders user satisfaction.

[0368] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0369] In this invention, the server includes means for accepting input of themes and keywords, means for collecting information related to the themes and keywords from a communication network and an internal database, means for evaluating the reliability of the collected information and determining its authenticity, means for classifying the evaluated information by type, means for the user to set the viewing target, story genre, and emotion to be conveyed, means for analyzing the user's emotional state in real time, means for generating a story structure based on the set viewing conditions and the analyzed emotional state, means for setting the intention of each cut and composing scenes based on the generated story structure, and means for rendering and outputting the final documentary video. This makes it possible to generate and provide high-quality, reliable documentary video in real time based on the viewing conditions and emotional state set by the user.

[0370] The "means for accepting input of themes and keywords" is an interface that allows a user to input subjects of interest or specific words into the system.

[0371] "Means for collecting information related to the theme or keyword from communication networks and internal databases" refers to a function for collecting information from the Internet and obtaining related information from internally stored databases.

[0372] "Means for assessing the reliability of collected information and determining its authenticity" refers to techniques for analyzing the reliability of the data obtained and determining whether the data is accurate.

[0373] The "means for classifying evaluated information by type" refers to a method for classifying information whose reliability has been evaluated into different categories based on its content.

[0374] "A means for users to set the target audience, story genre, and emotions they want to convey" is an interface that allows users to set the target audience, video genre, and emotions they want to convey through the video.

[0375] The "means for analyzing the user's emotional state in real time" is a technology for detecting and analyzing the user's current emotions in real time.

[0376] The "means for generating a story structure based on the set viewing conditions and the analyzed emotional state" is a method for automatically constructing an optimal story based on the viewing conditions set by the user and the analyzed emotions.

[0377] "Means for setting the intention of each cut based on the generated story structure and composing scenes" refers to a technique for setting the content and purpose of each scene based on the generated story structure and then visualizing it.

[0378] "Means for rendering and outputting the final documentary footage" refers to the function for integrating all scenes to generate the final footage and outputting it in a specified format.

[0379] The present invention is a system for generating high-quality documentary videos based on user viewing preferences and real-time emotion analysis. The system operates by combining multiple hardware and software technologies and is implemented in the following configuration.

[0380] Hardware and software used

[0381] Server: Responsible for information gathering, reliability assessment, data classification, sentiment analysis, story generation, scene composition, and rendering.

[0382] Smart devices (smartphones, smart TVs): Collect data on user input, viewing, and sentiment analysis.

[0383] Communications Network: Collecting and distributing data through internet connections and internal networks.

[0384] Specific software technologies used include:

[0385] Natural Language Processing (NLP) engine: Used to understand and analyze text and assess the reliability of information.

[0386] Machine learning model: Used for user sentiment analysis.

[0387] Template matching algorithm: Used to select the best story template.

[0388] Rendering engine: Used to generate the final documentary footage.

[0389] Cloud storage and databases: Used to store information and manage access.

[0390] Specific examples of the system

[0391] 1. Enter your topic or keywords:

[0392] Users input a theme or keyword, such as "marine pollution," through the interface of their smart device.

[0393] The interface is designed to be intuitive.

[0394] Example: User types "ocean pollution".

[0395] 2. Information Collection:

[0396] The server collects relevant information from the Internet and an internal database via a communication network.

[0397] Using web scraping and API integration, text, images, videos, audio data, etc. are obtained.

[0398] Example: A server collects academic papers, news articles, and visual materials on "ocean pollution."

[0399] 3. Reliability assessment:

[0400] The server uses NLP techniques to evaluate the reliability of the collected information.

[0401] The reliability score is calculated by analyzing the source of the information, the number of citations, and related research data.

[0402] Use social media analytics tools to filter out fake news.

[0403] Example: An NLP engine evaluates the reliability of collected news articles and filters out articles with low scores.

[0404] 4. Information classification:

[0405] The server categorizes the evaluated information.

[0406] Categories include "scientific data," "news articles," "policy reports," and "personal interviews."

[0407] Example: Collected information is divided into categories such as "scientific data" and "personal interviews."

[0408] 5. Viewing Preferences and Sentiment Analysis:

[0409] Users set the target audience, story genre, and the emotion they want to convey.

[0410] The server uses machine learning models to analyze emotions in real time based on the user's facial expressions and input speed collected from the smart device.

[0411] Example: User selects the documentary genre and wants a thrilling story.

[0412] 6. Story generation and scene composition:

[0413] The server generates an optimal story based on the set viewing conditions and the analyzed emotional state.

[0414] A template matching algorithm is used to select an appropriate story template.

[0415] The intention of each cut is set based on the generated story structure, and scenes are composed.

[0416] Example: The server automatically generates scenes with the theme of "tension" based on emotion analysis.

[0417] 7. Video rendering and output:

[0418] The server renders the final documentary footage and outputs it in the format you specify (e.g. MP4, MOV).

[0419] Users can watch on their smart devices or web browsers.

[0420] Example: A completed documentary film is streamed to a smart TV.

[0421] Prompt Sentence Examples

[0422] Please enter a topic (e.g., "Marine Pollution").

[0423] "Assessing reliability of visual data..."

[0424] "Analyzing user sentiment... Anything else you'd like to add?"

[0425] As a result, the present invention provides a system that efficiently generates and provides high-quality documentary footage while taking into consideration the user's emotions.

[0426] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0427] Step 1:

[0428] The user inputs themes and keywords through the interface of the smart device.

[0429] Input: User types in the topic "marine pollution."

[0430] Specific operation: The smart device interface receives input from the user and sends themes and keywords to the server.

[0431] Step 2:

[0432] The server collects information related to themes and keywords from the communication network and an internal database.

[0433] Input: Query regarding "marine pollution."

[0434] Data processing: Use web scraping and API integration to obtain relevant information.

[0435] Output: Collected text, images, video and audio data.

[0436] What it does: The server collects academic papers, news articles, visual materials, etc. related to "marine pollution."

[0437] Step 3:

[0438] The server evaluates the reliability of the collected information using natural language processing (NLP) techniques.

[0439] Input: Collected information.

[0440] Data calculations: Analysis of sources, citations, and related research data. Calculation of reliability scores.

[0441] Output: Information rated as reliable and its score.

[0442] How it works: The NLP engine evaluates the credibility of news articles and papers and assigns them a credibility score.

[0443] Step 4:

[0444] The server classifies the information whose reliability has been evaluated by type.

[0445] Input: Reliability-assessed information.

[0446] Data processing: Using classification algorithms, information is sorted into categories such as "scientific data," "news articles," "policy reports," and "personal interviews."

[0447] Output: Information broken down by category.

[0448] What it does: The server automatically categorizes collected information into appropriate categories based on trustworthiness.

[0449] Step 5:

[0450] Users set the target audience, story genre, and the emotion they want to convey through the interface of their smart device.

[0451] Input: User viewing preferences (e.g., "General Audience," "Documentary," "Intense").

[0452] Output: The set viewing conditions.

[0453] Specific operation: The user operates the smart device and sets up viewing.

[0454] Step 6:

[0455] The server uses a machine learning model to analyze emotions in real time based on the user's facial expressions and input speed collected from the smart device.

[0456] Input: Real-time facial expression data and input speed of the user.

[0457] Data computation: Sentiment analysis with machine learning models.

[0458] Output: User's emotional state (e.g., sense of urgency).

[0459] Specific operation: The server analyzes the user's emotions in real time, and collects and processes the data.

[0460] Step 7:

[0461] The server generates an optimal story based on the set viewing conditions and the analyzed emotional state.

[0462] Input: User's viewing preferences and emotional state, information broken down into categories.

[0463] Data calculation: A template matching algorithm is used to select the most suitable story template.

[0464] Output: The generated story structure.

[0465] What happens: The server uses a template matching algorithm to select an appropriate story template and generate a story.

[0466] Step 8:

[0467] Based on the story structure generated by the server, the intention of each cut is set and scenes are composed.

[0468] Input: The generated story structure.

[0469] Data processing: Automatically place appropriate video clips, narration, and music.

[0470] Output: The composed scene.

[0471] Specific operation: The server arranges video clips and music based on the automatically generated story structure to create scenes.

[0472] Step 9:

[0473] The server renders and outputs the final documentary footage.

[0474] Input: The composed scene.

[0475] Data calculation: Generate the final image using the rendering engine.

[0476] Output: Final documentary footage (e.g. MP4, MOV).

[0477] What it does: The server stitches all the scenes together, renders the final documentary footage, and outputs it in the specified format.

[0478] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0479] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0480] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0481] [Second embodiment]

[0482] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0483] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0484] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0485] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0486] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0487] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0488] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0489] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0490] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0491] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0492] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0493] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0494] This system provides a technical means for efficiently generating documentary footage. A specific embodiment will now be described.

[0495] The user first inputs the documentary's theme or keywords, which are the initial instructions for the system, for example, "marine pollution."

[0496] The server then collects information based on the input topic and keywords, searches the internet and internal databases for relevant information, and collects the necessary text, images, video, and audio data. The collected data is then stored in storage.

[0497] To assess the reliability of the collected information, the server uses natural language processing (NLP) techniques to calculate a credibility score for each piece of information based on its source, number of citations, related research data, etc. It also uses social media analysis tools to filter out fake news and inaccurate information.

[0498] Once the information has been assessed for trustworthiness, it is then sorted by the server into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews." Each piece of information is automatically labeled with the relevant category.

[0499] Next, the user sets the target audience, the genre of the story, the emotion they want to convey, etc. These settings are input through a user interface.

[0500] The server selects a story template based on the input settings and generates the story structure. This process uses a template matching algorithm to select the most suitable story template. The selected template is then incorporated with the collected information to generate the storyboard.

[0501] The next step is to set the intention for each cut and compose the scene. The device receives instructions from the user and selects the appropriate video clip for each cut. At this stage, video clips are dragged and dropped, narration is added, and transition effects are set.

[0502] Finally, the server renders all the scenes and outputs the final documentary video, which is generated in the specified output format (e.g. MP4, MOV) and made available to the user for download or viewing.

[0503] As a concrete example, consider a user creating a documentary on the theme of "marine pollution." First, the user enters "marine pollution," and the server collects the latest research papers, news articles, government reports, and footage from environmental organizations. After assessing the reliability and eliminating fake news, the information is classified into categories such as "scientific data," "news," and "policy responses." The user sets the target audience as general viewers, the story genre as documentary, and the emotion they want to convey as tension. The server composes the story based on a template, and the device places the footage and narration for each cut. Finally, the server integrates all the scenes and outputs the completed documentary footage.

[0504] In this way, the present invention allows for efficient and accurate generation of high quality documentary footage.

[0505] The processing flow will be explained below.

[0506] Step 1:

[0507] The user inputs the documentary's theme or keywords. For example, "marine pollution."

[0508] Step 2:

[0509] The server collects information based on themes and keywords. This includes retrieving relevant text, images, video, and audio data from the internet and internal databases. The server gathers the necessary data through web scraping and API integration.

[0510] Step 3:

[0511] The server evaluates the credibility of the collected information, using natural language processing (NLP) techniques to analyze each piece of information's source, number of citations, and related research data to calculate a credibility score, and also uses social media analysis tools to identify and remove fake news and inaccurate information.

[0512] Step 4:

[0513] The server categorizes the evaluated information into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and automatically labels and organizes the information.

[0514] Step 5:

[0515] The user sets the target audience, the genre of the story, and the emotion they want to convey. These settings are input through a user interface. For example, the target audience is set to general viewers, the genre is set to documentary, and the emotion they want to convey is set to tension.

[0516] Step 6:

[0517] The server selects a story template based on the input settings, uses a template matching algorithm to select the most suitable template, and places the collected information into the storyboard.

[0518] Step 7:

[0519] The device determines the intent of each cut and composes the scene. The user uses drag and drop to place the appropriate video clips for each cut, add narration, and set transition effects. The device then makes the final adjustments to the scene.

[0520] Step 8:

[0521] The server stitches all the scenes together and renders the final documentary video, generating the final video file in the specified output format (e.g. MP4, MOV) for users to download or watch.

[0522] Through the above steps, users can efficiently generate high-quality documentary footage.

[0523] Example 1

[0524] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0525] In today's information society, extracting reliable information from vast amounts of data and efficiently creating video content based on that information is a major challenge. While documentary video production in particular requires accurate and reliable information, it also requires a great deal of time and effort, including information gathering, reliability assessment, story construction, and video editing. Therefore, there is a demand for technological tools that allow users to easily generate high-quality documentary video.

[0526] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0527] In this invention, the server includes means for accepting input of a theme or keyword, means for collecting information related to the theme or keyword from a network and a database, means for evaluating the reliability of the collected information using natural language processing and determining its authenticity, means for categorizing the evaluated information, means for the user to set the target audience, story genre, and emotion to be conveyed, means for generating a story structure based on the set viewing conditions, means for setting the intention of each cut and composing scenes based on the generated story structure, means for rendering and outputting the final video content, means for identifying inaccurate information using a social media analysis tool, and means for selecting a story template optimal for the viewing conditions using a template matching algorithm. This allows the user to efficiently generate high-quality video content based on reliable information while eliminating cumbersome tasks.

[0528] "Theme" is a keyword that the user inputs as the central topic of the documentary video.

[0529] "Keywords" are relevant search terms that the system uses to gather information.

[0530] "Information" refers to any data, including text, images, video, and audio data, collected from networks and databases.

[0531] "Reliability" is the criterion for assessing whether the collected information is accurate and trustworthy.

[0532] "Natural language processing (NLP)" is a technology that analyzes collected text data and evaluates its reliability and categorizes it.

[0533] "Category" refers to the field or type used to classify collected information, such as "scientific data," "news articles," and "policy reports."

[0534] "Target audience" refers to the main audience of the documentary footage, and includes general viewers and experts.

[0535] "Story genre" refers to the type or style of documentary footage, including, for example, documentaries and educational films.

[0536] The "emotions to be conveyed" refer to the emotions and atmosphere that the viewer wants to convey through the video, and include, for example, tension and emotion.

[0537] "Story construction" involves organizing information and planning the flow of the footage based on set viewing conditions.

[0538] A "scene" refers to an individual scene or cut from a documentary film.

[0539] "Rendering" is the process of generating the entire image based on the collected information and the user's settings.

[0540] "Social Media Analytics Tools" are tools used to collect data from social media sites and identify inaccurate information.

[0541] The "template matching algorithm" is an algorithm for selecting the optimal story template based on the set viewing conditions.

[0542] "Video Content" refers to the final video file created based on the collected information.

[0543] This system provides a technical means for efficiently generating documentary footage. A specific embodiment will now be described.

[0544] System configuration

[0545] The system consists of the following main components: a user interface (GUI), a data collection module, a natural language processing (NLP) module, a template matching algorithm, a video editing module, and a rendering module.

[0546] User Interface

[0547] The user first inputs the documentary's theme and keywords through the GUI. This input is the initial instruction for the system, and for example, a word such as "marine pollution" is entered.

[0548] Data Collection Module

[0549] The server collects information based on the input topic or keyword, searches for related information from the network or internal database, and collects the necessary text, image, video, and audio data. For example, it uses Google Custom Search API to collect information from the Internet and pulls the necessary information from the internal database.

[0550] Natural Language Processing (NLP) Module

[0551] The server uses NLP technology to evaluate the reliability of the collected information. Specifically, it calculates a reliability score for each piece of information based on its source, number of citations, and related research data. It also uses social media analysis tools to filter out fake news and inaccurate information. NLP libraries such as spaCy and NLTK are used.

[0552] Categorizing information

[0553] Once the information has been assessed for reliability, it is sorted by the server into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and each piece of information is automatically labeled with the relevant category.

[0554] Viewing condition settings

[0555] Next, the user sets the target audience, the genre of the story, the emotions they want to convey, etc. These settings are again entered through the GUI and sent to the server. For example, possible settings include "general audience," "documentary," and "tension."

[0556] Selecting and configuring a story template

[0557] The server selects a story template based on the user's input settings and generates the story structure using a template matching algorithm. The selected template is then combined with the collected information to generate the storyboard.

[0558] Video editing

[0559] The device receives user instructions and selects the appropriate video clip for each cut. At this stage, video clips are dragged and dropped, narration is added, and transition effects are set. A preview of the video can also be viewed on the GUI.

[0560] Video rendering and output

[0561] Finally, the server renders all the scenes and outputs the final documentary video. The video is generated using a rendering tool such as FFmpeg in the specified format (e.g., MP4, MOV) and made available to users for download or viewing.

[0562] Specific examples

[0563] Consider a case where a user wants to create a documentary on the theme of "marine pollution." The user first enters "marine pollution," and the server uses the Google Custom Search API and an internal database to collect the latest research papers, news articles, government reports, and footage from environmental organizations. The collected information is evaluated for reliability using NLP, and inaccurate information is eliminated. Each piece of information is categorized as "scientific data," "news," "policy response," etc. The user sets viewing conditions such as "general audience," "documentary," and "urgency," and the server composes a story based on the template. The device arranges the footage and narration for each cut, and finally, the server integrates all the scenes and outputs the completed documentary video.

[0564] Prompt Sentence Examples

[0565] As a prompt sentence, provide the following sentence as input to the generative AI model:

[0566] Theme: Marine pollution

[0567] Keywords: plastics, climate change, marine life

[0568] Target audience: General audience

[0569] Story genre: Documentary

[0570] Emotions to convey: Urgency

[0571] Sources: networks, databases, news articles, reports, research papers

[0572] In this way, the present invention allows users to effectively and accurately generate high quality documentary footage.

[0573] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0574] Step 1: User enters topic and keywords

[0575] The user logs into the system's user interface (GUI) and is taken to a screen where they can enter a theme and keywords. The user enters the keywords "marine pollution" and "plastic" in the text field and clicks the "Search" button.

[0576] Input: topic (marine pollution) and keyword (plastic)

[0577] Output: The entered theme and keywords are sent to the server.

[0578] Step 2: Server collects information

[0579] The server receives the input topic and keywords and collects related information from the network and internal database. The server uses the Google Custom Search API to perform an internet search for the query "ocean pollution plastic" and scrapes useful information (text, images, video, audio data). It also queries the internal database to pull out related information.

[0580] Input: topic (marine pollution) and keyword (plastic)

[0581] Output: A list of collected text, image, video, and audio data

[0582] Step 3: The server evaluates trustworthiness using natural language processing (NLP)

[0583] The server applies NLP techniques to the collected information, specifically using the spaCy and NLTK libraries to calculate a credibility score for each piece of information. The server also evaluates the source, number of citations, and related research data for each piece of information, and uses social media analysis tools to filter out inaccurate information.

[0584] Input: A list of collected text, image, video, and audio data

[0585] Output: List of information with confidence scores

[0586] Step 4: The server categorizes the information

[0587] The server categorizes the information whose reliability has been evaluated. It analyzes the content of the text data and automatically categorizes it into categories such as "scientific data," "news articles," "policy reports," and "personal interviews." The server updates and saves this categorization information in a database.

[0588] Input: List of information with confidence scores

[0589] Output: A list of information sorted by category

[0590] Step 5: User sets target audience, genre, and emotion

[0591] The user again uses the GUI to set the target audience, the genre of the story, and the emotion they want to convey. Users select settings such as "general audience," "documentary," or "urgency" from drop-down menus and click the "Set" button.

[0592] Input: audience, story genre, emotion you want to convey

[0593] Output: Configuration information sent to the server

[0594] Step 6: The server selects a story template and generates a story structure.

[0595] The server runs a template matching algorithm based on the viewing conditions you set and selects the most suitable story template. The collected information is automatically incorporated into the selected template to generate a storyboard.

[0596] Input: audience, story genre, emotion to convey, list of information categorized

[0597] Output: Storyboard

[0598] Step 7: Select and edit video clips on your device

[0599] The user uses the device to select and edit video clips based on the generated storyboard. They drag and drop the appropriate video clips for each scene, add narration, and set transition effects. Editing proceeds while checking the video preview on the GUI.

[0600] Input: Storyboard, video clips

[0601] Output: Edited video sequence

[0602] Step 8: The server renders the entire scene and outputs the final video.

[0603] The server uses a rendering tool such as FFmpeg to combine all the scenes and render the video in the specified output format (e.g. MP4, MOV). The final video file is provided to the user, and a download link is displayed on the GUI.

[0604] Input: Edited video sequence

[0605] Output: Final video file (e.g. MP4, MOV)

[0606] Through the above steps, the user can efficiently generate high-quality documentary footage.

[0607] (Application example 1)

[0608] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0609] In today's world, documentary video production requires a high level of skill and time. Gathering accurate information tailored to the theme, assessing its reliability, and structuring a story that meets the needs of viewers are particularly challenging. Furthermore, amid a proliferation of fake information, there is a demand for video production that uses accurate and reliable information. Rendering and outputting the final video work is also a laborious process that requires increased efficiency. To address these challenges, there is an urgent need to provide a system that allows users to easily and efficiently generate and distribute high-quality documentary video.

[0610] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0611] In this invention, the server includes means for accepting input of a theme or keyword, means for collecting information related to the theme or keyword from a communication network and an internal database, means for evaluating the reliability of the collected information and determining its veracity, means for classifying the evaluated information by category, means for the user to set the target audience, story genre, and desired emotion, means for generating a story structure based on the set viewing conditions, means for setting the intention of each cut and composing scenes based on the generated story structure, means for rendering and outputting the final video work, and means for providing a user interface for the user to input a theme and automatically generate and distribute documentary video based on the collected information, thereby enabling users to easily and efficiently generate and distribute high-quality documentary video.

[0612] "Theme" is a concept that allows the user to indicate the central subject or interest of the documentary footage.

[0613] "Keywords" are indicators that contain important words or terms related to a topic and are used to narrow down information.

[0614] A "communications network" is a digital network for sending and receiving information, including the Internet.

[0615] An "internal database" is a source of information held within the system and is a storage device for storing and managing a portion of the collected data.

[0616] "Credibility" is a measure of the accuracy and veracity of collected information, and is calculated using NLP technology.

[0617] "Veracity" is the quality that determines whether information is based on actual, verified facts.

[0618] "User" refers to an individual or organization that performs operations to create documentary footage.

[0619] "Target audience" refers to the audience or target group that intends to watch the documentary footage.

[0620] "Story genre" is a classification that indicates the format and style of documentary footage.

[0621] "Emotions you want to convey" refers to the emotions and atmosphere you want viewers to feel through the documentary footage.

[0622] The "viewing conditions" are parameters related to the viewer set by the user, such as the viewing target, the genre of the story, and the emotion to be conveyed.

[0623] "Story structure" is the framework for designing the overall flow and development of a documentary film.

[0624] "The intention of each cut" refers to the purpose and message of each scene or cut in the video.

[0625] "Scenes" are the individual scenes or cuts that make up documentary footage.

[0626] "Video work" is a general term for the final documentary video.

[0627] "Rendering" is the process of converting an edited video clip into an output format as a single file.

[0628] "Output" refers to the act of saving or distributing the final video work in a specified format.

[0629] "User interface" refers to the screen and operating means through which a user interacts with and operates a system.

[0630] This invention relates to a system that efficiently generates and distributes documentary footage by allowing users to input themes and keywords. The system is configured as follows:

[0631] First, users input the documentary's theme and keywords through the smartphone application interface. For example, they can specify a theme such as "marine pollution." To accept this input, the system uses libraries such as React Native as its front end.

[0632] The server receives the input topic and keywords and collects related information from communication networks and internal databases. In this step, a crawler is used to collect text, images, video, and audio data from academic papers, news sites, government sites, and other sources on the Internet. The collected data is then stored in storage. The software used is Python and technologies such as Beautiful Soup and Scrapy.

[0633] Next, the reliability of the collected information is evaluated. TensorFlow or PyTorch is used to calculate a reliability score for the information using natural language processing (NLP) techniques. Furthermore, fake news and inaccurate information is filtered out using social media analysis tools. This analysis uses techniques such as sentiment analysis and named entity recognition (NER).

[0634] The evaluated information is then sorted into categories on the server, such as "scientific data," "news articles," "policy responses," and "personal interviews." The information is sorted into categories using labeling technology based on machine learning models.

[0635] Next, the user sets the target audience, story genre, and desired emotion on the application. These settings are necessary for the server to generate the optimal story structure. A template matching algorithm is used to select a story template and automatically generate the storyboard.

[0636] Based on the generated storyboard, the intention of each cut is set and the scene is composed. At this stage, the user inserts video clips, adds narration, and sets transition effects through the interface. The front-end editing function uses HTML5, CSS3, and JavaScript.

[0637] FFmpeg is used to render the final video and save it in the output format, which is then made available to users for download.

[0638] As a concrete example, consider a case where a user is creating a documentary on the theme of "marine pollution." When the user enters "marine pollution," the server collects the latest scientific data, news articles, and policy response information from the Internet. It evaluates the reliability of the collected information and categorizes it into categories such as "scientific data," "news," and "policy response." The user sets the target audience as general viewers, the story genre as documentary, and the emotion they want to convey as tension. The server composes the story based on a template and edits each cut through the user interface. Finally, the server integrates all the scenes and outputs the completed documentary video in MP4 format.

[0639] Examples of prompts include:

[0640] "A user wants to create a documentary about marine pollution. They need to gather the latest research, news, policy reports, etc., create a storyboard based on that information, and finally output it in MP4 format."

[0641] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0642] Step 1:

[0643] Users enter the documentary's theme and keywords into the smartphone application.

[0644] Input: Theme or keyword (e.g., marine pollution)

[0645] Output: Themes and keywords are sent to the server

[0646] Step 2:

[0647] The server collects related information from the communication network and internal database based on the input topic and keywords.

[0648] Input: Theme or keyword

[0649] Data processing: Using crawlers to collect information from academic papers, news sites, government sites, etc. on the Internet and store it in storage.

[0650] Output: Collected text, images, video, and audio data

[0651] Step 3:

[0652] The server uses NLP techniques to assess the reliability of the collected information.

[0653] Input: Collected information

[0654] Data Computation: Use TensorFlow or PyTorch to calculate the credibility score of information. Use social media analysis tools to filter out fake news and inaccurate information.

[0655] Output: reliable information

[0656] Step 4:

[0657] The server categorizes the information whose reliability has been evaluated.

[0658] Input: Information that has undergone a reliability assessment

[0659] Data processing: Using machine learning models, information is labeled into categories such as "scientific data," "news articles," "policy responses," and "personal interviews."

[0660] Output: Information sorted by category

[0661] Step 5:

[0662] Users set their viewing target, story genre, and the emotion they want to convey on the application.

[0663] Input: Target audience, story genre, and desired emotion (e.g., general audience, documentary, suspense)

[0664] Output: The configured viewing conditions are sent to the server.

[0665] Step 6:

[0666] The server selects the most suitable story template based on the set viewing conditions and generates a storyboard.

[0667] Input: Viewing conditions, information categorized by category

[0668] Data calculation: Using a template matching algorithm, a storyboard is automatically generated by selecting the story template that best suits the viewing conditions.

[0669] Output: Generated storyboard

[0670] Step 7:

[0671] The user sets the intention for each cut based on the generated storyboard and composes the scene.

[0672] Input: Generated storyboard

[0673] Specific operations: The user inserts a video clip, adds narration, and sets transition effects in the application interface.

[0674] Output: Edited storyboard

[0675] Step 8:

[0676] The server renders the final video production and outputs it in the specified format.

[0677] Input: Edited storyboard

[0678] Specific operation: All scenes are integrated using FFmpeg and the final video is rendered in MP4 format or similar.

[0679] Output: The finished documentary footage is made available to users in a downloadable format

[0680] As an example, the following prompt sentence is used:

[0681] "A user wants to create a documentary about marine pollution. They need to gather the latest research, news, policy reports, etc., create a storyboard based on that information, and finally output it in MP4 format."

[0682] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0683] The present invention is a system that collects related information based on specific themes and keywords, evaluates and classifies reliability, and generates high-quality documentary footage by recognizing the user's viewing preferences and emotions.

[0684] System Configuration

[0685] 1. How to input themes and keywords

[0686] The user inputs the documentary's theme or keywords. For example, "marine pollution." The user interface (UI) is designed to be intuitive.

[0687] 2. Information gathering methods

[0688] The server collects relevant information from the internet and internal databases using web scraping and API integration to obtain text, images, video, and audio data.

[0689] 3. Reliability assessment methods

[0690] The server uses natural language processing (NLP) techniques to assess the reliability of collected information, analyzing its source, number of citations, and related research data to calculate a credibility score. It also uses social media analysis tools to identify and filter out fake news.

[0691] 4. Information classification means

[0692] The server categorizes the evaluated information into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and automatically labels each piece of information with the appropriate category.

[0693] 5. Viewing setting input method

[0694] The user sets the target audience (e.g., general audience), the genre of the story (e.g., documentary), and the emotion they want to convey (e.g., tension). These settings are made through a user interface.

[0695] 6. Emotion Engine

[0696] The emotion engine analyzes the user's emotions in real time as they type, and reflects this in the story structure and scene settings. Emotion analysis is performed based on the user's facial expressions, tone of voice, and typing speed.

[0697] 7. Story Generation Methods

[0698] The server selects the optimal story template based on the viewing preferences and the analysis results of the emotion engine, and generates a story structure. Using a template matching algorithm, the collected information is arranged on the storyboard.

[0699] 8. Scene Composition Methods

[0700] The device allows the user to compose scenes by setting the intention for each cut. Using drag and drop, the user places the appropriate video clips for each cut, adds narration, and sets transition effects. Based on the emotional state recognized by the emotion engine, appropriate narration and music are automatically selected and placed.

[0701] 9. Video rendering and output methods

[0702] The server stitches all the scenes together and renders the final documentary video. The video file is generated in the specified output format (e.g. MP4, MOV) and made available to the user for download or viewing.

[0703] Specific examples

[0704] If a user wants to create a documentary on the theme of "marine pollution," they first input the theme. The server collects related information and evaluates and classifies it for reliability. The user then selects a general audience, a documentary genre, and a desire to convey a sense of tension. The emotion engine recognizes the user's sense of tension as they input, and reflects this in the story structure and scene settings. Appropriate video clips, narration, and music are selected, and finally, the server renders all the scenes, outputting the finished documentary video.

[0705] In this way, the present invention provides a system that efficiently generates high-quality documentary footage while taking into account the user's emotions.

[0706] The processing flow will be explained below.

[0707] Step 1:

[0708] The user inputs the documentary's theme or keywords. For example, "marine pollution."

[0709] Step 2:

[0710] The server collects information based on themes and keywords. The information is collected using web scraping and API integration to obtain text, images, video, and audio data from the internet and internal databases.

[0711] Step 3:

[0712] The server evaluates the credibility of the collected information, using natural language processing (NLP) techniques to analyze the information's source, number of citations, and related research data to calculate a credibility score, and also uses social media analysis tools to identify and remove fake news and inaccurate information.

[0713] Step 4:

[0714] The server categorizes the evaluated information into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and each piece of information is automatically labeled with the appropriate category.

[0715] Step 5:

[0716] The user sets the target audience, the genre of the story, and the emotion they want to convey. These settings are entered through a highly visible user interface. For example, the target audience can be set to general viewers, the genre to documentary, and the emotion they want to convey to be tension.

[0717] Step 6:

[0718] The emotion engine analyzes the user's emotions in real time, analyzing facial expressions, tone of voice, and input speed when entering settings to recognize the user's current emotional state.

[0719] Step 7:

[0720] The server selects the optimal story template based on the user's viewing preferences and the analysis results of the emotion engine. Using a template matching algorithm, it selects the template that best suits the preferences and analysis data, and places the collected information on the storyboard.

[0721] Step 8:

[0722] The device sets the intention for each cut and composes the scene. The user uses drag and drop to place the appropriate video clips for each cut, add narration, and set transition effects. Appropriate narration and music are also automatically selected and placed based on the emotional state recognized by the emotion engine.

[0723] Step 9:

[0724] The server stitches all the scenes together and renders the final documentary video. The video file is generated in the specified output format (e.g. MP4, MOV) and made available to the user for download or viewing.

[0725] By following these steps, users can efficiently generate high-quality documentary footage that reflects emotions.

[0726] Example 2

[0727] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0728] In conventional documentary video production systems, the processes of information collection, credibility assessment, information classification, viewing settings, emotion analysis, story generation, scene composition, and video rendering are all performed manually, which is extremely time-consuming and labor-intensive. Furthermore, they lack the ability to analyze user emotions in real time and reflect them in video production, resulting in a lack of emotional impact on viewers. Furthermore, with the rise of fake news, unreliable information is often mixed in, making it difficult to produce high-quality documentary videos based on accurate information.

[0729] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for accepting input of themes and keywords, means for collecting information from a network and an internal database, means for evaluating the reliability of the collected information and determining its authenticity, means for identifying fake information using social media analysis technology, means for categorizing the evaluated information, means for the user to set the viewing target, story genre, and emotion to be conveyed, means for analyzing the user's emotions in real time using an emotion analysis means and reflecting the results in scene settings, means for selecting a story template optimal for the viewing conditions using a template matching algorithm, means for setting the intention of each scene based on the generated story structure and composing scenes, and means for synthesizing and outputting the final documentary video. This enables automatic collection and reliability evaluation of information, appropriate information classification, real-time analysis of user emotions, and automatic selection of the optimal story template, thereby enabling efficient generation of high-quality documentary video with emotional impact.

[0730] The "means for accepting input of themes and keywords" is a means for providing an interface that allows a user to input themes and keywords necessary for producing a documentary video.

[0731] "Means for collecting information from networks and internal databases" refers to means for automatically collecting information related to the entered topic or keyword via the Internet or internal databases.

[0732] "Means for assessing the reliability of collected information and determining its authenticity" refers to means for assessing the reliability of collected information using natural language processing technology and source analysis, and determining the authenticity of the information.

[0733] "Methods for identifying fake information using social media analysis technology" refers to methods that use technology to analyze data on social media and identify misinformation and fake news.

[0734] A "means of categorizing evaluated information" is a means of classifying information that has been collected and evaluated for reliability into specific categories, such as "scientific data" or "news articles."

[0735] "A means for users to set the target audience, story genre, and emotions they want to convey" is a means for providing an interface that allows users to set the audience, genre, and emotions they want to convey for the documentary footage.

[0736] "Means for analyzing a user's emotions in real time using emotion analysis means and reflecting the emotions in scene settings" refers to means for analyzing a user's facial expressions and tone of voice, grasping the user's emotional state in real time, and reflecting the results in scene settings.

[0737] "Means for selecting a story template optimal for viewing conditions using a template matching algorithm" refers to means for executing an algorithm that selects the optimal template from among pre-prepared story templates based on the user's viewing settings and emotion analysis results.

[0738] "Means for setting the intention of each scene based on the generated story structure and composing scenes" refers to means for setting the purpose and intention of each scene based on the generated story and arranging corresponding images and narration.

[0739] The "means for compositing and outputting the final documentary footage" refers to a means for combining all the scenes into one, generating the final documentary footage, and outputting it in a specified format.

[0740] This invention is a system that collects related information based on specific themes and keywords, evaluates and classifies its reliability, recognizes the user's viewing preferences and emotions, and automatically generates high-quality documentary footage.

[0741] System Configuration

[0742] The system has the following main components:

[0743] 1. How to input themes and keywords

[0744] Users input the documentary's theme and keywords through a user interface (UI), which is designed for intuitive operation.

[0745] 2. Information gathering methods

[0746] The server collects relevant information from the internet and internal databases using web scraping tools (e.g., BeautifulSoup) and API integration (e.g., Twitter API), which allows it to obtain text, images, video, and audio data.

[0747] 3. Reliability assessment methods

[0748] The server evaluates the reliability of the collected information using natural language processing (NLP) techniques (e.g., BERT, GPT-3), and also identifies and eliminates fake information using social media analysis techniques.

[0749] 4. Information classification means

[0750] The server categorizes the evaluated information using a clustering algorithm (e.g., K-means) to automatically categorize it into categories such as "scientific data," "news articles," "policy reports," and "personal interviews."

[0751] 5. Viewing setting input method

[0752] Users can set their target audience, the genre of the story, and the emotion they want to convey through a user interface.

[0753] 6. Emotion analysis method

[0754] The server is equipped with an emotion analysis engine that analyzes the user's emotions in real time using software that analyzes facial expressions (e.g., OpenCV) and voice tone (e.g., Deepgram), and also combines the user's input speed for analysis.

[0755] 7. Story Generation Methods

[0756] The server runs a template matching algorithm to select the most appropriate story template based on viewing preferences and sentiment analysis results, and places the collected information into the appropriate storyboard.

[0757] 8. Scene Composition Methods

[0758] The user uses the drag-and-drop function on the device to place appropriate video clips for each cut on a storyboard, and then sets narration and transition effects. Appropriate narration and music are automatically selected based on the emotional state recognized by the emotion analysis means.

[0759] 9. Video rendering and output methods

[0760] The server combines all the scenes to generate the final documentary video and renders it as a video file in the specified format, such as MP4 or MOV.

[0761] Specific examples

[0762] Consider a scenario where a user is creating a documentary on the theme of "marine pollution." First, the user inputs the theme. The server collects related information, evaluates its reliability, and classifies it. Next, the user sets the target audience as "general audience," the genre as "documentary," and the emotion they want to convey as "tension." The emotion analysis engine recognizes the user's emotions in real time as they are input, and reflects them in the story structure and scene settings. Appropriate video clips, narration, and music are selected, and the server renders all scenes to output the completed documentary video.

[0763] Prompt Sentence Examples

[0764] "Collect information to create a documentary film on the following topic, assess its reliability, classify it, generate a story, and output the final film. The user settings are for general audiences, the genre is documentary, and the emotion you want to convey is a sense of urgency. The topic is 'Marine Pollution.'"

[0765] This system makes it possible to efficiently generate high-quality documentary footage while taking into account the user's emotions.

[0766] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0767] Step 1:

[0768] Accepts input of themes and keywords.

[0769] The user inputs themes and keywords through a user interface (UI).

[0770] Specific behavior: The user enters "Marine Pollution" in the topic field and clicks the submit button.

[0771] Input: Theme or keyword (e.g. "Marine pollution")

[0772] Output: Themes and keywords entered

[0773] Step 2:

[0774] Automatic Collection of Information.

[0775] The server collects information from the Internet and internal databases based on the entered themes and keywords.

[0776] Specific operation: The server uses a web scraping tool (e.g., BeautifulSoup) to retrieve articles from Google search results and uses an API (e.g., Twitter API) to collect related tweets.

[0777] Input: Theme or keyword (e.g. "Marine pollution")

[0778] Output: Collected information (text, images, video, audio data)

[0779] Step 3:

[0780] Evaluating the reliability of information.

[0781] The server evaluates the reliability of the collected information using natural language processing (NLP) techniques.

[0782] How it works: The server uses NLP models (e.g., BERT, GPT-3) to calculate a credibility score for each article, and also uses social media analysis techniques to identify and remove misinformation.

[0783] Input: Collected information

[0784] Output: Information with assessed reliability (with score)

[0785] Step 4:

[0786] Classification of information.

[0787] The server categorizes the evaluated information.

[0788] What it does: The server uses a clustering algorithm (e.g., K-means) to classify information into categories such as "scientific data," "news articles," "policy reports," and "personal interviews."

[0789] Input: Information with reliability assessment

[0790] Output: Categorically labeled information

[0791] Step 5:

[0792] Enter viewing settings.

[0793] The user inputs the target audience, the genre of the story, and the emotion they want to convey.

[0794] Specific operation: The user selects the target audience as "general audience," the genre as "documentary," and the emotion to be conveyed as "tension."

[0795] Input: target audience, story genre, emotion you want to convey

[0796] Output: Viewing Settings

[0797] Step 6:

[0798] Performing sentiment analysis.

[0799] The server analyzes the user's emotions in real time.

[0800] How it works: The server uses facial expression analysis software (e.g., OpenCV) and voice analysis software (e.g., Deepgram) to detect the user's emotions. The user's typing speed is also used for analysis.

[0801] Input: User's facial expression data, tone of voice, typing speed

[0802] Output: Real-time analyzed emotion data

[0803] Step 7:

[0804] Story generation.

[0805] The server selects a story template based on viewing preferences and emotion analysis results, and generates a story structure.

[0806] Specific operation: The server uses a template matching algorithm to select an appropriate story template and places the collected information on the storyboard.

[0807] Input: Viewing settings, sentiment analysis results, and information categorized by category

[0808] Output: Generated story structure

[0809] Step 8:

[0810] Scene placement and adjustment.

[0811] The terminal allows the user to set the intention for each cut and compose a scene.

[0812] How it works: Users use the drag-and-drop function to place video clips on the storyboard and set narration and transition effects. Appropriate narration and music are automatically selected based on emotional data.

[0813] Input: Generated story structure, video clips, narration, transition effects

[0814] Output: Composed Scene

[0815] Step 9:

[0816] Video rendering and output.

[0817] The server stitches all the scenes together and renders the final documentary footage.

[0818] What it does: The server uses video editing software (e.g., FFmpeg) to combine and render all the scenes into a single video file in MP4 or MOV format.

[0819] Input: Composed Scene

[0820] Output: Final documentary footage file

[0821] (Application example 2)

[0822] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0823] In today's world, users are seeking to generate and view high-quality documentary videos based on their own interests. However, conventional methods have difficulty collecting reliable information and providing it in an optimal format based on the viewer's emotions and preferences. Furthermore, there is no system that can analyze the viewer's emotional state in real time and dynamically change the video content accordingly, which hinders user satisfaction.

[0824] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0825] In this invention, the server includes means for accepting input of themes and keywords, means for collecting information related to the themes and keywords from a communication network and an internal database, means for evaluating the reliability of the collected information and determining its authenticity, means for classifying the evaluated information by type, means for the user to set the viewing target, story genre, and emotion to be conveyed, means for analyzing the user's emotional state in real time, means for generating a story structure based on the set viewing conditions and the analyzed emotional state, means for setting the intention of each cut and composing scenes based on the generated story structure, and means for rendering and outputting the final documentary video. This makes it possible to generate and provide high-quality, reliable documentary video in real time based on the viewing conditions and emotional state set by the user.

[0826] The "means for accepting input of themes and keywords" is an interface that allows a user to input subjects of interest or specific words into the system.

[0827] "Means for collecting information related to the theme or keyword from communication networks and internal databases" refers to a function for collecting information from the Internet and obtaining related information from internally stored databases.

[0828] "Means for assessing the reliability of collected information and determining its authenticity" refers to techniques for analyzing the reliability of the data obtained and determining whether the data is accurate.

[0829] The "means for classifying evaluated information by type" refers to a method for classifying information whose reliability has been evaluated into different categories based on its content.

[0830] "A means for users to set the target audience, story genre, and emotions they want to convey" is an interface that allows users to set the target audience, video genre, and emotions they want to convey through the video.

[0831] The "means for analyzing the user's emotional state in real time" is a technology for detecting and analyzing the user's current emotions in real time.

[0832] The "means for generating a story structure based on the set viewing conditions and the analyzed emotional state" is a method for automatically constructing an optimal story based on the viewing conditions set by the user and the analyzed emotions.

[0833] "Means for setting the intention of each cut based on the generated story structure and composing scenes" refers to a technique for setting the content and purpose of each scene based on the generated story structure and then visualizing it.

[0834] "Means for rendering and outputting the final documentary footage" refers to the function for integrating all scenes to generate the final footage and outputting it in a specified format.

[0835] The present invention is a system for generating high-quality documentary videos based on user viewing preferences and real-time emotion analysis. The system operates by combining multiple hardware and software technologies and is implemented in the following configuration.

[0836] Hardware and software used

[0837] Server: Responsible for information gathering, reliability assessment, data classification, sentiment analysis, story generation, scene composition, and rendering.

[0838] Smart devices (smartphones, smart TVs): Collect data on user input, viewing, and sentiment analysis.

[0839] Communications Network: Collecting and distributing data through internet connections and internal networks.

[0840] Specific software technologies used include:

[0841] Natural Language Processing (NLP) engine: Used to understand and analyze text and assess the reliability of information.

[0842] Machine learning model: Used for user sentiment analysis.

[0843] Template matching algorithm: Used to select the best story template.

[0844] Rendering engine: Used to generate the final documentary footage.

[0845] Cloud storage and databases: Used to store information and manage access.

[0846] Specific examples of the system

[0847] 1. Enter your topic or keywords:

[0848] Users input a theme or keyword, such as "marine pollution," through the interface of their smart device.

[0849] The interface is designed to be intuitive.

[0850] Example: User types "ocean pollution".

[0851] 2. Information Collection:

[0852] The server collects relevant information from the Internet and an internal database via a communication network.

[0853] Using web scraping and API integration, text, images, videos, audio data, etc. are obtained.

[0854] Example: A server collects academic papers, news articles, and visual materials on "ocean pollution."

[0855] 3. Reliability assessment:

[0856] The server uses NLP techniques to evaluate the reliability of the collected information.

[0857] The reliability score is calculated by analyzing the source of the information, the number of citations, and related research data.

[0858] Use social media analytics tools to filter out fake news.

[0859] Example: An NLP engine evaluates the reliability of collected news articles and filters out articles with low scores.

[0860] 4. Information classification:

[0861] The server categorizes the evaluated information.

[0862] Categories include "scientific data," "news articles," "policy reports," and "personal interviews."

[0863] Example: Collected information is divided into categories such as "scientific data" and "personal interviews."

[0864] 5. Viewing Preferences and Sentiment Analysis:

[0865] Users set the target audience, story genre, and the emotion they want to convey.

[0866] The server uses machine learning models to analyze emotions in real time based on the user's facial expressions and input speed collected from the smart device.

[0867] Example: User selects the documentary genre and wants a thrilling story.

[0868] 6. Story generation and scene composition:

[0869] The server generates an optimal story based on the set viewing conditions and the analyzed emotional state.

[0870] A template matching algorithm is used to select an appropriate story template.

[0871] The intention of each cut is set based on the generated story structure, and scenes are composed.

[0872] Example: The server automatically generates scenes with the theme of "tension" based on emotion analysis.

[0873] 7. Video rendering and output:

[0874] The server renders the final documentary footage and outputs it in the format you specify (e.g. MP4, MOV).

[0875] Users can watch on their smart devices or web browsers.

[0876] Example: A completed documentary film is streamed to a smart TV.

[0877] Prompt Sentence Examples

[0878] Please enter a topic (e.g., "Marine Pollution").

[0879] "Assessing reliability of visual data..."

[0880] "Analyzing user sentiment... Anything else you'd like to add?"

[0881] As a result, the present invention provides a system that efficiently generates and provides high-quality documentary footage while taking into consideration the user's emotions.

[0882] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0883] Step 1:

[0884] The user inputs themes and keywords through the interface of the smart device.

[0885] Input: User types in the topic "marine pollution."

[0886] Specific operation: The smart device interface receives input from the user and sends themes and keywords to the server.

[0887] Step 2:

[0888] The server collects information related to themes and keywords from the communication network and an internal database.

[0889] Input: Query regarding "marine pollution."

[0890] Data processing: Use web scraping and API integration to obtain relevant information.

[0891] Output: Collected text, images, video and audio data.

[0892] What it does: The server collects academic papers, news articles, visual materials, etc. related to "marine pollution."

[0893] Step 3:

[0894] The server evaluates the reliability of the collected information using natural language processing (NLP) techniques.

[0895] Input: Collected information.

[0896] Data calculations: Analysis of sources, citations, and related research data. Calculation of reliability scores.

[0897] Output: Information rated as reliable and its score.

[0898] How it works: The NLP engine evaluates the credibility of news articles and papers and assigns them a credibility score.

[0899] Step 4:

[0900] The server classifies the information whose reliability has been evaluated by type.

[0901] Input: Reliability-assessed information.

[0902] Data processing: Using classification algorithms, information is sorted into categories such as "scientific data," "news articles," "policy reports," and "personal interviews."

[0903] Output: Information broken down by category.

[0904] What it does: The server automatically categorizes collected information into appropriate categories based on trustworthiness.

[0905] Step 5:

[0906] Users set the target audience, story genre, and the emotion they want to convey through the interface of their smart device.

[0907] Input: User viewing preferences (e.g., "General Audience," "Documentary," "Intense").

[0908] Output: The set viewing conditions.

[0909] Specific operation: The user operates the smart device and sets up viewing.

[0910] Step 6:

[0911] The server uses a machine learning model to analyze emotions in real time based on the user's facial expressions and input speed collected from the smart device.

[0912] Input: Real-time facial expression data and input speed of the user.

[0913] Data computation: Sentiment analysis with machine learning models.

[0914] Output: User's emotional state (e.g., sense of urgency).

[0915] Specific operation: The server analyzes the user's emotions in real time, and collects and processes the data.

[0916] Step 7:

[0917] The server generates an optimal story based on the set viewing conditions and the analyzed emotional state.

[0918] Input: User's viewing preferences and emotional state, information broken down into categories.

[0919] Data calculation: A template matching algorithm is used to select the most suitable story template.

[0920] Output: The generated story structure.

[0921] What happens: The server uses a template matching algorithm to select an appropriate story template and generate a story.

[0922] Step 8:

[0923] Based on the story structure generated by the server, the intention of each cut is set and scenes are composed.

[0924] Input: The generated story structure.

[0925] Data processing: Automatically place appropriate video clips, narration, and music.

[0926] Output: The composed scene.

[0927] Specific operation: The server arranges video clips and music based on the automatically generated story structure to create scenes.

[0928] Step 9:

[0929] The server renders and outputs the final documentary footage.

[0930] Input: The composed scene.

[0931] Data calculation: Generate the final image using the rendering engine.

[0932] Output: Final documentary footage (e.g. MP4, MOV).

[0933] What it does: The server stitches all the scenes together, renders the final documentary footage, and outputs it in the specified format.

[0934] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0935] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0936] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0937] [Third embodiment]

[0938] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0939] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0940] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0941] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0942] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0943] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0944] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0945] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0946] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0947] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0948] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0949] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0950] This system provides a technical means for efficiently generating documentary footage. A specific embodiment will now be described.

[0951] The user first inputs the documentary's theme or keywords, which are the initial instructions for the system, for example, "marine pollution."

[0952] The server then collects information based on the input topic and keywords, searches the internet and internal databases for relevant information, and collects the necessary text, images, video, and audio data. The collected data is then stored in storage.

[0953] To assess the reliability of the collected information, the server uses natural language processing (NLP) techniques to calculate a credibility score for each piece of information based on its source, number of citations, related research data, etc. It also uses social media analysis tools to filter out fake news and inaccurate information.

[0954] Once the information has been assessed for trustworthiness, it is then sorted by the server into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews." Each piece of information is automatically labeled with the relevant category.

[0955] Next, the user sets the target audience, the genre of the story, the emotion they want to convey, etc. These settings are input through a user interface.

[0956] The server selects a story template based on the input settings and generates the story structure. This process uses a template matching algorithm to select the most suitable story template. The selected template is then incorporated with the collected information to generate the storyboard.

[0957] The next step is to set the intention for each cut and compose the scene. The device receives instructions from the user and selects the appropriate video clip for each cut. At this stage, video clips are dragged and dropped, narration is added, and transition effects are set.

[0958] Finally, the server renders all the scenes and outputs the final documentary video, which is generated in the specified output format (e.g. MP4, MOV) and made available to the user for download or viewing.

[0959] As a concrete example, consider a user creating a documentary on the theme of "marine pollution." First, the user enters "marine pollution," and the server collects the latest research papers, news articles, government reports, and footage from environmental organizations. After assessing the reliability and eliminating fake news, the information is classified into categories such as "scientific data," "news," and "policy responses." The user sets the target audience as general viewers, the story genre as documentary, and the emotion they want to convey as tension. The server composes the story based on a template, and the device places the footage and narration for each cut. Finally, the server integrates all the scenes and outputs the completed documentary footage.

[0960] In this way, the present invention allows for efficient and accurate generation of high quality documentary footage.

[0961] The processing flow will be explained below.

[0962] Step 1:

[0963] The user inputs the documentary's theme or keywords. For example, "marine pollution."

[0964] Step 2:

[0965] The server collects information based on themes and keywords. This includes retrieving relevant text, images, video, and audio data from the internet and internal databases. The server gathers the necessary data through web scraping and API integration.

[0966] Step 3:

[0967] The server evaluates the credibility of the collected information, using natural language processing (NLP) techniques to analyze each piece of information's source, number of citations, and related research data to calculate a credibility score, and also uses social media analysis tools to identify and remove fake news and inaccurate information.

[0968] Step 4:

[0969] The server categorizes the evaluated information into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and automatically labels and organizes the information.

[0970] Step 5:

[0971] The user sets the target audience, the genre of the story, and the emotion they want to convey. These settings are input through a user interface. For example, the target audience is set to general viewers, the genre is set to documentary, and the emotion they want to convey is set to tension.

[0972] Step 6:

[0973] The server selects a story template based on the input settings, uses a template matching algorithm to select the most suitable template, and places the collected information into the storyboard.

[0974] Step 7:

[0975] The device determines the intent of each cut and composes the scene. The user uses drag and drop to place the appropriate video clips for each cut, add narration, and set transition effects. The device then makes the final adjustments to the scene.

[0976] Step 8:

[0977] The server stitches all the scenes together and renders the final documentary video, generating the final video file in the specified output format (e.g. MP4, MOV) for users to download or watch.

[0978] Through the above steps, users can efficiently generate high-quality documentary footage.

[0979] Example 1

[0980] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0981] In today's information society, extracting reliable information from vast amounts of data and efficiently creating video content based on that information is a major challenge. While documentary video production in particular requires accurate and reliable information, it also requires a great deal of time and effort, including information gathering, reliability assessment, story construction, and video editing. Therefore, there is a demand for technological tools that allow users to easily generate high-quality documentary video.

[0982] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0983] In this invention, the server includes means for accepting input of a theme or keyword, means for collecting information related to the theme or keyword from a network and a database, means for evaluating the reliability of the collected information using natural language processing and determining its authenticity, means for categorizing the evaluated information, means for the user to set the target audience, story genre, and emotion to be conveyed, means for generating a story structure based on the set viewing conditions, means for setting the intention of each cut and composing scenes based on the generated story structure, means for rendering and outputting the final video content, means for identifying inaccurate information using a social media analysis tool, and means for selecting a story template optimal for the viewing conditions using a template matching algorithm. This allows the user to efficiently generate high-quality video content based on reliable information while eliminating cumbersome tasks.

[0984] "Theme" is a keyword that the user inputs as the central topic of the documentary video.

[0985] "Keywords" are relevant search terms that the system uses to gather information.

[0986] "Information" refers to any data, including text, images, video, and audio data, collected from networks and databases.

[0987] "Reliability" is the criterion for assessing whether the collected information is accurate and trustworthy.

[0988] "Natural language processing (NLP)" is a technology that analyzes collected text data and evaluates its reliability and categorizes it.

[0989] "Category" refers to the field or type used to classify collected information, such as "scientific data," "news articles," and "policy reports."

[0990] "Target audience" refers to the main audience of the documentary footage, and includes general viewers and experts.

[0991] "Story genre" refers to the type or style of documentary footage, including, for example, documentaries and educational films.

[0992] The "emotions to be conveyed" refer to the emotions and atmosphere that the viewer wants to convey through the video, and include, for example, tension and emotion.

[0993] "Story construction" involves organizing information and planning the flow of the footage based on set viewing conditions.

[0994] A "scene" refers to an individual scene or cut from a documentary film.

[0995] "Rendering" is the process of generating the entire image based on the collected information and the user's settings.

[0996] "Social Media Analytics Tools" are tools used to collect data from social media sites and identify inaccurate information.

[0997] The "template matching algorithm" is an algorithm for selecting the optimal story template based on the set viewing conditions.

[0998] "Video Content" refers to the final video file created based on the collected information.

[0999] This system provides a technical means for efficiently generating documentary footage. A specific embodiment will now be described.

[1000] System configuration

[1001] The system consists of the following main components: a user interface (GUI), a data collection module, a natural language processing (NLP) module, a template matching algorithm, a video editing module, and a rendering module.

[1002] User Interface

[1003] The user first inputs the documentary's theme and keywords through the GUI. This input is the initial instruction for the system, and for example, a word such as "marine pollution" is entered.

[1004] Data Collection Module

[1005] The server collects information based on the input topic or keyword, searches for related information from the network or internal database, and collects the necessary text, image, video, and audio data. For example, it uses Google Custom Search API to collect information from the Internet and pulls the necessary information from the internal database.

[1006] Natural Language Processing (NLP) Module

[1007] The server uses NLP technology to evaluate the reliability of the collected information. Specifically, it calculates a reliability score for each piece of information based on its source, number of citations, and related research data. It also uses social media analysis tools to filter out fake news and inaccurate information. NLP libraries such as spaCy and NLTK are used.

[1008] Categorizing information

[1009] Once the information has been assessed for reliability, it is sorted by the server into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and each piece of information is automatically labeled with the relevant category.

[1010] Viewing condition settings

[1011] Next, the user sets the target audience, the genre of the story, the emotions they want to convey, etc. These settings are again entered through the GUI and sent to the server. For example, possible settings include "general audience," "documentary," and "tension."

[1012] Selecting and configuring a story template

[1013] The server selects a story template based on the user's input settings and generates the story structure using a template matching algorithm. The selected template is then combined with the collected information to generate the storyboard.

[1014] Video editing

[1015] The device receives user instructions and selects the appropriate video clip for each cut. At this stage, video clips are dragged and dropped, narration is added, and transition effects are set. A preview of the video can also be viewed on the GUI.

[1016] Video rendering and output

[1017] Finally, the server renders all the scenes and outputs the final documentary video. The video is generated using a rendering tool such as FFmpeg in the specified format (e.g., MP4, MOV) and made available to users for download or viewing.

[1018] Specific examples

[1019] Consider a case where a user wants to create a documentary on the theme of "marine pollution." The user first enters "marine pollution," and the server uses the Google Custom Search API and an internal database to collect the latest research papers, news articles, government reports, and footage from environmental organizations. The collected information is evaluated for reliability using NLP, and inaccurate information is eliminated. Each piece of information is categorized as "scientific data," "news," "policy response," etc. The user sets viewing conditions such as "general audience," "documentary," and "urgency," and the server composes a story based on the template. The device arranges the footage and narration for each cut, and finally, the server integrates all the scenes and outputs the completed documentary video.

[1020] Prompt Sentence Examples

[1021] As a prompt sentence, provide the following sentence as input to the generative AI model:

[1022] Theme: Marine pollution

[1023] Keywords: plastics, climate change, marine life

[1024] Target audience: General audience

[1025] Story genre: Documentary

[1026] Emotions to convey: Urgency

[1027] Sources: networks, databases, news articles, reports, research papers

[1028] In this way, the present invention allows users to effectively and accurately generate high quality documentary footage.

[1029] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1030] Step 1: User enters topic and keywords

[1031] The user logs into the system's user interface (GUI) and is taken to a screen where they can enter a theme and keywords. The user enters the keywords "marine pollution" and "plastic" in the text field and clicks the "Search" button.

[1032] Input: topic (marine pollution) and keyword (plastic)

[1033] Output: The entered theme and keywords are sent to the server.

[1034] Step 2: Server collects information

[1035] The server receives the input topic and keywords and collects related information from the network and internal database. The server uses the Google Custom Search API to perform an internet search for the query "ocean pollution plastic" and scrapes useful information (text, images, video, audio data). It also queries the internal database to pull out related information.

[1036] Input: topic (marine pollution) and keyword (plastic)

[1037] Output: A list of collected text, image, video, and audio data

[1038] Step 3: The server evaluates trustworthiness using natural language processing (NLP)

[1039] The server applies NLP techniques to the collected information, specifically using the spaCy and NLTK libraries to calculate a credibility score for each piece of information. The server also evaluates the source, number of citations, and related research data for each piece of information, and uses social media analysis tools to filter out inaccurate information.

[1040] Input: A list of collected text, image, video, and audio data

[1041] Output: List of information with confidence scores

[1042] Step 4: The server categorizes the information

[1043] The server categorizes the information whose reliability has been evaluated. It analyzes the content of the text data and automatically categorizes it into categories such as "scientific data," "news articles," "policy reports," and "personal interviews." The server updates and saves this categorization information in a database.

[1044] Input: List of information with confidence scores

[1045] Output: A list of information sorted by category

[1046] Step 5: User sets target audience, genre, and emotion

[1047] The user again uses the GUI to set the target audience, the genre of the story, and the emotion they want to convey. Users select settings such as "general audience," "documentary," or "urgency" from drop-down menus and click the "Set" button.

[1048] Input: audience, story genre, emotion you want to convey

[1049] Output: Configuration information sent to the server

[1050] Step 6: The server selects a story template and generates a story structure.

[1051] The server runs a template matching algorithm based on the viewing conditions you set and selects the most suitable story template. The collected information is automatically incorporated into the selected template to generate a storyboard.

[1052] Input: audience, story genre, emotion to convey, list of information categorized

[1053] Output: Storyboard

[1054] Step 7: Select and edit video clips on your device

[1055] The user uses the device to select and edit video clips based on the generated storyboard. They drag and drop the appropriate video clips for each scene, add narration, and set transition effects. Editing proceeds while checking the video preview on the GUI.

[1056] Input: Storyboard, video clips

[1057] Output: Edited video sequence

[1058] Step 8: The server renders the entire scene and outputs the final video.

[1059] The server uses a rendering tool such as FFmpeg to combine all the scenes and render the video in the specified output format (e.g. MP4, MOV). The final video file is provided to the user, and a download link is displayed on the GUI.

[1060] Input: Edited video sequence

[1061] Output: Final video file (e.g. MP4, MOV)

[1062] Through the above steps, the user can efficiently generate high-quality documentary footage.

[1063] (Application example 1)

[1064] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1065] In today's world, documentary video production requires a high level of skill and time. Gathering accurate information tailored to the theme, assessing its reliability, and structuring a story that meets the needs of viewers are particularly challenging. Furthermore, amid a proliferation of fake information, there is a demand for video production that uses accurate and reliable information. Rendering and outputting the final video work is also a laborious process that requires increased efficiency. To address these challenges, there is an urgent need to provide a system that allows users to easily and efficiently generate and distribute high-quality documentary video.

[1066] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1067] In this invention, the server includes means for accepting input of a theme or keyword, means for collecting information related to the theme or keyword from a communication network and an internal database, means for evaluating the reliability of the collected information and determining its veracity, means for classifying the evaluated information by category, means for the user to set the target audience, story genre, and desired emotion, means for generating a story structure based on the set viewing conditions, means for setting the intention of each cut and composing scenes based on the generated story structure, means for rendering and outputting the final video work, and means for providing a user interface for the user to input a theme and automatically generate and distribute documentary video based on the collected information, thereby enabling users to easily and efficiently generate and distribute high-quality documentary video.

[1068] "Theme" is a concept that allows the user to indicate the central subject or interest of the documentary footage.

[1069] "Keywords" are indicators that contain important words or terms related to a topic and are used to narrow down information.

[1070] A "communications network" is a digital network for sending and receiving information, including the Internet.

[1071] An "internal database" is a source of information held within the system and is a storage device for storing and managing a portion of the collected data.

[1072] "Credibility" is a measure of the accuracy and veracity of collected information, and is calculated using NLP technology.

[1073] "Veracity" is the quality that determines whether information is based on actual, verified facts.

[1074] "User" refers to an individual or organization that performs operations to create documentary footage.

[1075] "Target audience" refers to the audience or target group that intends to watch the documentary footage.

[1076] "Story genre" is a classification that indicates the format and style of documentary footage.

[1077] "Emotions you want to convey" refers to the emotions and atmosphere you want viewers to feel through the documentary footage.

[1078] The "viewing conditions" are parameters related to the viewer set by the user, such as the viewing target, the genre of the story, and the emotion to be conveyed.

[1079] "Story structure" is the framework for designing the overall flow and development of a documentary film.

[1080] "The intention of each cut" refers to the purpose and message of each scene or cut in the video.

[1081] "Scenes" are the individual scenes or cuts that make up documentary footage.

[1082] "Video work" is a general term for the final documentary video.

[1083] "Rendering" is the process of converting an edited video clip into an output format as a single file.

[1084] "Output" refers to the act of saving or distributing the final video work in a specified format.

[1085] "User interface" refers to the screen and operating means through which a user interacts with and operates a system.

[1086] This invention relates to a system that efficiently generates and distributes documentary footage by allowing users to input themes and keywords. The system is configured as follows:

[1087] First, users input the documentary's theme and keywords through the smartphone application interface. For example, they can specify a theme such as "marine pollution." To accept this input, the system uses libraries such as React Native as its front end.

[1088] The server receives the input topic and keywords and collects related information from communication networks and internal databases. In this step, a crawler is used to collect text, images, video, and audio data from academic papers, news sites, government sites, and other sources on the Internet. The collected data is then stored in storage. The software used is Python and technologies such as Beautiful Soup and Scrapy.

[1089] Next, the reliability of the collected information is evaluated. TensorFlow or PyTorch is used to calculate a reliability score for the information using natural language processing (NLP) techniques. Furthermore, fake news and inaccurate information is filtered out using social media analysis tools. This analysis uses techniques such as sentiment analysis and named entity recognition (NER).

[1090] The evaluated information is then sorted into categories on the server, such as "scientific data," "news articles," "policy responses," and "personal interviews." The information is sorted into categories using labeling technology based on machine learning models.

[1091] Next, the user sets the target audience, story genre, and desired emotion on the application. These settings are necessary for the server to generate the optimal story structure. A template matching algorithm is used to select a story template and automatically generate the storyboard.

[1092] Based on the generated storyboard, the intention of each cut is set and the scene is composed. At this stage, the user inserts video clips, adds narration, and sets transition effects through the interface. The front-end editing function uses HTML5, CSS3, and JavaScript.

[1093] FFmpeg is used to render the final video and save it in the output format, which is then made available to users for download.

[1094] As a concrete example, consider a case where a user is creating a documentary on the theme of "marine pollution." When the user enters "marine pollution," the server collects the latest scientific data, news articles, and policy response information from the Internet. It evaluates the reliability of the collected information and categorizes it into categories such as "scientific data," "news," and "policy response." The user sets the target audience as general viewers, the story genre as documentary, and the emotion they want to convey as tension. The server composes the story based on a template and edits each cut through the user interface. Finally, the server integrates all the scenes and outputs the completed documentary video in MP4 format.

[1095] Examples of prompts include:

[1096] "A user wants to create a documentary about marine pollution. They need to gather the latest research, news, policy reports, etc., create a storyboard based on that information, and finally output it in MP4 format."

[1097] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1098] Step 1:

[1099] Users enter the documentary's theme and keywords into the smartphone application.

[1100] Input: Theme or keyword (e.g., marine pollution)

[1101] Output: Themes and keywords are sent to the server

[1102] Step 2:

[1103] The server collects related information from the communication network and internal database based on the input topic and keywords.

[1104] Input: Theme or keyword

[1105] Data processing: Using crawlers to collect information from academic papers, news sites, government sites, etc. on the Internet and store it in storage.

[1106] Output: Collected text, images, video, and audio data

[1107] Step 3:

[1108] The server uses NLP techniques to assess the reliability of the collected information.

[1109] Input: Collected information

[1110] Data Computation: Use TensorFlow or PyTorch to calculate the credibility score of information. Use social media analysis tools to filter out fake news and inaccurate information.

[1111] Output: reliable information

[1112] Step 4:

[1113] The server categorizes the information whose reliability has been evaluated.

[1114] Input: Information that has undergone a reliability assessment

[1115] Data processing: Using machine learning models, information is labeled into categories such as "scientific data," "news articles," "policy responses," and "personal interviews."

[1116] Output: Information sorted by category

[1117] Step 5:

[1118] Users set their viewing target, story genre, and the emotion they want to convey on the application.

[1119] Input: Target audience, story genre, and desired emotion (e.g., general audience, documentary, suspense)

[1120] Output: The configured viewing conditions are sent to the server.

[1121] Step 6:

[1122] The server selects the most suitable story template based on the set viewing conditions and generates a storyboard.

[1123] Input: Viewing conditions, information categorized by category

[1124] Data calculation: Using a template matching algorithm, a storyboard is automatically generated by selecting the story template that best suits the viewing conditions.

[1125] Output: Generated storyboard

[1126] Step 7:

[1127] The user sets the intention for each cut based on the generated storyboard and composes the scene.

[1128] Input: Generated storyboard

[1129] Specific operations: The user inserts a video clip, adds narration, and sets transition effects in the application interface.

[1130] Output: Edited storyboard

[1131] Step 8:

[1132] The server renders the final video production and outputs it in the specified format.

[1133] Input: Edited storyboard

[1134] Specific operation: All scenes are integrated using FFmpeg and the final video is rendered in MP4 format or similar.

[1135] Output: The finished documentary footage is made available to users in a downloadable format

[1136] As an example, the following prompt sentence is used:

[1137] "A user wants to create a documentary about marine pollution. They need to gather the latest research, news, policy reports, etc., create a storyboard based on that information, and finally output it in MP4 format."

[1138] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1139] The present invention is a system that collects related information based on specific themes and keywords, evaluates and classifies reliability, and generates high-quality documentary footage by recognizing the user's viewing preferences and emotions.

[1140] System Configuration

[1141] 1. How to input themes and keywords

[1142] The user inputs the documentary's theme or keywords. For example, "marine pollution." The user interface (UI) is designed to be intuitive.

[1143] 2. Information gathering methods

[1144] The server collects relevant information from the internet and internal databases using web scraping and API integration to obtain text, images, video, and audio data.

[1145] 3. Reliability assessment methods

[1146] The server uses natural language processing (NLP) techniques to assess the reliability of collected information, analyzing its source, number of citations, and related research data to calculate a credibility score. It also uses social media analysis tools to identify and filter out fake news.

[1147] 4. Information classification means

[1148] The server categorizes the evaluated information into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and automatically labels each piece of information with the appropriate category.

[1149] 5. Viewing setting input method

[1150] The user sets the target audience (e.g., general audience), the genre of the story (e.g., documentary), and the emotion they want to convey (e.g., tension). These settings are made through a user interface.

[1151] 6. Emotion Engine

[1152] The emotion engine analyzes the user's emotions in real time as they type, and reflects this in the story structure and scene settings. Emotion analysis is performed based on the user's facial expressions, tone of voice, and typing speed.

[1153] 7. Story Generation Methods

[1154] The server selects the optimal story template based on the viewing preferences and the analysis results of the emotion engine, and generates a story structure. Using a template matching algorithm, the collected information is arranged on the storyboard.

[1155] 8. Scene Composition Methods

[1156] The device allows the user to compose scenes by setting the intention for each cut. Using drag and drop, the user places the appropriate video clips for each cut, adds narration, and sets transition effects. Based on the emotional state recognized by the emotion engine, appropriate narration and music are automatically selected and placed.

[1157] 9. Video rendering and output methods

[1158] The server stitches all the scenes together and renders the final documentary video. The video file is generated in the specified output format (e.g. MP4, MOV) and made available to the user for download or viewing.

[1159] Specific examples

[1160] If a user wants to create a documentary on the theme of "marine pollution," they first input the theme. The server collects related information and evaluates and classifies it for reliability. The user then selects a general audience, a documentary genre, and a desire to convey a sense of tension. The emotion engine recognizes the user's sense of tension as they input, and reflects this in the story structure and scene settings. Appropriate video clips, narration, and music are selected, and finally, the server renders all the scenes, outputting the finished documentary video.

[1161] In this way, the present invention provides a system that efficiently generates high-quality documentary footage while taking into account the user's emotions.

[1162] The processing flow will be explained below.

[1163] Step 1:

[1164] The user inputs the documentary's theme or keywords. For example, "marine pollution."

[1165] Step 2:

[1166] The server collects information based on themes and keywords. The information is collected using web scraping and API integration to obtain text, images, video, and audio data from the internet and internal databases.

[1167] Step 3:

[1168] The server evaluates the credibility of the collected information, using natural language processing (NLP) techniques to analyze the information's source, number of citations, and related research data to calculate a credibility score, and also uses social media analysis tools to identify and remove fake news and inaccurate information.

[1169] Step 4:

[1170] The server categorizes the evaluated information into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and each piece of information is automatically labeled with the appropriate category.

[1171] Step 5:

[1172] The user sets the target audience, the genre of the story, and the emotion they want to convey. These settings are entered through a highly visible user interface. For example, the target audience can be set to general viewers, the genre to documentary, and the emotion they want to convey to be tension.

[1173] Step 6:

[1174] The emotion engine analyzes the user's emotions in real time, analyzing facial expressions, tone of voice, and input speed when entering settings to recognize the user's current emotional state.

[1175] Step 7:

[1176] The server selects the optimal story template based on the user's viewing preferences and the analysis results of the emotion engine. Using a template matching algorithm, it selects the template that best suits the preferences and analysis data, and places the collected information on the storyboard.

[1177] Step 8:

[1178] The device sets the intention for each cut and composes the scene. The user uses drag and drop to place the appropriate video clips for each cut, add narration, and set transition effects. Appropriate narration and music are also automatically selected and placed based on the emotional state recognized by the emotion engine.

[1179] Step 9:

[1180] The server stitches all the scenes together and renders the final documentary video. The video file is generated in the specified output format (e.g. MP4, MOV) and made available to the user for download or viewing.

[1181] By following these steps, users can efficiently generate high-quality documentary footage that reflects emotions.

[1182] Example 2

[1183] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1184] In conventional documentary video production systems, the processes of information collection, credibility assessment, information classification, viewing settings, emotion analysis, story generation, scene composition, and video rendering are all performed manually, which is extremely time-consuming and labor-intensive. Furthermore, they lack the ability to analyze user emotions in real time and reflect them in video production, resulting in a lack of emotional impact on viewers. Furthermore, with the rise of fake news, unreliable information is often mixed in, making it difficult to produce high-quality documentary videos based on accurate information.

[1185] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for accepting input of themes and keywords, means for collecting information from a network and an internal database, means for evaluating the reliability of the collected information and determining its authenticity, means for identifying fake information using social media analysis technology, means for categorizing the evaluated information, means for the user to set the viewing target, story genre, and emotion to be conveyed, means for analyzing the user's emotions in real time using an emotion analysis means and reflecting the results in scene settings, means for selecting a story template optimal for the viewing conditions using a template matching algorithm, means for setting the intention of each scene based on the generated story structure and composing scenes, and means for synthesizing and outputting the final documentary video. This enables automatic collection and reliability evaluation of information, appropriate information classification, real-time analysis of user emotions, and automatic selection of the optimal story template, thereby enabling efficient generation of high-quality documentary video with emotional impact.

[1186] The "means for accepting input of themes and keywords" is a means for providing an interface that allows a user to input themes and keywords necessary for producing a documentary video.

[1187] "Means for collecting information from networks and internal databases" refers to means for automatically collecting information related to the entered topic or keyword via the Internet or internal databases.

[1188] "Means for assessing the reliability of collected information and determining its authenticity" refers to means for assessing the reliability of collected information using natural language processing technology and source analysis, and determining the authenticity of the information.

[1189] "Methods for identifying fake information using social media analysis technology" refers to methods that use technology to analyze data on social media and identify misinformation and fake news.

[1190] A "means of categorizing evaluated information" is a means of classifying information that has been collected and evaluated for reliability into specific categories, such as "scientific data" or "news articles."

[1191] "A means for users to set the target audience, story genre, and emotions they want to convey" is a means for providing an interface that allows users to set the audience, genre, and emotions they want to convey for the documentary footage.

[1192] "Means for analyzing a user's emotions in real time using emotion analysis means and reflecting the emotions in scene settings" refers to means for analyzing a user's facial expressions and tone of voice, grasping the user's emotional state in real time, and reflecting the results in scene settings.

[1193] "Means for selecting a story template optimal for viewing conditions using a template matching algorithm" refers to means for executing an algorithm that selects the optimal template from among pre-prepared story templates based on the user's viewing settings and emotion analysis results.

[1194] "Means for setting the intention of each scene based on the generated story structure and composing scenes" refers to means for setting the purpose and intention of each scene based on the generated story and arranging corresponding images and narration.

[1195] The "means for compositing and outputting the final documentary footage" refers to a means for combining all the scenes into one, generating the final documentary footage, and outputting it in a specified format.

[1196] This invention is a system that collects related information based on specific themes and keywords, evaluates and classifies its reliability, recognizes the user's viewing preferences and emotions, and automatically generates high-quality documentary footage.

[1197] System Configuration

[1198] The system has the following main components:

[1199] 1. How to input themes and keywords

[1200] Users input the documentary's theme and keywords through a user interface (UI), which is designed for intuitive operation.

[1201] 2. Information gathering methods

[1202] The server collects relevant information from the internet and internal databases using web scraping tools (e.g., BeautifulSoup) and API integration (e.g., Twitter API), which allows it to obtain text, images, video, and audio data.

[1203] 3. Reliability assessment methods

[1204] The server evaluates the reliability of the collected information using natural language processing (NLP) techniques (e.g., BERT, GPT-3), and also identifies and eliminates fake information using social media analysis techniques.

[1205] 4. Information classification means

[1206] The server categorizes the evaluated information using a clustering algorithm (e.g., K-means) to automatically categorize it into categories such as "scientific data," "news articles," "policy reports," and "personal interviews."

[1207] 5. Viewing setting input method

[1208] Users can set their target audience, the genre of the story, and the emotion they want to convey through a user interface.

[1209] 6. Emotion analysis method

[1210] The server is equipped with an emotion analysis engine that analyzes the user's emotions in real time using software that analyzes facial expressions (e.g., OpenCV) and voice tone (e.g., Deepgram), and also combines the user's input speed for analysis.

[1211] 7. Story Generation Methods

[1212] The server runs a template matching algorithm to select the most appropriate story template based on viewing preferences and sentiment analysis results, and places the collected information into the appropriate storyboard.

[1213] 8. Scene Composition Methods

[1214] The user uses the drag-and-drop function on the device to place appropriate video clips for each cut on a storyboard, and then sets narration and transition effects. Appropriate narration and music are automatically selected based on the emotional state recognized by the emotion analysis means.

[1215] 9. Video rendering and output methods

[1216] The server combines all the scenes to generate the final documentary video and renders it as a video file in the specified format, such as MP4 or MOV.

[1217] Specific examples

[1218] Consider a scenario where a user is creating a documentary on the theme of "marine pollution." First, the user inputs the theme. The server collects related information, evaluates its reliability, and classifies it. Next, the user sets the target audience as "general audience," the genre as "documentary," and the emotion they want to convey as "tension." The emotion analysis engine recognizes the user's emotions in real time as they are input, and reflects them in the story structure and scene settings. Appropriate video clips, narration, and music are selected, and the server renders all scenes to output the completed documentary video.

[1219] Prompt Sentence Examples

[1220] "Collect information to create a documentary film on the following topic, assess its reliability, classify it, generate a story, and output the final film. The user settings are for general audiences, the genre is documentary, and the emotion you want to convey is a sense of urgency. The topic is 'Marine Pollution.'"

[1221] This system makes it possible to efficiently generate high-quality documentary footage while taking into account the user's emotions.

[1222] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1223] Step 1:

[1224] Accepts input of themes and keywords.

[1225] The user inputs themes and keywords through a user interface (UI).

[1226] Specific behavior: The user enters "Marine Pollution" in the topic field and clicks the submit button.

[1227] Input: Theme or keyword (e.g. "Marine pollution")

[1228] Output: Themes and keywords entered

[1229] Step 2:

[1230] Automatic Collection of Information.

[1231] The server collects information from the Internet and internal databases based on the entered themes and keywords.

[1232] Specific operation: The server uses a web scraping tool (e.g., BeautifulSoup) to retrieve articles from Google search results and uses an API (e.g., Twitter API) to collect related tweets.

[1233] Input: Theme or keyword (e.g. "Marine pollution")

[1234] Output: Collected information (text, images, video, audio data)

[1235] Step 3:

[1236] Evaluating the reliability of information.

[1237] The server evaluates the reliability of the collected information using natural language processing (NLP) techniques.

[1238] How it works: The server uses NLP models (e.g., BERT, GPT-3) to calculate a credibility score for each article, and also uses social media analysis techniques to identify and remove misinformation.

[1239] Input: Collected information

[1240] Output: Information with assessed reliability (with score)

[1241] Step 4:

[1242] Classification of information.

[1243] The server categorizes the evaluated information.

[1244] What it does: The server uses a clustering algorithm (e.g., K-means) to classify information into categories such as "scientific data," "news articles," "policy reports," and "personal interviews."

[1245] Input: Information with reliability assessment

[1246] Output: Categorically labeled information

[1247] Step 5:

[1248] Enter viewing settings.

[1249] The user inputs the target audience, the genre of the story, and the emotion they want to convey.

[1250] Specific operation: The user selects the target audience as "general audience," the genre as "documentary," and the emotion to be conveyed as "tension."

[1251] Input: target audience, story genre, emotion you want to convey

[1252] Output: Viewing Settings

[1253] Step 6:

[1254] Performing sentiment analysis.

[1255] The server analyzes the user's emotions in real time.

[1256] How it works: The server uses facial expression analysis software (e.g., OpenCV) and voice analysis software (e.g., Deepgram) to detect the user's emotions. The user's typing speed is also used for analysis.

[1257] Input: User's facial expression data, tone of voice, typing speed

[1258] Output: Real-time analyzed emotion data

[1259] Step 7:

[1260] Story generation.

[1261] The server selects a story template based on viewing preferences and emotion analysis results, and generates a story structure.

[1262] Specific operation: The server uses a template matching algorithm to select an appropriate story template and places the collected information on the storyboard.

[1263] Input: Viewing settings, sentiment analysis results, and information categorized by category

[1264] Output: Generated story structure

[1265] Step 8:

[1266] Scene placement and adjustment.

[1267] The terminal allows the user to set the intention for each cut and compose a scene.

[1268] How it works: Users use the drag-and-drop function to place video clips on the storyboard and set narration and transition effects. Appropriate narration and music are automatically selected based on emotional data.

[1269] Input: Generated story structure, video clips, narration, transition effects

[1270] Output: Composed Scene

[1271] Step 9:

[1272] Video rendering and output.

[1273] The server stitches all the scenes together and renders the final documentary footage.

[1274] What it does: The server uses video editing software (e.g., FFmpeg) to combine and render all the scenes into a single video file in MP4 or MOV format.

[1275] Input: Composed Scene

[1276] Output: Final documentary footage file

[1277] (Application example 2)

[1278] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1279] In today's world, users are seeking to generate and view high-quality documentary videos based on their own interests. However, conventional methods have difficulty collecting reliable information and providing it in an optimal format based on the viewer's emotions and preferences. Furthermore, there is no system that can analyze the viewer's emotional state in real time and dynamically change the video content accordingly, which hinders user satisfaction.

[1280] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1281] In this invention, the server includes means for accepting input of themes and keywords, means for collecting information related to the themes and keywords from a communication network and an internal database, means for evaluating the reliability of the collected information and determining its authenticity, means for classifying the evaluated information by type, means for the user to set the viewing target, story genre, and emotion to be conveyed, means for analyzing the user's emotional state in real time, means for generating a story structure based on the set viewing conditions and the analyzed emotional state, means for setting the intention of each cut and composing scenes based on the generated story structure, and means for rendering and outputting the final documentary video. This makes it possible to generate and provide high-quality, reliable documentary video in real time based on the viewing conditions and emotional state set by the user.

[1282] The "means for accepting input of themes and keywords" is an interface that allows a user to input subjects of interest or specific words into the system.

[1283] "Means for collecting information related to the theme or keyword from communication networks and internal databases" refers to a function for collecting information from the Internet and obtaining related information from internally stored databases.

[1284] "Means for assessing the reliability of collected information and determining its authenticity" refers to techniques for analyzing the reliability of the data obtained and determining whether the data is accurate.

[1285] The "means for classifying evaluated information by type" refers to a method for classifying information whose reliability has been evaluated into different categories based on its content.

[1286] "A means for users to set the target audience, story genre, and emotions they want to convey" is an interface that allows users to set the target audience, video genre, and emotions they want to convey through the video.

[1287] The "means for analyzing the user's emotional state in real time" is a technology for detecting and analyzing the user's current emotions in real time.

[1288] The "means for generating a story structure based on the set viewing conditions and the analyzed emotional state" is a method for automatically constructing an optimal story based on the viewing conditions set by the user and the analyzed emotions.

[1289] "Means for setting the intention of each cut based on the generated story structure and composing scenes" refers to a technique for setting the content and purpose of each scene based on the generated story structure and then visualizing it.

[1290] "Means for rendering and outputting the final documentary footage" refers to the function for integrating all scenes to generate the final footage and outputting it in a specified format.

[1291] The present invention is a system for generating high-quality documentary videos based on user viewing preferences and real-time emotion analysis. The system operates by combining multiple hardware and software technologies and is implemented in the following configuration.

[1292] Hardware and software used

[1293] Server: Responsible for information gathering, reliability assessment, data classification, sentiment analysis, story generation, scene composition, and rendering.

[1294] Smart devices (smartphones, smart TVs): Collect data on user input, viewing, and sentiment analysis.

[1295] Communications Network: Collecting and distributing data through internet connections and internal networks.

[1296] Specific software technologies used include:

[1297] Natural Language Processing (NLP) engine: Used to understand and analyze text and assess the reliability of information.

[1298] Machine learning model: Used for user sentiment analysis.

[1299] Template matching algorithm: Used to select the best story template.

[1300] Rendering engine: Used to generate the final documentary footage.

[1301] Cloud storage and databases: Used to store information and manage access.

[1302] Specific examples of the system

[1303] 1. Enter your topic or keywords:

[1304] Users input a theme or keyword, such as "marine pollution," through the interface of their smart device.

[1305] The interface is designed to be intuitive.

[1306] Example: User types "ocean pollution".

[1307] 2. Information Collection:

[1308] The server collects relevant information from the Internet and an internal database via a communication network.

[1309] Using web scraping and API integration, text, images, videos, audio data, etc. are obtained.

[1310] Example: A server collects academic papers, news articles, and visual materials on "ocean pollution."

[1311] 3. Reliability assessment:

[1312] The server uses NLP techniques to evaluate the reliability of the collected information.

[1313] The reliability score is calculated by analyzing the source of the information, the number of citations, and related research data.

[1314] Use social media analytics tools to filter out fake news.

[1315] Example: An NLP engine evaluates the reliability of collected news articles and filters out articles with low scores.

[1316] 4. Information classification:

[1317] The server categorizes the evaluated information.

[1318] Categories include "scientific data," "news articles," "policy reports," and "personal interviews."

[1319] Example: Collected information is divided into categories such as "scientific data" and "personal interviews."

[1320] 5. Viewing Preferences and Sentiment Analysis:

[1321] Users set the target audience, story genre, and the emotion they want to convey.

[1322] The server uses machine learning models to analyze emotions in real time based on the user's facial expressions and input speed collected from the smart device.

[1323] Example: User selects the documentary genre and wants a thrilling story.

[1324] 6. Story generation and scene composition:

[1325] The server generates an optimal story based on the set viewing conditions and the analyzed emotional state.

[1326] A template matching algorithm is used to select an appropriate story template.

[1327] The intention of each cut is set based on the generated story structure, and scenes are composed.

[1328] Example: The server automatically generates scenes with the theme of "tension" based on emotion analysis.

[1329] 7. Video rendering and output:

[1330] The server renders the final documentary footage and outputs it in the format you specify (e.g. MP4, MOV).

[1331] Users can watch on their smart devices or web browsers.

[1332] Example: A completed documentary film is streamed to a smart TV.

[1333] Prompt Sentence Examples

[1334] Please enter a topic (e.g., "Marine Pollution").

[1335] "Assessing reliability of visual data..."

[1336] "Analyzing user sentiment... Anything else you'd like to add?"

[1337] As a result, the present invention provides a system that efficiently generates and provides high-quality documentary footage while taking into consideration the user's emotions.

[1338] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1339] Step 1:

[1340] The user inputs themes and keywords through the interface of the smart device.

[1341] Input: User types in the topic "marine pollution."

[1342] Specific operation: The smart device interface receives input from the user and sends themes and keywords to the server.

[1343] Step 2:

[1344] The server collects information related to themes and keywords from the communication network and an internal database.

[1345] Input: Query regarding "marine pollution."

[1346] Data processing: Use web scraping and API integration to obtain relevant information.

[1347] Output: Collected text, images, video and audio data.

[1348] What it does: The server collects academic papers, news articles, visual materials, etc. related to "marine pollution."

[1349] Step 3:

[1350] The server evaluates the reliability of the collected information using natural language processing (NLP) techniques.

[1351] Input: Collected information.

[1352] Data calculations: Analysis of sources, citations, and related research data. Calculation of reliability scores.

[1353] Output: Information rated as reliable and its score.

[1354] How it works: The NLP engine evaluates the credibility of news articles and papers and assigns them a credibility score.

[1355] Step 4:

[1356] The server classifies the information whose reliability has been evaluated by type.

[1357] Input: Reliability-assessed information.

[1358] Data processing: Using classification algorithms, information is sorted into categories such as "scientific data," "news articles," "policy reports," and "personal interviews."

[1359] Output: Information broken down by category.

[1360] What it does: The server automatically categorizes collected information into appropriate categories based on trustworthiness.

[1361] Step 5:

[1362] Users set the target audience, story genre, and the emotion they want to convey through the interface of their smart device.

[1363] Input: User viewing preferences (e.g., "General Audience," "Documentary," "Intense").

[1364] Output: The set viewing conditions.

[1365] Specific operation: The user operates the smart device and sets up viewing.

[1366] Step 6:

[1367] The server uses a machine learning model to analyze emotions in real time based on the user's facial expressions and input speed collected from the smart device.

[1368] Input: Real-time facial expression data and input speed of the user.

[1369] Data computation: Sentiment analysis with machine learning models.

[1370] Output: User's emotional state (e.g., sense of urgency).

[1371] Specific operation: The server analyzes the user's emotions in real time, and collects and processes the data.

[1372] Step 7:

[1373] The server generates an optimal story based on the set viewing conditions and the analyzed emotional state.

[1374] Input: User's viewing preferences and emotional state, information broken down into categories.

[1375] Data calculation: A template matching algorithm is used to select the most suitable story template.

[1376] Output: The generated story structure.

[1377] What happens: The server uses a template matching algorithm to select an appropriate story template and generate a story.

[1378] Step 8:

[1379] Based on the story structure generated by the server, the intention of each cut is set and scenes are composed.

[1380] Input: The generated story structure.

[1381] Data processing: Automatically place appropriate video clips, narration, and music.

[1382] Output: The composed scene.

[1383] Specific operation: The server arranges video clips and music based on the automatically generated story structure to create scenes.

[1384] Step 9:

[1385] The server renders and outputs the final documentary footage.

[1386] Input: The composed scene.

[1387] Data calculation: Generate the final image using the rendering engine.

[1388] Output: Final documentary footage (e.g. MP4, MOV).

[1389] What it does: The server stitches all the scenes together, renders the final documentary footage, and outputs it in the specified format.

[1390] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1391] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1392] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1393] [Fourth embodiment]

[1394] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1395] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1396] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1397] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1398] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1399] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1400] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1401] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1402] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1403] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1404] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1405] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1406] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1407] This system provides a technical means for efficiently generating documentary footage. A specific embodiment will now be described.

[1408] The user first inputs the documentary's theme or keywords, which are the initial instructions for the system, for example, "marine pollution."

[1409] The server then collects information based on the input topic and keywords, searches the internet and internal databases for relevant information, and collects the necessary text, images, video, and audio data. The collected data is then stored in storage.

[1410] To assess the reliability of the collected information, the server uses natural language processing (NLP) techniques to calculate a credibility score for each piece of information based on its source, number of citations, related research data, etc. It also uses social media analysis tools to filter out fake news and inaccurate information.

[1411] Once the information has been assessed for trustworthiness, it is then sorted by the server into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews." Each piece of information is automatically labeled with the relevant category.

[1412] Next, the user sets the target audience, the genre of the story, the emotion they want to convey, etc. These settings are input through a user interface.

[1413] The server selects a story template based on the input settings and generates the story structure. This process uses a template matching algorithm to select the most suitable story template. The selected template is then incorporated with the collected information to generate the storyboard.

[1414] The next step is to set the intention for each cut and compose the scene. The device receives instructions from the user and selects the appropriate video clip for each cut. At this stage, video clips are dragged and dropped, narration is added, and transition effects are set.

[1415] Finally, the server renders all the scenes and outputs the final documentary video, which is generated in the specified output format (e.g. MP4, MOV) and made available to the user for download or viewing.

[1416] As a concrete example, consider a user creating a documentary on the theme of "marine pollution." First, the user enters "marine pollution," and the server collects the latest research papers, news articles, government reports, and footage from environmental organizations. After assessing the reliability and eliminating fake news, the information is classified into categories such as "scientific data," "news," and "policy responses." The user sets the target audience as general viewers, the story genre as documentary, and the emotion they want to convey as tension. The server composes the story based on a template, and the device places the footage and narration for each cut. Finally, the server integrates all the scenes and outputs the completed documentary footage.

[1417] In this way, the present invention allows for efficient and accurate generation of high quality documentary footage.

[1418] The processing flow will be explained below.

[1419] Step 1:

[1420] The user inputs the documentary's theme or keywords. For example, "marine pollution."

[1421] Step 2:

[1422] The server collects information based on themes and keywords. This includes retrieving relevant text, images, video, and audio data from the internet and internal databases. The server gathers the necessary data through web scraping and API integration.

[1423] Step 3:

[1424] The server evaluates the credibility of the collected information, using natural language processing (NLP) techniques to analyze each piece of information's source, number of citations, and related research data to calculate a credibility score, and also uses social media analysis tools to identify and remove fake news and inaccurate information.

[1425] Step 4:

[1426] The server categorizes the evaluated information into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and automatically labels and organizes the information.

[1427] Step 5:

[1428] The user sets the target audience, the genre of the story, and the emotion they want to convey. These settings are input through a user interface. For example, the target audience is set to general viewers, the genre is set to documentary, and the emotion they want to convey is set to tension.

[1429] Step 6:

[1430] The server selects a story template based on the input settings, uses a template matching algorithm to select the most suitable template, and places the collected information into the storyboard.

[1431] Step 7:

[1432] The device determines the intent of each cut and composes the scene. The user uses drag and drop to place the appropriate video clips for each cut, add narration, and set transition effects. The device then makes the final adjustments to the scene.

[1433] Step 8:

[1434] The server stitches all the scenes together and renders the final documentary video, generating the final video file in the specified output format (e.g. MP4, MOV) for users to download or watch.

[1435] Through the above steps, users can efficiently generate high-quality documentary footage.

[1436] Example 1

[1437] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1438] In today's information society, extracting reliable information from vast amounts of data and efficiently creating video content based on that information is a major challenge. While documentary video production in particular requires accurate and reliable information, it also requires a great deal of time and effort, including information gathering, reliability assessment, story construction, and video editing. Therefore, there is a demand for technological tools that allow users to easily generate high-quality documentary video.

[1439] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1440] In this invention, the server includes means for accepting input of a theme or keyword, means for collecting information related to the theme or keyword from a network and a database, means for evaluating the reliability of the collected information using natural language processing and determining its authenticity, means for categorizing the evaluated information, means for the user to set the target audience, story genre, and emotion to be conveyed, means for generating a story structure based on the set viewing conditions, means for setting the intention of each cut and composing scenes based on the generated story structure, means for rendering and outputting the final video content, means for identifying inaccurate information using a social media analysis tool, and means for selecting a story template optimal for the viewing conditions using a template matching algorithm. This allows the user to efficiently generate high-quality video content based on reliable information while eliminating cumbersome tasks.

[1441] "Theme" is a keyword that the user inputs as the central topic of the documentary video.

[1442] "Keywords" are relevant search terms that the system uses to gather information.

[1443] "Information" refers to any data, including text, images, video, and audio data, collected from networks and databases.

[1444] "Reliability" is the criterion for assessing whether the collected information is accurate and trustworthy.

[1445] "Natural language processing (NLP)" is a technology that analyzes collected text data and evaluates its reliability and categorizes it.

[1446] "Category" refers to the field or type used to classify collected information, such as "scientific data," "news articles," and "policy reports."

[1447] "Target audience" refers to the main audience of the documentary footage, and includes general viewers and experts.

[1448] "Story genre" refers to the type or style of documentary footage, including, for example, documentaries and educational films.

[1449] The "emotions to be conveyed" refer to the emotions and atmosphere that the viewer wants to convey through the video, and include, for example, tension and emotion.

[1450] "Story construction" involves organizing information and planning the flow of the footage based on set viewing conditions.

[1451] A "scene" refers to an individual scene or cut from a documentary film.

[1452] "Rendering" is the process of generating the entire image based on the collected information and the user's settings.

[1453] "Social Media Analytics Tools" are tools used to collect data from social media sites and identify inaccurate information.

[1454] The "template matching algorithm" is an algorithm for selecting the optimal story template based on the set viewing conditions.

[1455] "Video Content" refers to the final video file created based on the collected information.

[1456] This system provides a technical means for efficiently generating documentary footage. A specific embodiment will now be described.

[1457] System configuration

[1458] The system consists of the following main components: a user interface (GUI), a data collection module, a natural language processing (NLP) module, a template matching algorithm, a video editing module, and a rendering module.

[1459] User Interface

[1460] The user first inputs the documentary's theme and keywords through the GUI. This input is the initial instruction for the system, and for example, a word such as "marine pollution" is entered.

[1461] Data Collection Module

[1462] The server collects information based on the input topic or keyword, searches for related information from the network or internal database, and collects the necessary text, image, video, and audio data. For example, it uses Google Custom Search API to collect information from the Internet and pulls the necessary information from the internal database.

[1463] Natural Language Processing (NLP) Module

[1464] The server uses NLP technology to evaluate the reliability of the collected information. Specifically, it calculates a reliability score for each piece of information based on its source, number of citations, and related research data. It also uses social media analysis tools to filter out fake news and inaccurate information. NLP libraries such as spaCy and NLTK are used.

[1465] Categorizing information

[1466] Once the information has been assessed for reliability, it is sorted by the server into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and each piece of information is automatically labeled with the relevant category.

[1467] Viewing condition settings

[1468] Next, the user sets the target audience, the genre of the story, the emotions they want to convey, etc. These settings are again entered through the GUI and sent to the server. For example, possible settings include "general audience," "documentary," and "tension."

[1469] Selecting and configuring a story template

[1470] The server selects a story template based on the user's input settings and generates the story structure using a template matching algorithm. The selected template is then combined with the collected information to generate the storyboard.

[1471] Video editing

[1472] The device receives user instructions and selects the appropriate video clip for each cut. At this stage, video clips are dragged and dropped, narration is added, and transition effects are set. A preview of the video can also be viewed on the GUI.

[1473] Video rendering and output

[1474] Finally, the server renders all the scenes and outputs the final documentary video. The video is generated using a rendering tool such as FFmpeg in the specified format (e.g., MP4, MOV) and made available to users for download or viewing.

[1475] Specific examples

[1476] Consider a case where a user wants to create a documentary on the theme of "marine pollution." The user first enters "marine pollution," and the server uses the Google Custom Search API and an internal database to collect the latest research papers, news articles, government reports, and footage from environmental organizations. The collected information is evaluated for reliability using NLP, and inaccurate information is eliminated. Each piece of information is categorized as "scientific data," "news," "policy response," etc. The user sets viewing conditions such as "general audience," "documentary," and "urgency," and the server composes a story based on the template. The device arranges the footage and narration for each cut, and finally, the server integrates all the scenes and outputs the completed documentary video.

[1477] Prompt Sentence Examples

[1478] As a prompt sentence, provide the following sentence as input to the generative AI model:

[1479] Theme: Marine pollution

[1480] Keywords: plastics, climate change, marine life

[1481] Target audience: General audience

[1482] Story genre: Documentary

[1483] Emotions to convey: Urgency

[1484] Sources: networks, databases, news articles, reports, research papers

[1485] In this way, the present invention allows users to effectively and accurately generate high quality documentary footage.

[1486] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1487] Step 1: User enters topic and keywords

[1488] The user logs into the system's user interface (GUI) and is taken to a screen where they can enter a theme and keywords. The user enters the keywords "marine pollution" and "plastic" in the text field and clicks the "Search" button.

[1489] Input: topic (marine pollution) and keyword (plastic)

[1490] Output: The entered theme and keywords are sent to the server.

[1491] Step 2: Server collects information

[1492] The server receives the input topic and keywords and collects related information from the network and internal database. The server uses the Google Custom Search API to perform an internet search for the query "ocean pollution plastic" and scrapes useful information (text, images, video, audio data). It also queries the internal database to pull out related information.

[1493] Input: topic (marine pollution) and keyword (plastic)

[1494] Output: A list of collected text, image, video, and audio data

[1495] Step 3: The server evaluates trustworthiness using natural language processing (NLP)

[1496] The server applies NLP techniques to the collected information, specifically using the spaCy and NLTK libraries to calculate a credibility score for each piece of information. The server also evaluates the source, number of citations, and related research data for each piece of information, and uses social media analysis tools to filter out inaccurate information.

[1497] Input: A list of collected text, image, video, and audio data

[1498] Output: List of information with confidence scores

[1499] Step 4: The server categorizes the information

[1500] The server categorizes the information whose reliability has been evaluated. It analyzes the content of the text data and automatically categorizes it into categories such as "scientific data," "news articles," "policy reports," and "personal interviews." The server updates and saves this categorization information in a database.

[1501] Input: List of information with confidence scores

[1502] Output: A list of information sorted by category

[1503] Step 5: User sets target audience, genre, and emotion

[1504] The user again uses the GUI to set the target audience, the genre of the story, and the emotion they want to convey. Users select settings such as "general audience," "documentary," or "urgency" from drop-down menus and click the "Set" button.

[1505] Input: audience, story genre, emotion you want to convey

[1506] Output: Configuration information sent to the server

[1507] Step 6: The server selects a story template and generates a story structure.

[1508] The server runs a template matching algorithm based on the viewing conditions you set and selects the most suitable story template. The collected information is automatically incorporated into the selected template to generate a storyboard.

[1509] Input: audience, story genre, emotion to convey, list of information categorized

[1510] Output: Storyboard

[1511] Step 7: Select and edit video clips on your device

[1512] The user uses the device to select and edit video clips based on the generated storyboard. They drag and drop the appropriate video clips for each scene, add narration, and set transition effects. Editing proceeds while checking the video preview on the GUI.

[1513] Input: Storyboard, video clips

[1514] Output: Edited video sequence

[1515] Step 8: The server renders the entire scene and outputs the final video.

[1516] The server uses a rendering tool such as FFmpeg to combine all the scenes and render the video in the specified output format (e.g. MP4, MOV). The final video file is provided to the user, and a download link is displayed on the GUI.

[1517] Input: Edited video sequence

[1518] Output: Final video file (e.g. MP4, MOV)

[1519] Through the above steps, the user can efficiently generate high-quality documentary footage.

[1520] (Application example 1)

[1521] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1522] In today's world, documentary video production requires a high level of skill and time. Gathering accurate information tailored to the theme, assessing its reliability, and structuring a story that meets the needs of viewers are particularly challenging. Furthermore, amid a proliferation of fake information, there is a demand for video production that uses accurate and reliable information. Rendering and outputting the final video work is also a laborious process that requires increased efficiency. To address these challenges, there is an urgent need to provide a system that allows users to easily and efficiently generate and distribute high-quality documentary video.

[1523] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1524] In this invention, the server includes means for accepting input of a theme or keyword, means for collecting information related to the theme or keyword from a communication network and an internal database, means for evaluating the reliability of the collected information and determining its veracity, means for classifying the evaluated information by category, means for the user to set the target audience, story genre, and desired emotion, means for generating a story structure based on the set viewing conditions, means for setting the intention of each cut and composing scenes based on the generated story structure, means for rendering and outputting the final video work, and means for providing a user interface for the user to input a theme and automatically generate and distribute documentary video based on the collected information, thereby enabling users to easily and efficiently generate and distribute high-quality documentary video.

[1525] "Theme" is a concept that allows the user to indicate the central subject or interest of the documentary footage.

[1526] "Keywords" are indicators that contain important words or terms related to a topic and are used to narrow down information.

[1527] A "communications network" is a digital network for sending and receiving information, including the Internet.

[1528] An "internal database" is a source of information held within the system and is a storage device for storing and managing a portion of the collected data.

[1529] "Credibility" is a measure of the accuracy and veracity of collected information, and is calculated using NLP technology.

[1530] "Veracity" is the quality that determines whether information is based on actual, verified facts.

[1531] "User" refers to an individual or organization that performs operations to create documentary footage.

[1532] "Target audience" refers to the audience or target group that intends to watch the documentary footage.

[1533] "Story genre" is a classification that indicates the format and style of documentary footage.

[1534] "Emotions you want to convey" refers to the emotions and atmosphere you want viewers to feel through the documentary footage.

[1535] The "viewing conditions" are parameters related to the viewer set by the user, such as the viewing target, the genre of the story, and the emotion to be conveyed.

[1536] "Story structure" is the framework for designing the overall flow and development of a documentary film.

[1537] "The intention of each cut" refers to the purpose and message of each scene or cut in the video.

[1538] "Scenes" are the individual scenes or cuts that make up documentary footage.

[1539] "Video work" is a general term for the final documentary video.

[1540] "Rendering" is the process of converting an edited video clip into an output format as a single file.

[1541] "Output" refers to the act of saving or distributing the final video work in a specified format.

[1542] "User interface" refers to the screen and operating means through which a user interacts with and operates a system.

[1543] This invention relates to a system that efficiently generates and distributes documentary footage by allowing users to input themes and keywords. The system is configured as follows:

[1544] First, users input the documentary's theme and keywords through the smartphone application interface. For example, they can specify a theme such as "marine pollution." To accept this input, the system uses libraries such as React Native as its front end.

[1545] The server receives the input topic and keywords and collects related information from communication networks and internal databases. In this step, a crawler is used to collect text, images, video, and audio data from academic papers, news sites, government sites, and other sources on the Internet. The collected data is then stored in storage. The software used is Python and technologies such as Beautiful Soup and Scrapy.

[1546] Next, the reliability of the collected information is evaluated. TensorFlow or PyTorch is used to calculate a reliability score for the information using natural language processing (NLP) techniques. Furthermore, fake news and inaccurate information is filtered out using social media analysis tools. This analysis uses techniques such as sentiment analysis and named entity recognition (NER).

[1547] The evaluated information is then sorted into categories on the server, such as "scientific data," "news articles," "policy responses," and "personal interviews." The information is sorted into categories using labeling technology based on machine learning models.

[1548] Next, the user sets the target audience, story genre, and desired emotion on the application. These settings are necessary for the server to generate the optimal story structure. A template matching algorithm is used to select a story template and automatically generate the storyboard.

[1549] Based on the generated storyboard, the intention of each cut is set and the scene is composed. At this stage, the user inserts video clips, adds narration, and sets transition effects through the interface. The front-end editing function uses HTML5, CSS3, and JavaScript.

[1550] FFmpeg is used to render the final video and save it in the output format, which is then made available to users for download.

[1551] As a concrete example, consider a case where a user is creating a documentary on the theme of "marine pollution." When the user enters "marine pollution," the server collects the latest scientific data, news articles, and policy response information from the Internet. It evaluates the reliability of the collected information and categorizes it into categories such as "scientific data," "news," and "policy response." The user sets the target audience as general viewers, the story genre as documentary, and the emotion they want to convey as tension. The server composes the story based on a template and edits each cut through the user interface. Finally, the server integrates all the scenes and outputs the completed documentary video in MP4 format.

[1552] Examples of prompts include:

[1553] "A user wants to create a documentary about marine pollution. They need to gather the latest research, news, policy reports, etc., create a storyboard based on that information, and finally output it in MP4 format."

[1554] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1555] Step 1:

[1556] Users enter the documentary's theme and keywords into the smartphone application.

[1557] Input: Theme or keyword (e.g., marine pollution)

[1558] Output: Themes and keywords are sent to the server

[1559] Step 2:

[1560] The server collects related information from the communication network and internal database based on the input topic and keywords.

[1561] Input: Theme or keyword

[1562] Data processing: Using crawlers to collect information from academic papers, news sites, government sites, etc. on the Internet and store it in storage.

[1563] Output: Collected text, images, video, and audio data

[1564] Step 3:

[1565] The server uses NLP techniques to assess the reliability of the collected information.

[1566] Input: Collected information

[1567] Data Computation: Use TensorFlow or PyTorch to calculate the credibility score of information. Use social media analysis tools to filter out fake news and inaccurate information.

[1568] Output: reliable information

[1569] Step 4:

[1570] The server categorizes the information whose reliability has been evaluated.

[1571] Input: Information that has undergone a reliability assessment

[1572] Data processing: Using machine learning models, information is labeled into categories such as "scientific data," "news articles," "policy responses," and "personal interviews."

[1573] Output: Information sorted by category

[1574] Step 5:

[1575] Users set their viewing target, story genre, and the emotion they want to convey on the application.

[1576] Input: Target audience, story genre, and desired emotion (e.g., general audience, documentary, suspense)

[1577] Output: The configured viewing conditions are sent to the server.

[1578] Step 6:

[1579] The server selects the most suitable story template based on the set viewing conditions and generates a storyboard.

[1580] Input: Viewing conditions, information categorized by category

[1581] Data calculation: Using a template matching algorithm, a storyboard is automatically generated by selecting the story template that best suits the viewing conditions.

[1582] Output: Generated storyboard

[1583] Step 7:

[1584] The user sets the intention for each cut based on the generated storyboard and composes the scene.

[1585] Input: Generated storyboard

[1586] Specific operations: The user inserts a video clip, adds narration, and sets transition effects in the application interface.

[1587] Output: Edited storyboard

[1588] Step 8:

[1589] The server renders the final video production and outputs it in the specified format.

[1590] Input: Edited storyboard

[1591] Specific operation: All scenes are integrated using FFmpeg and the final video is rendered in MP4 format or similar.

[1592] Output: The finished documentary footage is made available to users in a downloadable format

[1593] As an example, the following prompt sentence is used:

[1594] "A user wants to create a documentary about marine pollution. They need to gather the latest research, news, policy reports, etc., create a storyboard based on that information, and finally output it in MP4 format."

[1595] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1596] The present invention is a system that collects related information based on specific themes and keywords, evaluates and classifies reliability, and generates high-quality documentary footage by recognizing the user's viewing preferences and emotions.

[1597] System Configuration

[1598] 1. How to input themes and keywords

[1599] The user inputs the documentary's theme or keywords. For example, "marine pollution." The user interface (UI) is designed to be intuitive.

[1600] 2. Information gathering methods

[1601] The server collects relevant information from the internet and internal databases using web scraping and API integration to obtain text, images, video, and audio data.

[1602] 3. Reliability assessment methods

[1603] The server uses natural language processing (NLP) techniques to assess the reliability of collected information, analyzing its source, number of citations, and related research data to calculate a credibility score. It also uses social media analysis tools to identify and filter out fake news.

[1604] 4. Information classification means

[1605] The server categorizes the evaluated information into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and automatically labels each piece of information with the appropriate category.

[1606] 5. Viewing setting input method

[1607] The user sets the target audience (e.g., general audience), the genre of the story (e.g., documentary), and the emotion they want to convey (e.g., tension). These settings are made through a user interface.

[1608] 6. Emotion Engine

[1609] The emotion engine analyzes the user's emotions in real time as they type, and reflects this in the story structure and scene settings. Emotion analysis is performed based on the user's facial expressions, tone of voice, and typing speed.

[1610] 7. Story Generation Methods

[1611] The server selects the optimal story template based on the viewing preferences and the analysis results of the emotion engine, and generates a story structure. Using a template matching algorithm, the collected information is arranged on the storyboard.

[1612] 8. Scene Composition Methods

[1613] The device allows the user to compose scenes by setting the intention for each cut. Using drag and drop, the user places the appropriate video clips for each cut, adds narration, and sets transition effects. Based on the emotional state recognized by the emotion engine, appropriate narration and music are automatically selected and placed.

[1614] 9. Video rendering and output methods

[1615] The server stitches all the scenes together and renders the final documentary video. The video file is generated in the specified output format (e.g. MP4, MOV) and made available to the user for download or viewing.

[1616] Specific examples

[1617] If a user wants to create a documentary on the theme of "marine pollution," they first input the theme. The server collects related information and evaluates and classifies it for reliability. The user then selects a general audience, a documentary genre, and a desire to convey a sense of tension. The emotion engine recognizes the user's sense of tension as they input, and reflects this in the story structure and scene settings. Appropriate video clips, narration, and music are selected, and finally, the server renders all the scenes, outputting the finished documentary video.

[1618] In this way, the present invention provides a system that efficiently generates high-quality documentary footage while taking into account the user's emotions.

[1619] The processing flow will be explained below.

[1620] Step 1:

[1621] The user inputs the documentary's theme or keywords. For example, "marine pollution."

[1622] Step 2:

[1623] The server collects information based on themes and keywords. The information is collected using web scraping and API integration to obtain text, images, video, and audio data from the internet and internal databases.

[1624] Step 3:

[1625] The server evaluates the credibility of the collected information, using natural language processing (NLP) techniques to analyze the information's source, number of citations, and related research data to calculate a credibility score, and also uses social media analysis tools to identify and remove fake news and inaccurate information.

[1626] Step 4:

[1627] The server categorizes the evaluated information into categories, such as "scientific data," "news articles," "policy reports," and "personal interviews," and each piece of information is automatically labeled with the appropriate category.

[1628] Step 5:

[1629] The user sets the target audience, the genre of the story, and the emotion they want to convey. These settings are entered through a highly visible user interface. For example, the target audience can be set to general viewers, the genre to documentary, and the emotion they want to convey to be tension.

[1630] Step 6:

[1631] The emotion engine analyzes the user's emotions in real time, analyzing facial expressions, tone of voice, and input speed when entering settings to recognize the user's current emotional state.

[1632] Step 7:

[1633] The server selects the optimal story template based on the user's viewing preferences and the analysis results of the emotion engine. Using a template matching algorithm, it selects the template that best suits the preferences and analysis data, and places the collected information on the storyboard.

[1634] Step 8:

[1635] The device sets the intention for each cut and composes the scene. The user uses drag and drop to place the appropriate video clips for each cut, add narration, and set transition effects. Appropriate narration and music are also automatically selected and placed based on the emotional state recognized by the emotion engine.

[1636] Step 9:

[1637] The server stitches all the scenes together and renders the final documentary video. The video file is generated in the specified output format (e.g. MP4, MOV) and made available to the user for download or viewing.

[1638] By following these steps, users can efficiently generate high-quality documentary footage that reflects emotions.

[1639] Example 2

[1640] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1641] In conventional documentary video production systems, the processes of information collection, credibility assessment, information classification, viewing settings, emotion analysis, story generation, scene composition, and video rendering are all performed manually, which is extremely time-consuming and labor-intensive. Furthermore, they lack the ability to analyze user emotions in real time and reflect them in video production, resulting in a lack of emotional impact on viewers. Furthermore, with the rise of fake news, unreliable information is often mixed in, making it difficult to produce high-quality documentary videos based on accurate information.

[1642] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for accepting input of themes and keywords, means for collecting information from a network and an internal database, means for evaluating the reliability of the collected information and determining its authenticity, means for identifying fake information using social media analysis technology, means for categorizing the evaluated information, means for the user to set the viewing target, story genre, and emotion to be conveyed, means for analyzing the user's emotions in real time using an emotion analysis means and reflecting the results in scene settings, means for selecting a story template optimal for the viewing conditions using a template matching algorithm, means for setting the intention of each scene based on the generated story structure and composing scenes, and means for synthesizing and outputting the final documentary video. This enables automatic collection and reliability evaluation of information, appropriate information classification, real-time analysis of user emotions, and automatic selection of the optimal story template, thereby enabling efficient generation of high-quality documentary video with emotional impact.

[1643] The "means for accepting input of themes and keywords" is a means for providing an interface that allows a user to input themes and keywords necessary for producing a documentary video.

[1644] "Means for collecting information from networks and internal databases" refers to means for automatically collecting information related to the entered topic or keyword via the Internet or internal databases.

[1645] "Means for assessing the reliability of collected information and determining its authenticity" refers to means for assessing the reliability of collected information using natural language processing technology and source analysis, and determining the authenticity of the information.

[1646] "Methods for identifying fake information using social media analysis technology" refers to methods that use technology to analyze data on social media and identify misinformation and fake news.

[1647] A "means of categorizing evaluated information" is a means of classifying information that has been collected and evaluated for reliability into specific categories, such as "scientific data" or "news articles."

[1648] "A means for users to set the target audience, story genre, and emotions they want to convey" is a means for providing an interface that allows users to set the audience, genre, and emotions they want to convey for the documentary footage.

[1649] "Means for analyzing a user's emotions in real time using emotion analysis means and reflecting the emotions in scene settings" refers to means for analyzing a user's facial expressions and tone of voice, grasping the user's emotional state in real time, and reflecting the results in scene settings.

[1650] "Means for selecting a story template optimal for viewing conditions using a template matching algorithm" refers to means for executing an algorithm that selects the optimal template from among pre-prepared story templates based on the user's viewing settings and emotion analysis results.

[1651] "Means for setting the intention of each scene based on the generated story structure and composing scenes" refers to means for setting the purpose and intention of each scene based on the generated story and arranging corresponding images and narration.

[1652] The "means for compositing and outputting the final documentary footage" refers to a means for combining all the scenes into one, generating the final documentary footage, and outputting it in a specified format.

[1653] This invention is a system that collects related information based on specific themes and keywords, evaluates and classifies its reliability, recognizes the user's viewing preferences and emotions, and automatically generates high-quality documentary footage.

[1654] System Configuration

[1655] The system has the following main components:

[1656] 1. How to input themes and keywords

[1657] Users input the documentary's theme and keywords through a user interface (UI), which is designed for intuitive operation.

[1658] 2. Information gathering methods

[1659] The server collects relevant information from the internet and internal databases using web scraping tools (e.g., BeautifulSoup) and API integration (e.g., Twitter API), which allows it to obtain text, images, video, and audio data.

[1660] 3. Reliability assessment methods

[1661] The server evaluates the reliability of the collected information using natural language processing (NLP) techniques (e.g., BERT, GPT-3), and also identifies and eliminates fake information using social media analysis techniques.

[1662] 4. Information classification means

[1663] The server categorizes the evaluated information using a clustering algorithm (e.g., K-means) to automatically categorize it into categories such as "scientific data," "news articles," "policy reports," and "personal interviews."

[1664] 5. Viewing setting input method

[1665] Users can set their target audience, the genre of the story, and the emotion they want to convey through a user interface.

[1666] 6. Emotion analysis method

[1667] The server is equipped with an emotion analysis engine that analyzes the user's emotions in real time using software that analyzes facial expressions (e.g., OpenCV) and voice tone (e.g., Deepgram), and also combines the user's input speed for analysis.

[1668] 7. Story Generation Methods

[1669] The server runs a template matching algorithm to select the most appropriate story template based on viewing preferences and sentiment analysis results, and places the collected information into the appropriate storyboard.

[1670] 8. Scene Composition Methods

[1671] The user uses the drag-and-drop function on the device to place appropriate video clips for each cut on a storyboard, and then sets narration and transition effects. Appropriate narration and music are automatically selected based on the emotional state recognized by the emotion analysis means.

[1672] 9. Video rendering and output methods

[1673] The server combines all the scenes to generate the final documentary video and renders it as a video file in the specified format, such as MP4 or MOV.

[1674] Specific examples

[1675] Consider a scenario where a user is creating a documentary on the theme of "marine pollution." First, the user inputs the theme. The server collects related information, evaluates its reliability, and classifies it. Next, the user sets the target audience as "general audience," the genre as "documentary," and the emotion they want to convey as "tension." The emotion analysis engine recognizes the user's emotions in real time as they are input, and reflects them in the story structure and scene settings. Appropriate video clips, narration, and music are selected, and the server renders all scenes to output the completed documentary video.

[1676] Prompt Sentence Examples

[1677] "Collect information to create a documentary film on the following topic, assess its reliability, classify it, generate a story, and output the final film. The user settings are for general audiences, the genre is documentary, and the emotion you want to convey is a sense of urgency. The topic is 'Marine Pollution.'"

[1678] This system makes it possible to efficiently generate high-quality documentary footage while taking into account the user's emotions.

[1679] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1680] Step 1:

[1681] Accepts input of themes and keywords.

[1682] The user inputs themes and keywords through a user interface (UI).

[1683] Specific behavior: The user enters "Marine Pollution" in the topic field and clicks the submit button.

[1684] Input: Theme or keyword (e.g. "Marine pollution")

[1685] Output: Themes and keywords entered

[1686] Step 2:

[1687] Automatic Collection of Information.

[1688] The server collects information from the Internet and internal databases based on the entered themes and keywords.

[1689] Specific operation: The server uses a web scraping tool (e.g., BeautifulSoup) to retrieve articles from Google search results and uses an API (e.g., Twitter API) to collect related tweets.

[1690] Input: Theme or keyword (e.g. "Marine pollution")

[1691] Output: Collected information (text, images, video, audio data)

[1692] Step 3:

[1693] Evaluating the reliability of information.

[1694] The server evaluates the reliability of the collected information using natural language processing (NLP) techniques.

[1695] How it works: The server uses NLP models (e.g., BERT, GPT-3) to calculate a credibility score for each article, and also uses social media analysis techniques to identify and remove misinformation.

[1696] Input: Collected information

[1697] Output: Information with assessed reliability (with score)

[1698] Step 4:

[1699] Classification of information.

[1700] The server categorizes the evaluated information.

[1701] What it does: The server uses a clustering algorithm (e.g., K-means) to classify information into categories such as "scientific data," "news articles," "policy reports," and "personal interviews."

[1702] Input: Information with reliability assessment

[1703] Output: Categorically labeled information

[1704] Step 5:

[1705] Enter viewing settings.

[1706] The user inputs the target audience, the genre of the story, and the emotion they want to convey.

[1707] Specific operation: The user selects the target audience as "general audience," the genre as "documentary," and the emotion to be conveyed as "tension."

[1708] Input: target audience, story genre, emotion you want to convey

[1709] Output: Viewing Settings

[1710] Step 6:

[1711] Performing sentiment analysis.

[1712] The server analyzes the user's emotions in real time.

[1713] How it works: The server uses facial expression analysis software (e.g., OpenCV) and voice analysis software (e.g., Deepgram) to detect the user's emotions. The user's typing speed is also used for analysis.

[1714] Input: User's facial expression data, tone of voice, typing speed

[1715] Output: Real-time analyzed emotion data

[1716] Step 7:

[1717] Story generation.

[1718] The server selects a story template based on viewing preferences and emotion analysis results, and generates a story structure.

[1719] Specific operation: The server uses a template matching algorithm to select an appropriate story template and places the collected information on the storyboard.

[1720] Input: Viewing settings, sentiment analysis results, and information categorized by category

[1721] Output: Generated story structure

[1722] Step 8:

[1723] Scene placement and adjustment.

[1724] The terminal allows the user to set the intention for each cut and compose a scene.

[1725] How it works: Users use the drag-and-drop function to place video clips on the storyboard and set narration and transition effects. Appropriate narration and music are automatically selected based on emotional data.

[1726] Input: Generated story structure, video clips, narration, transition effects

[1727] Output: Composed Scene

[1728] Step 9:

[1729] Video rendering and output.

[1730] The server stitches all the scenes together and renders the final documentary footage.

[1731] What it does: The server uses video editing software (e.g., FFmpeg) to combine and render all the scenes into a single video file in MP4 or MOV format.

[1732] Input: Composed Scene

[1733] Output: Final documentary footage file

[1734] (Application example 2)

[1735] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1736] In today's world, users are seeking to generate and view high-quality documentary videos based on their own interests. However, conventional methods have difficulty collecting reliable information and providing it in an optimal format based on the viewer's emotions and preferences. Furthermore, there is no system that can analyze the viewer's emotional state in real time and dynamically change the video content accordingly, which hinders user satisfaction.

[1737] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1738] In this invention, the server includes means for accepting input of themes and keywords, means for collecting information related to the themes and keywords from a communication network and an internal database, means for evaluating the reliability of the collected information and determining its authenticity, means for classifying the evaluated information by type, means for the user to set the viewing target, story genre, and emotion to be conveyed, means for analyzing the user's emotional state in real time, means for generating a story structure based on the set viewing conditions and the analyzed emotional state, means for setting the intention of each cut and composing scenes based on the generated story structure, and means for rendering and outputting the final documentary video. This makes it possible to generate and provide high-quality, reliable documentary video in real time based on the viewing conditions and emotional state set by the user.

[1739] The "means for accepting input of themes and keywords" is an interface that allows a user to input subjects of interest or specific words into the system.

[1740] "Means for collecting information related to the theme or keyword from communication networks and internal databases" refers to a function for collecting information from the Internet and obtaining related information from internally stored databases.

[1741] "Means for assessing the reliability of collected information and determining its authenticity" refers to techniques for analyzing the reliability of the data obtained and determining whether the data is accurate.

[1742] The "means for classifying evaluated information by type" refers to a method for classifying information whose reliability has been evaluated into different categories based on its content.

[1743] "A means for users to set the target audience, story genre, and emotions they want to convey" is an interface that allows users to set the target audience, video genre, and emotions they want to convey through the video.

[1744] The "means for analyzing the user's emotional state in real time" is a technology for detecting and analyzing the user's current emotions in real time.

[1745] The "means for generating a story structure based on the set viewing conditions and the analyzed emotional state" is a method for automatically constructing an optimal story based on the viewing conditions set by the user and the analyzed emotions.

[1746] "Means for setting the intention of each cut based on the generated story structure and composing scenes" refers to a technique for setting the content and purpose of each scene based on the generated story structure and then visualizing it.

[1747] "Means for rendering and outputting the final documentary footage" refers to the function for integrating all scenes to generate the final footage and outputting it in a specified format.

[1748] The present invention is a system for generating high-quality documentary videos based on user viewing preferences and real-time emotion analysis. The system operates by combining multiple hardware and software technologies and is implemented in the following configuration.

[1749] Hardware and software used

[1750] Server: Responsible for information gathering, reliability assessment, data classification, sentiment analysis, story generation, scene composition, and rendering.

[1751] Smart devices (smartphones, smart TVs): Collect data on user input, viewing, and sentiment analysis.

[1752] Communications Network: Collecting and distributing data through internet connections and internal networks.

[1753] Specific software technologies used include:

[1754] Natural Language Processing (NLP) engine: Used to understand and analyze text and assess the reliability of information.

[1755] Machine learning model: Used for user sentiment analysis.

[1756] Template matching algorithm: Used to select the best story template.

[1757] Rendering engine: Used to generate the final documentary footage.

[1758] Cloud storage and databases: Used to store information and manage access.

[1759] Specific examples of the system

[1760] 1. Enter your topic or keywords:

[1761] Users input a theme or keyword, such as "marine pollution," through the interface of their smart device.

[1762] The interface is designed to be intuitive.

[1763] Example: User types "ocean pollution".

[1764] 2. Information Collection:

[1765] The server collects relevant information from the Internet and an internal database via a communication network.

[1766] Using web scraping and API integration, text, images, videos, audio data, etc. are obtained.

[1767] Example: A server collects academic papers, news articles, and visual materials on "ocean pollution."

[1768] 3. Reliability assessment:

[1769] The server uses NLP techniques to evaluate the reliability of the collected information.

[1770] The reliability score is calculated by analyzing the source of the information, the number of citations, and related research data.

[1771] Use social media analytics tools to filter out fake news.

[1772] Example: An NLP engine evaluates the reliability of collected news articles and filters out articles with low scores.

[1773] 4. Information classification:

[1774] The server categorizes the evaluated information.

[1775] Categories include "scientific data," "news articles," "policy reports," and "personal interviews."

[1776] Example: Collected information is divided into categories such as "scientific data" and "personal interviews."

[1777] 5. Viewing Preferences and Sentiment Analysis:

[1778] Users set the target audience, story genre, and the emotion they want to convey.

[1779] The server uses machine learning models to analyze emotions in real time based on the user's facial expressions and input speed collected from the smart device.

[1780] Example: User selects the documentary genre and wants a thrilling story.

[1781] 6. Story generation and scene composition:

[1782] The server generates an optimal story based on the set viewing conditions and the analyzed emotional state.

[1783] A template matching algorithm is used to select an appropriate story template.

[1784] The intention of each cut is set based on the generated story structure, and scenes are composed.

[1785] Example: The server automatically generates scenes with the theme of "tension" based on emotion analysis.

[1786] 7. Video rendering and output:

[1787] The server renders the final documentary footage and outputs it in the format you specify (e.g. MP4, MOV).

[1788] Users can watch on their smart devices or web browsers.

[1789] Example: A completed documentary film is streamed to a smart TV.

[1790] Prompt Sentence Examples

[1791] Please enter a topic (e.g., "Marine Pollution").

[1792] "Assessing reliability of visual data..."

[1793] "Analyzing user sentiment... Anything else you'd like to add?"

[1794] As a result, the present invention provides a system that efficiently generates and provides high-quality documentary footage while taking into consideration the user's emotions.

[1795] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1796] Step 1:

[1797] The user inputs themes and keywords through the interface of the smart device.

[1798] Input: User types in the topic "marine pollution."

[1799] Specific operation: The smart device interface receives input from the user and sends themes and keywords to the server.

[1800] Step 2:

[1801] The server collects information related to themes and keywords from the communication network and an internal database.

[1802] Input: Query regarding "marine pollution."

[1803] Data processing: Use web scraping and API integration to obtain relevant information.

[1804] Output: Collected text, images, video and audio data.

[1805] What it does: The server collects academic papers, news articles, visual materials, etc. related to "marine pollution."

[1806] Step 3:

[1807] The server evaluates the reliability of the collected information using natural language processing (NLP) techniques.

[1808] Input: Collected information.

[1809] Data calculations: Analysis of sources, citations, and related research data. Calculation of reliability scores.

[1810] Output: Information rated as reliable and its score.

[1811] How it works: The NLP engine evaluates the credibility of news articles and papers and assigns them a credibility score.

[1812] Step 4:

[1813] The server classifies the information whose reliability has been evaluated by type.

[1814] Input: Reliability-assessed information.

[1815] Data processing: Using classification algorithms, information is sorted into categories such as "scientific data," "news articles," "policy reports," and "personal interviews."

[1816] Output: Information broken down by category.

[1817] What it does: The server automatically categorizes collected information into appropriate categories based on trustworthiness.

[1818] Step 5:

[1819] Users set the target audience, story genre, and the emotion they want to convey through the interface of their smart device.

[1820] Input: User viewing preferences (e.g., "General Audience," "Documentary," "Intense").

[1821] Output: The set viewing conditions.

[1822] Specific operation: The user operates the smart device and sets up viewing.

[1823] Step 6:

[1824] The server uses a machine learning model to analyze emotions in real time based on the user's facial expressions and input speed collected from the smart device.

[1825] Input: Real-time facial expression data and input speed of the user.

[1826] Data computation: Sentiment analysis with machine learning models.

[1827] Output: User's emotional state (e.g., sense of urgency).

[1828] Specific operation: The server analyzes the user's emotions in real time, and collects and processes the data.

[1829] Step 7:

[1830] The server generates an optimal story based on the set viewing conditions and the analyzed emotional state.

[1831] Input: User's viewing preferences and emotional state, information broken down into categories.

[1832] Data calculation: A template matching algorithm is used to select the most suitable story template.

[1833] Output: The generated story structure.

[1834] What happens: The server uses a template matching algorithm to select an appropriate story template and generate a story.

[1835] Step 8:

[1836] Based on the story structure generated by the server, the intention of each cut is set and scenes are composed.

[1837] Input: The generated story structure.

[1838] Data processing: Automatically place appropriate video clips, narration, and music.

[1839] Output: The composed scene.

[1840] Specific operation: The server arranges video clips and music based on the automatically generated story structure to create scenes.

[1841] Step 9:

[1842] The server renders and outputs the final documentary footage.

[1843] Input: The composed scene.

[1844] Data calculation: Generate the final image using the rendering engine.

[1845] Output: Final documentary footage (e.g. MP4, MOV).

[1846] What it does: The server stitches all the scenes together, renders the final documentary footage, and outputs it in the specified format.

[1847] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1848] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1849] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1850] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1851] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1852] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1853] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1854] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1855] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1856] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1857] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1858] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1859] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1860] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1861] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1862] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1863] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1864] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1865] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1866] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1867] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1868] The following is further disclosed regarding the above embodiment.

[1869] (Claim 1)

[1870] a means for accepting input of themes and keywords;

[1871] A means of collecting information related to the theme or keyword from the Internet and internal databases;

[1872] A means of assessing the reliability of the collected information and determining its authenticity;

[1873] a means of categorizing the assessed information;

[1874] A way for users to set the audience, story genre, and emotion they want to convey;

[1875] A means for generating a story structure based on set viewing conditions;

[1876] A means to set the intention of each cut based on the generated story structure and compose the scene;

[1877] A means to render and output the final documentary footage;

[1878] A system including:

[1879] (Claim 2)

[1880] 10. The system of claim 1, further comprising means for identifying fake information using social media analysis tools.

[1881] (Claim 3)

[1882] 10. The system of claim 1, further comprising means for selecting a story template that best suits viewing conditions using a template matching algorithm.

[1883] "Example 1"

[1884] (Claim 1)

[1885] a means for accepting input of themes and keywords;

[1886] A means for collecting information related to the theme or keyword from networks and databases;

[1887] A method for evaluating the reliability of collected information using natural language processing and determining its authenticity;

[1888] a means of categorizing the assessed information;

[1889] A way for users to set the audience, story genre, and emotion they want to convey;

[1890] A means for generating a story structure based on set viewing conditions;

[1891] A means to set the intention of each cut based on the generated story structure and compose the scene;

[1892] a means for rendering and outputting the final video content;

[1893] A system including:

[1894] (Claim 2)

[1895] 10. The system of claim 1, further comprising means for identifying inaccurate information using social media analysis tools.

[1896] (Claim 3)

[1897] 10. The system of claim 1, further comprising means for selecting a story template that best suits viewing conditions using a template matching algorithm.

[1898] "Application Example 1"

[1899] (Claim 1)

[1900] a means for accepting input of themes and keywords;

[1901] means for collecting information related to the subject or keyword from communication networks and internal databases;

[1902] a means of assessing the reliability and determining the veracity of the collected information;

[1903] a means of categorizing the assessed information;

[1904] A way for users to set the audience, story genre, and emotion they want to convey;

[1905] A means for generating a story structure based on set viewing conditions;

[1906] A means to set the intention of each cut based on the generated story structure and compose the scene;

[1907] A means to render and output the final video work;

[1908] A means for providing a user interface that allows users to input a theme and automatically generate and distribute documentary footage based on the collected information;

[1909] A system including:

[1910] (Claim 2)

[1911] 10. The system of claim 1, further comprising means for filtering out misinformation using social media analysis tools.

[1912] (Claim 3)

[1913] 10. The system of claim 1, further comprising means for selecting a story template that best suits viewing conditions using a template matching algorithm.

[1914] "Example 2: Combining Emotion Engines"

[1915] (Claim 1)

[1916] a means for accepting input of themes and keywords;

[1917] A means for collecting information related to the subject or keyword from the network and internal database;

[1918] A means of assessing the reliability of the collected information and determining its authenticity;

[1919] a means of categorizing the assessed information;

[1920] A way for users to set the audience, story genre, and emotion they want to convey;

[1921] A means for generating a story structure based on set viewing conditions;

[1922] A means for setting the intention of each scene based on the generated story structure and composing the scene;

[1923] A means for analyzing a user's emotions in real time using an emotion analysis means and reflecting the emotions in the scene settings;

[1924] A means to composite and output the final documentary footage;

[1925] A system including:

[1926] (Claim 2)

[1927] 10. The system of claim 1, further comprising means for identifying fake information using social media analysis techniques.

[1928] (Claim 3)

[1929] 10. The system of claim 1, further comprising means for selecting a story template that best suits viewing conditions using a template matching algorithm.

[1930] "Application example 2 when combining emotion engines"

[1931] (Claim 1)

[1932] a means for accepting input of themes and keywords;

[1933] means for collecting information related to the subject or keyword from communication networks and internal databases;

[1934] A means of assessing the reliability of the collected information and determining its authenticity;

[1935] a means for classifying the evaluated information by type;

[1936] A way for users to set the audience, story genre, and emotion they want to convey;

[1937] means for analyzing the user's emotional state in real time;

[1938] means for generating a story structure based on the set viewing conditions and the analyzed emotional state;

[1939] A means to set the intention of each cut based on the generated story structure and compose the scene;

[1940] A means to render and output the final documentary footage;

[1941] A system including:

[1942] (Claim 2)

[1943] 10. The system of claim 1, further comprising means for identifying fake information using social media analysis tools.

[1944] (Claim 3)

[1945] 10. The system of claim 1, further comprising means for selecting a story template that best suits the viewing conditions and emotional state of the user using a template matching algorithm. [Explanation of symbols]

[1946] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for accepting input of themes and keywords; A means of collecting information related to the theme or keyword from the Internet and internal databases; A means of assessing the reliability of the collected information and determining its authenticity; a means of categorizing the assessed information; A way for users to set the audience, story genre, and emotion they want to convey; A means for generating a story structure based on set viewing conditions; A means to set the intention of each cut based on the generated story structure and compose the scene; A means to render and output the final documentary footage; A system including:

2. The system of claim 1 , further comprising means for identifying fake information using social media analysis tools.

3. 10. The system of claim 1, further comprising means for selecting a story template that best suits viewing conditions using a template matching algorithm.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A