system
A system that analyzes novel text to generate promotional videos, addressing consumer uncertainty by providing a visual experience, enhancing purchase decisions and simplifying promotional efforts.
Patent Information
- Application Number
- JP2024140506
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-03-06
AI Technical Summary
Consumers hesitate to purchase novels due to uncertainty about the story content, lacking a way to experience the atmosphere and story beforehand.
A system that receives input text data, performs natural language analysis, generates a promotional video integrating specified visual elements, and displays it on a designated device to provide a visual experience of the novel's opening.
Enables consumers to visually experience the novel's atmosphere, aiding purchase decisions and reducing the effort required for promotional material creation.
Smart Images

Figure 2026037481000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When buying a novel at a bookstore, some people hesitate to buy it because they don't really know what the story is about, or they end up feeling like they made a mistake. There is a need for a way to solve this problem and let consumers get a feel for the atmosphere and story of a novel beforehand. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by providing a system including: means for receiving input text data; means for natural language analysis of the received text data; means for generating a promotional video by integrating specified visual elements based on the analyzed text data; means for transmitting the generated promotional video to a predetermined display device; and means for playing the transmitted promotional video. This allows consumers to visually experience the opening part of a novel as a video, which can be used as information for making a purchase decision.
[0006] "Input text data" refers to data of a part or the entire text of a novel or the like that is input by a user into the system.
[0007] "Means for receiving" refers to a device or software that has the function of receiving data from the outside.
[0008] "Natural language analysis" is a technology that linguistically analyzes input text and understands its meaning and structure.
[0009] "Visual elements" are visual elements such as images, videos, text, and animations.
[0010] A "promotional video" is a short video created for a specific purpose that visually conveys the content and atmosphere of a novel.
[0011] The "predetermined display device" is a display device such as a digital signage or a pop-up screen that has been designated in advance.
[0012] A "transmitting means" is a device or software that has the function of sending data to other devices or terminals.
[0013] The "playback means" is a function that displays the received video data as a moving image on a display device. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] The system for implementing this invention receives text data of a novel entered by a user, analyzes the text, generates a promotional video, and finally displays it on a digital signage or pop-up screen. This system is mainly composed of three elements: a server, a terminal, and a display device.
[0036] System configuration
[0037] 1. User Input
[0038] A user uses a device (e.g., a PC, tablet, or smartphone) to input text data, such as the opening of a novel, which is then entered into the device through a dedicated web form or application.
[0039] User: Enter the opening line of a novel and click the submit button.
[0040] 2. Sending text data
[0041] The user's device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[0042] Terminal: Sends text data to the server.
[0043] 3. Text Analysis
[0044] The server uses natural language processing (NLP) algorithms to analyze the received text data. NLP analyzes the content of the text and extracts key elements (e.g., characters, setting, theme, etc.).
[0045] Server: Analyzes the text data and extracts key elements.
[0046] 4. Image Generation
[0047] Based on the analysis results, the server creates a storyboard and selects corresponding visual elements (images, animations, text, etc.), then uses a video generation engine to render the promotional video.
[0048] Server: Creates storyboards, selects visual elements, and generates footage.
[0049] 5. Video data transmission
[0050] The server encodes the generated promotional video data and transmits it to the designated display device, also using a secure communication protocol.
[0051] Server: Encodes video data and sends it to the display device.
[0052] 6. Video display
[0053] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video.
[0054] Terminal: Displays the received video data.
[0055] Specific examples
[0056] A user uses an interactive terminal installed in a bookstore to input the opening part of the novel "Adventure in the Mysterious Forest."
[0057] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0058] The terminal transmits this text data to the server.
[0059] The server analyzes the text and extracts the keywords "boy," "mysterious forest," "animals," and "adventure."
[0060] The server creates a storyboard based on these keywords and selects appropriate visual elements to generate the promotional video.
[0061] The generated promotional video is sent to the bookstore's digital signage and displayed.
[0062] This system allows bookstore customers to visually get a feel for the atmosphere of a novel in advance and use it as a reference when making a purchase.
[0063] The processing flow will be explained below.
[0064] Step 1:
[0065] The user accesses a dedicated web form or application using a terminal, enters text data to be used in the promotional video, such as the opening of a novel, and clicks the submit button.
[0066] Step 2:
[0067] The device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[0068] Step 3:
[0069] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.).
[0070] Step 4:
[0071] The server creates a storyboard based on the analysis results, which includes information on the scenes and characters required for the promotional video.
[0072] Step 5:
[0073] The server selects corresponding visual elements (images, animations, text, etc.) according to the storyboard. These visual elements are integrated by the video generation engine and rendered as a promotional video.
[0074] Step 6:
[0075] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or pop-up screen). This transmission also uses a secure communication protocol.
[0076] Step 7:
[0077] The display device decodes the received video data and plays it as a promotional video, allowing bookstore visitors to visually experience the opening part of the novel.
[0078] Example 1
[0079] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0080] Traditionally, video generation for promoting novels and stories has been done manually, requiring a great deal of time and effort. Furthermore, specialized knowledge is required to select appropriate visual elements based on the content, making it difficult for anyone to easily create promotional videos. There has also been a lack of technology to instantly display generated videos on digital signage or pop-up screens. There is a need to solve these issues and automate the effective promotion of novels and stories quickly.
[0081] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0082] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for generating a promotional video by integrating specified visual elements based on the analyzed text data, means for transmitting the generated promotional video to a predetermined display device, means for playing the transmitted promotional video, and means for displaying the promotional video on a digital signage or pop-up screen. This makes it possible to automatically generate a promotional video based on the content of a novel or story and display it on a digital signage or pop-up screen quickly and effectively.
[0083] The "means for receiving input text data" is a function for transmitting text data input by a user using a terminal to a server and receiving the data.
[0084] "Means for natural language analysis" refers to algorithms or programs that analyze input text data, understand its content, and extract key elements and keywords.
[0085] The "means for generating promotional videos" is a function that combines the necessary visual elements based on the analyzed text data to create videos.
[0086] The "means for transmitting to a display device" is a communication function for encoding the generated promotional video data into an appropriate format and sending it to a display device such as a digital signage or a pop-up screen.
[0087] The "means for playing promotional video" is a function for the display device to decode the video data received, convert it into a playable format, and display it.
[0088] "Means for displaying on digital signage or a pop-up screen" refers to a function for actually displaying the generated promotional video on a digital signage or a pop-up screen.
[0089] "Major elements" are important items such as characters, settings, and themes that are extracted from the input text data.
[0090] "Visual elements" are visual elements such as images, animations, and text that make up the promotional video.
[0091] This invention is a system that inputs text data for a novel or story, analyzes its content, automatically generates a promotional video, and displays it on a display device. This system is mainly composed of three elements: a server, a terminal, and a display device. Each element works together to automate the promotional video generation process.
[0092] Hardware and software used
[0093] server:
[0094] Natural language processing is performed using Python's "spaCy" library and "Hugging Face"'s "Transformers" library.
[0095] Generate promotional videos using Unity and Blender.
[0096] Device:
[0097] A device such as a computer, tablet, or smartphone that accepts user input.
[0098] The ability to enter text data using a dedicated web form or application and click a submit button.
[0099] Display device:
[0100] Display devices such as digital signage and pop-up screens.
[0101] A decoding library and video player application for decoding and displaying received promotional videos.
[0102] System Operation
[0103] User Input
[0104] A user uses a terminal to input text data, such as the opening of a novel, which is received through a dedicated web form or application.
[0105] example:
[0106] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0107]
[0108] User: Enter the opening line of the novel and click the submit button.
[0109] Server processing
[0110] The server analyzes the received text data using natural language processing algorithms. Through this analysis, key elements such as characters, setting, and theme are extracted. A storyboard is then created based on the analysis results, and appropriate visual elements are selected. A promotional video is then generated based on the storyboard using Unity or Blender.
[0111] Transmission to a display device and display
[0112] The generated promotional video is encoded and transmitted from the server to the display device, which can then decode the received video data and play it back as the promotional video.
[0113] Specific examples
[0114] As an example, we will explain the creation of a promotional video for the novel "Adventure in the Mysterious Forest" using the following prompt sentence.
[0115] Example prompt sentence:
[0116] Prompt: Generate a promotional video for the novel "Adventure in the Mysterious Forest." Analyze the text below to extract characters, setting, and themes, and create a video with appropriate visual elements.
[0117] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0118] This system allows bookstore customers to visually get a feel for the atmosphere of a novel beforehand, which they can use as a reference when making a purchase. It also makes it possible to significantly reduce the effort required to create promotional materials.
[0119] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0120] Step 1:
[0121] The user uses a terminal to input the opening part of a novel. The input text data is received by a dedicated web form or application. Specifically, the user inputs the data in Unicode text format and clicks the submit button. The input data format is encoded in UTF-8.
[0122] input:
[0123] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0124] output:
[0125] UTF-8 encoded text data.
[0126] Step 2:
[0127] The device sends the entered text data to the server as an HTTP POST request. This transmission uses the HTTPS protocol to maintain data integrity. Specifically, the request body contains the text data, and the URL is the server's API endpoint.
[0128] input:
[0129] UTF-8 encoded text data.
[0130] output:
[0131] The HTTP POST request sent to the server.
[0132] Step 3:
[0133] The server analyzes the received text data using natural language processing (NLP) algorithms, such as the Python "spaCy" library and the "Transformers" library from "Hugging Face," to analyze the text content and extract key elements (e.g., characters, setting, and theme).
[0134] input:
[0135] The text data sent in the HTTP POST request.
[0136] output:
[0137] A list of the main elements (characters, setting, theme).
[0138] Step 4:
[0139] The server creates a storyboard based on the analyzed text data, determining the order and structure of scenes based on the extracted key elements, and selecting appropriate visual elements (images, animations, text, etc.).
[0140] input:
[0141] A list of the main elements (characters, setting, theme).
[0142] output:
[0143] Storyboard and selected visual elements.
[0144] Step 5:
[0145] The server generates promotional videos using Unity, Blender, etc. It combines storyboards and visual elements, renders each scene, and creates the final video file.
[0146] input:
[0147] Storyboards and visual elements.
[0148] output:
[0149] Generated promotional video file.
[0150] Step 6:
[0151] The server encodes and converts the generated promotional video data into a suitable format, and transmits the encoded video data to the appropriate display device using a secure communication protocol (HTTPS).
[0152] input:
[0153] Promotional video file.
[0154] output:
[0155] Encoded video data.
[0156] Step 7:
[0157] The display device decodes the received video data and plays it as a promotional video. It uses a specific decoding library to convert the received data into a playable format.
[0158] input:
[0159] Encoded video data.
[0160] output:
[0161] Decoded promotional video.
[0162] Through this series of processing steps, a promotional video can be automatically generated from the text of a novel entered by the user and quickly played back on a display device.
[0163] (Application example 1)
[0164] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0165] In the past, self-publishing authors and bloggers had to spend time and money creating promotional videos to effectively promote their work. It was also difficult for individuals and small publishers to produce high-quality advertising videos without large-scale production. This left many people without a means to widely publicize their work. The present invention solves this problem.
[0166] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0167] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for integrating specified visual elements based on the analyzed text data to generate a promotional video, means for generating a URL for the generated promotional video and transmitting it to a predetermined display device or user terminal, and means for playing and sharing the transmitted promotional video on the user terminal or on a social networking platform. This allows authors and bloggers to easily generate high-quality promotional videos for their works and promote them widely.
[0168] The "means for receiving input text data" is a function for transmitting text data input by a user to a server via a terminal, and for the server to receive the data.
[0169] "Means for natural language analysis" refers to a function that uses natural language processing technology to analyze received text data and understand the content and components of the text.
[0170] "Means for integrating visual elements to generate promotional videos" refers to a function that creates promotional videos by combining visual elements such as images, animations, and text based on information obtained through natural language analysis.
[0171] "Means for generating a URL for the generated promotional video and transmitting it to a specified display device or user terminal" refers to a function that generates an internet address (URL) for the created promotional video and transmits it to a display device or a terminal used by the user.
[0172] "Means for playing and sharing the transmitted promotional video on user devices or SNS platforms" refers to the function of making the promotional video available for viewing on user devices or SNS (social networking service) platforms and sharing it with other users.
[0173] The system for implementing this invention mainly comprises three main components: a user terminal, a server, and a display device. The role and specific processing of each component will be explained below.
[0174] User terminal
[0175] A user uses a device such as a smartphone, PC, or tablet to input text data for a novel or article. The input is done through a dedicated web form or application. For example, a user launches an application on their smartphone, enters the opening part of a novel in the text box, and clicks the submit button. This sends the input text data to the server as an HTTP POST request.
[0176] server
[0177] The server receives the input text data and analyzes it using natural language processing (NLP) algorithms. This analysis extracts key elements of the text (such as characters, setting, and theme). Specifically, an NLP engine (e.g., spaCy or Transformers) is used to understand the content of the text and identify important keywords and phrases. Next, visual elements (images, animations, text, etc.) are integrated to create a storyboard based on the analysis results. A video generation engine (e.g., FFmpeg) is used to generate a promotional video based on this storyboard.
[0178] After generating the video, the server generates a URL for the promotional video and sends this URL to the designated display device or user terminal. This process uses a secure communication protocol (e.g., HTTPS) to ensure data integrity, ensuring the secure transfer of video data.
[0179] Display device and playback
[0180] The video is played on a designated display device (digital signage, tablet, etc.) or user device using the URL of the received promotional video. Furthermore, the user device can share the generated promotional video on social media platforms (e.g., Facebook, Twitter), allowing the work to be disseminated to a wider audience.
[0181] Specific examples
[0182] For example, suppose a user uses a smartphone to input the opening part of the novel "Adventure in the Mysterious Forest."
[0183] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0184] When this text is entered, the server analyzes it and extracts keywords such as "boy," "mysterious forest," "animals," and "adventure." Based on the analysis results, the server creates a storyboard, selects appropriate visual elements, and generates a promotional video. The URL of the generated promotional video is sent to the user's device. Users can use the URL to share it on social media.
[0185] Prompt Sentence Examples
[0186] "Analyze the following text and extract its main elements. These elements can be characters, places, themes, etc."
[0187] Text: "One day, a young boy named Taro gets lost in a mysterious forest. There he meets various animals and embarks on a new adventure."
[0188] By following the above steps, this invention makes it possible to automatically generate high-quality promotional videos based on text data entered by the user and share them widely.
[0189] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0190] Step 1:
[0191] Users use devices such as smartphones, tablets, and PCs to enter the opening of a novel or the text they want to advertise. Input is done through an app or a web form and begins by clicking a submit button. The entered text data is sent to the server as an HTTP POST request.
[0192] Input: Text data (e.g., "One day, a young boy named Taro gets lost in a mysterious forest. There, he meets various animals and embarks on a new adventure.")
[0193] Output: A request with text data sent to the server
[0194] Step 2:
[0195] The server receives the text data sent by the user, stores it in a database, and passes it on to the next analysis step.
[0196] Input: Text data sent by the user
[0197] Output: Text data stored on the server
[0198] Step 3:
[0199] The server uses a natural language processing (NLP) engine (e.g., spaCy or Transformers) to analyze the received text data and extract key elements (e.g., characters, setting, theme, etc.).
[0200] Input: Text data stored on the server
[0201] Output: Extracted key elements (e.g., "boy," "mysterious forest," "animal," "adventure")
[0202] Step 4:
[0203] The server creates a storyboard based on the analysis results. The storyboard indicates which visual elements (images, animations, text, etc.) are used in which scenes. The visual elements are selected from a pre-asset library.
[0204] Input: Extracted key elements
[0205] Output: Storyboard
[0206] Step 5:
[0207] The server uses a video generation engine (e.g., FFmpeg) to generate promotional videos based on the storyboard, and the generated videos are stored on the server.
[0208] Input: Storyboard
[0209] Output: Promotional video
[0210] Step 6:
[0211] The server generates a URL for the generated promotional video and sends the URL to the user terminal. The communication uses a secure protocol (e.g., HTTPS).
[0212] Input: Promotional video
[0213] Output: Promotional video URL
[0214] Step 7:
[0215] The user's device will use the received URL to play the promotional video, and the user can also share the video on social media platforms or other media.
[0216] Input: Promotional video URL
[0217] Output: Promotional videos played and shared videos
[0218] The above steps enable the automatic generation and widespread sharing of promotional videos based on text data entered by the user.
[0219] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0220] The system for implementing this invention receives text data of a novel entered by a user, analyzes that text, generates a promotional video, and finally displays it on a digital signage or pop-up screen. It also incorporates an emotion engine that recognizes the user's emotions, and generates the promotional video based on the analysis results. This system is primarily composed of three elements: a server, a terminal, and a display device.
[0221] System configuration
[0222] 1. User Input
[0223] A user uses a device (e.g., a PC, tablet, or smartphone) to input text data, such as the opening of a novel, which is then entered into the device through a dedicated web form or application.
[0224] User: Enter the opening line of a novel and click the submit button.
[0225] 2. Sending text data
[0226] The user's device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[0227] Terminal: Sends text data to the server.
[0228] 3. Text Analysis
[0229] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.).
[0230] Server: Analyzes the text data and extracts key elements.
[0231] 4. Emotion analysis
[0232] The server uses an emotion engine to analyze the user's emotions from the input text data, and this analysis identifies the emotional tone and mood of the text.
[0233] Server: Analyzes user sentiment from text data.
[0234] 5. Image Generation
[0235] The server creates a storyboard based on the analysis results (text content and user emotions), which includes information on scenes and characters required for the promotional video.
[0236] The server selects corresponding visual elements (images, animations, text, etc.) according to the storyboard, and incorporates music and narration that match the user's emotions into the promotional video.Then, it uses a video generation engine to render the promotional video.
[0237] Server: Creates storyboards and generates footage by selecting music and narration that match the visual elements and emotions.
[0238] 6. Video data transmission
[0239] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or pop-up screen). This transmission also uses a secure communication protocol.
[0240] Server: Encodes video data and sends it to the display device.
[0241] 7. Video display
[0242] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video.
[0243] Terminal: Displays the received video data.
[0244] Specific examples
[0245] A user uses an interactive terminal installed in a bookstore to input the opening part of the novel "Adventure in the Mysterious Forest."
[0246] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0247] The terminal transmits this text data to the server.
[0248] The server analyzes the text and extracts the main elements: "boy," "mysterious forest," "animals," and "adventure."
[0249] The emotion engine analyzes the emotional tone of this text as "adventure" and "excitement."
[0250] The server creates a storyboard based on the analysis results, selecting the appropriate visual elements (forest, boy, animals) as well as adventurous music and narration.
[0251] The server combines these elements to generate a promotional video and transmits it to the bookstore's digital signage.
[0252] Digital signage plays this promotional video, visually conveying the atmosphere of the novel to customers.
[0253] This system allows bookstore visitors to get a visual and emotional feel for the novel beforehand, helping them make a purchasing decision.
[0254] The processing flow will be explained below.
[0255] Step 1:
[0256] The user accesses a dedicated web form or application using a terminal, enters text data to be used in the promotional video, such as the opening of a novel, and clicks the submit button.
[0257] Step 2:
[0258] The device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[0259] Step 3:
[0260] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.).
[0261] Step 4:
[0262] The server uses an emotion engine to analyze the user's emotions from the input text data, and this analysis identifies the emotional tone and mood of the text.
[0263] Step 5:
[0264] The server creates a storyboard based on the analysis results (text content and user emotions), which includes information on scenes and characters required for the promotional video.
[0265] Step 6:
[0266] The server selects corresponding visual elements (images, animations, text, etc.) according to the storyboard. It also integrates music and narration that match the user's emotions into the promotional video. These visual elements and emotional elements are integrated by the video generation engine and rendered as a promotional video.
[0267] Step 7:
[0268] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or pop-up screen). This transmission also uses a secure communication protocol.
[0269] Step 8:
[0270] The display device decodes the received video data and plays it as a promotional video, allowing bookstore visitors to visually and emotionally experience the atmosphere of the novel.
[0271] Specific examples
[0272] A user uses an interactive terminal installed in a bookstore to input the opening part of the novel "Adventure in the Mysterious Forest."
[0273] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0274] Step 1:
[0275] A user enters text into a web form and clicks the submit button.
[0276] Step 2:
[0277] The terminal sends this text data to the server as an HTTP POST request.
[0278] Step 3:
[0279] The server receives the text data and uses a natural language processing algorithm to extract key elements such as "boy," "mysterious forest," "animals," and "adventure."
[0280] Step 4:
[0281] The server uses an emotion engine to analyze the emotional tones contained in the text, such as "adventure" or "excitement."
[0282] Step 5:
[0283] Based on the analysis results, the server creates a storyboard that includes scenes of the boy getting lost in the forest and meeting animals.
[0284] Step 6:
[0285] The server selects the visual elements, adds adventurous music and narration, and renders the promotional video.
[0286] Step 7:
[0287] The server encodes the promotional video and transmits it to the bookstore's digital signage.
[0288] Step 8:
[0289] The digital signage plays the received video, visually conveying the atmosphere of the novel to customers.
[0290] The system allows shoppers to visually and emotionally experience the opening pages of the novel, helping them make purchasing decisions.
[0291] Example 2
[0292] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0293] Conventional promotional video generation systems simply convert text data entered by users into visual elements, making it difficult to generate videos that reflect the user's emotions and intentions. Furthermore, the accuracy of text analysis was insufficient, resulting in low-quality generated videos. Furthermore, even after the video was generated, it was often not properly transmitted to the display device in real time, resulting in display timing discrepancies.
[0294] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0295] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for analyzing emotions based on the analyzed text data, means for integrating specified visual elements based on the analyzed text data and the emotion analysis results to generate a promotional video, means for transmitting the generated promotional video to a predetermined display device, and means for playing the transmitted promotional video, thereby enabling the generation of high-quality promotional videos that reflect the user's input data and emotions.
[0296] "Input text data" refers to character information provided by a user to the system through a terminal.
[0297] "Natural language analysis" is the process by which a computer understands and analyzes human language.
[0298] "Sentiment analysis" is the process of identifying a user's emotional tone or mood from text data.
[0299] "Visual elements" refer to the visual elements such as images, animations, and text that make up the promotional video.
[0300] A "promotional video" is a short video content created for advertising or marketing purposes.
[0301] A "storyboard" is a blueprint that shows the scene composition and character placement when creating a video.
[0302] "Display device" refers to a device for playing promotional videos, such as digital signage or pop-up screens.
[0303] The system for implementing this invention receives text data of a novel entered by a user, analyzes the text, generates a promotional video, and finally displays it on a display device. This system is mainly composed of three elements: a server, a terminal, and a display device. The specific configuration and operation of this system are described below.
[0304] User Input
[0305] Users use their own devices (e.g., PCs, tablets, smartphones) to input text data, such as the opening of a novel, through a dedicated web form or application.
[0306] For example, suppose a user enters the opening line of the novel "Adventure in the Mysterious Forest" as follows:
[0307] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0308] Sending text data
[0309] The user's device sends the entered text data to the server as an HTTP POST request, using a secure communication protocol such as HTTPS.
[0310] Text analytics
[0311] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.) using tools such as the Google® Cloud Natural Language API.
[0312] Emotion analysis
[0313] The server uses an emotion engine to analyze the user's emotions from the input text data. This analysis identifies the emotional tone and mood of the text, for example, using emotion recognition tools such as IBM Watson® NLU.
[0314] Image Generation
[0315] The server creates a storyboard based on the text content and the results of user sentiment analysis. The storyboard includes information on the scenes and characters required for the promotional video. The server then selects visual elements (images, animations, text, etc.) according to the storyboard and incorporates music and narration that match the emotions into the promotional video. During this process, the promotional video is rendered using a video generation engine such as Adobe After Effects.
[0316] Video data transmission
[0317] The server then encodes the generated promotional video and transmits it to the designated display device, again using a secure communication protocol such as HTTPS.
[0318] Video display
[0319] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video, allowing customers to visually experience the atmosphere of the novel.
[0320] As a concrete example, we present a process for generating a promotional video for the opening of the novel "Adventure in the Mysterious Forest." This system allows bookstore visitors to visually and emotionally grasp the atmosphere of the novel beforehand, and can use this information to make a purchase decision.
[0321] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0322] Step 1:
[0323] User Input
[0324] Specific description:
[0325] Users use their own devices to input text data, such as the opening of a novel, through a dedicated web form or application.
[0326] input:
[0327] The text data to be input by the user (e.g., the opening part of a novel) is entered into the input field.
[0328] output:
[0329] Text data entered through the terminal is stored in a buffer.
[0330] Specific behavior:
[0331] A user visits a web form, enters text into an input field, and then clicks a "Submit" button when finished.
[0332] Step 2:
[0333] Sending text data
[0334] Specific description:
[0335] The user's device sends the entered text data to the server as an HTTP POST request, using a secure communication protocol such as HTTPS.
[0336] input:
[0337] Text data stored on the device.
[0338] output:
[0339] The text data is sent to the server as an HTTP POST request.
[0340] Specific behavior:
[0341] The device creates an HTTP POST request and sends the text data as a payload to the server.
[0342] Step 3:
[0343] Text analytics
[0344] Specific description:
[0345] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (such as characters, setting, and theme).
[0346] input:
[0347] The text data received by the server.
[0348] output:
[0349] Data from which key elements (characters, setting, theme, etc.) have been extracted.
[0350] Specific behavior:
[0351] The server stores the text data in a database, then uses NLP algorithms (e.g., Google Cloud Natural Language API) to analyze the text and extract key elements.
[0352] Step 4:
[0353] Emotion analysis
[0354] Specific description:
[0355] The server uses an emotion engine to analyze the user's emotions from the input text data, and this analysis identifies the emotional tone and mood of the text.
[0356] input:
[0357] The parsed text data.
[0358] output:
[0359] Data that identifies emotional tone and mood.
[0360] Specific behavior:
[0361] The server invokes an emotion recognition tool (e.g., IBM Watson NLU) to analyze the text data and identify the emotional tone.
[0362] Step 5:
[0363] Image Generation
[0364] Specific description:
[0365] The server creates a storyboard based on the analysis results (text content and user emotions), selects visual elements (images, animations, text, etc.) according to the storyboard, and incorporates music and narration that match the emotions into the promotional video. Finally, it renders the promotional video using a video generation engine.
[0366] input:
[0367] Data identified key elements and emotional tones.
[0368] output:
[0369] Promotional video data.
[0370] Specific behavior:
[0371] The server runs a script that generates a storyboard, selects appropriate images and animations based on the storyboard, adds music and narration, and renders the promotional video using a video generation engine (e.g., Adobe After Effects).
[0372] Step 6:
[0373] Video data transmission
[0374] Specific description:
[0375] The server then encodes the generated promotional video and transmits it to the designated display device, again using a secure communication protocol such as HTTPS.
[0376] input:
[0377] Promotional video data.
[0378] output:
[0379] The encoded video data is transmitted to a display device.
[0380] Specific behavior:
[0381] The server encodes the video file into the appropriate format and sends it to the IP address of the specified display device.
[0382] Step 7:
[0383] Video display
[0384] Specific description:
[0385] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video, allowing customers to visually experience the atmosphere of the novel.
[0386] input:
[0387] Video data sent to the display device.
[0388] output:
[0389] Promotional video displayed on display device.
[0390] Specific behavior:
[0391] The display device receives the video data sent from the server, decodes it, and plays it back using playback software.
[0392] (Application example 2)
[0393] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0394] In modern brick-and-mortar stores such as bookstores and convenience stores, there are limited means to effectively communicate the contents of a book to customers. In particular, it is difficult to visually and emotionally convey the appeal and emotional tone of the story, resulting in low book promotion effectiveness. Conventional methods only convey the contents of a book through text and still images, which means that customers do not receive sufficient information and it is difficult to stimulate their desire to purchase.
[0395] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0396] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for integrating designated visual elements and music and narration that match emotions based on the analyzed text data to generate a promotional video, means for transmitting the generated promotional video to a predetermined display device, and means for playing the transmitted promotional video, thereby making it possible to visually and emotionally convey the content and emotional tone of a book.
[0397] "Input text data" refers to novels or any other text data that a user inputs using a terminal.
[0398] The "receiving means" refers to a communication means for transmitting text data input by a user to a server and receiving the data.
[0399] "Means for natural language analysis" refers to means for analyzing received text data using natural language processing (NLP) algorithms to understand its content and components.
[0400] "Designated visual elements" refer to images or animations that are selected based on the content and emotional tone of the text data.
[0401] "Music and narration that matches the emotion" refers to music and narration selected to match the emotional tone analyzed from the text data.
[0402] The "means for generating a promotional video" refers to a means for generating a promotional video by integrating specified visual elements and music based on the analyzed text data and emotional tone.
[0403] The "means for transmitting to a display device" refers to a means for transmitting the generated promotional video data to a display device such as a digital signage or a smartphone.
[0404] The "means for playing" refers to a means for decoding the transmitted promotional video data and visually playing it back on a display device.
[0405] "Major elements" refer to important components of text data, such as characters, settings, and themes.
[0406] "Emotional tone" refers to the emotional tone or mood analyzed from text data.
[0407] A "storyboard" is a blueprint that visually organizes the scene and character information required to create a promotional video.
[0408] This invention provides a system for effectively generating and displaying promotional videos for books in brick-and-mortar stores such as bookstores, convenience stores, etc. This system is capable of processing user input, text data transmission, text analysis, emotion analysis, video generation, video data transmission, and video display.
[0409] Specifically, a user uses a smartphone application to input a novel or any other text data. This input data is securely sent to a server using the HTTPS protocol. The server then analyzes the received text data using natural language processing (NLP) algorithms. This analysis extracts the content and key elements of the text (such as characters, setting, and theme).
[0410] Next, the emotional engine analyzes the text for emotional tone, identifying the emotional tone and mood of the text. Based on this analysis, the server creates a storyboard and selects the specified visual elements (images and animations) as well as music and narration that match the emotion. This generates a promotional video.
[0411] The generated promotional video is encoded as video data and sent to a designated display device, such as a bookstore. This transmission is also done using a secure communication protocol. The display device decodes the received video data and plays it as a promotional video. This allows the appeal of the book to be conveyed visually and emotionally to store visitors.
[0412] A specific example is shown below.
[0413] Examples:
[0414] A user uses an interactive terminal installed in a bookstore to input the opening part of a novel. For example, "One day, a young boy named Taro gets lost in a mysterious forest. There, he meets various animals and embarks on a new adventure." The terminal then sends this text data to a server. The server analyzes the text and extracts the key elements—"boy," "mysterious forest," "animals," and "adventure." The emotion engine then interprets this text as "adventure" and "excitement." The server then creates a storyboard based on the analysis results and selects the appropriate visual elements (forest, boy, animals) as well as adventurous music and narration. Finally, the server integrates these elements to generate a promotional video, which is then sent to the bookstore's digital signage. The digital signage plays this promotional video, visually conveying the atmosphere of the novel to customers.
[0415] Example prompts to input to a generative AI model:
[0416] The opening of "Adventure in the Mysterious Forest":
[0417] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0418] Extract key elements from this text, analyze the emotional tone and generate a promotional video.
[0419] In this way, the present invention can enhance the effectiveness of book promotion in physical stores and stimulate customers' desire to purchase.
[0420] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0421] Step 1:
[0422] A user starts the smartphone application and enters the opening part of a novel or any other text data into the interface. The entered text data is temporarily saved within the application.
[0423] Input: Text data of the novel entered by the user.
[0424] Output: Text data that is temporarily stored on the device.
[0425] Step 2:
[0426] The terminal securely transmits the input text data to the server using the HTTPS protocol, where the data is sent to the server using an HTTP POST request.
[0427] Input: Text data stored in the device.
[0428] Output: The text data sent to the server.
[0429] Step 3:
[0430] The server applies natural language processing (NLP) algorithms to analyze the received text data, extracting key elements (characters, setting, theme, etc.) as a result of the analysis.
[0431] Input: The text data sent to the server.
[0432] Output: Key elements extracted by natural language processing.
[0433] Specific behavior: Performs text analysis using NLP libraries (e.g. spaCy, NLTK).
[0434] Step 4:
[0435] The server uses an emotion engine to analyze the emotional tone from the text data, and determines what emotion the input text evokes based on the emotional tone.
[0436] Input: The text data sent to the server.
[0437] Output: The emotional tone identified by the emotion engine (e.g., excitement, sadness, joy).
[0438] What it does: Identifies emotional tone using a sentiment analysis model (e.g., TextBlob, VADER).
[0439] Step 5:
[0440] The server creates a storyboard based on the analysis results, and selects corresponding visual elements (images, animations) as well as music and narration that match the emotion, thereby generating a promotional video.
[0441] Input: Key elements and emotional tone.
[0442] Output: The generated promo video.
[0443] Specific operation: Use a video generation engine (e.g. FFmpeg) to generate video according to the storyboard.
[0444] Step 6:
[0445] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or smartphone). This transmission also uses the HTTPS protocol.
[0446] Input: Generated promo video.
[0447] Output: The encoded video data sent to a display device.
[0448] Specific operation: Encodes video data and sends it via an HTTP POST request.
[0449] Step 7:
[0450] The display device decodes the received video data and visually reproduces it as a promotional video.
[0451] Input: Transmitted video data.
[0452] Output: The promotional video that will be played.
[0453] Specific operation: Decode and play video data using a decoding library (e.g., VLC, FFmpeg).
[0454] Through this series of steps, a promotional video generated based on the text data entered by the user is played on the display device, making it possible to provide visitors with a visually and emotionally effective promotion.
[0455] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0456] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0457] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0458] [Second embodiment]
[0459] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0460] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0461] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0462] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0463] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0464] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0465] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0466] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0467] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0468] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0469] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0470] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0471] The system for implementing this invention receives text data of a novel entered by a user, analyzes the text, generates a promotional video, and finally displays it on a digital signage or pop-up screen. This system is mainly composed of three elements: a server, a terminal, and a display device.
[0472] System configuration
[0473] 1. User Input
[0474] A user uses a device (e.g., a PC, tablet, or smartphone) to input text data, such as the opening of a novel, which is then entered into the device through a dedicated web form or application.
[0475] User: Enter the opening line of a novel and click the submit button.
[0476] 2. Sending text data
[0477] The user's device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[0478] Terminal: Sends text data to the server.
[0479] 3. Text Analysis
[0480] The server uses natural language processing (NLP) algorithms to analyze the received text data. NLP analyzes the content of the text and extracts key elements (e.g., characters, setting, theme, etc.).
[0481] Server: Analyzes the text data and extracts key elements.
[0482] 4. Image Generation
[0483] Based on the analysis results, the server creates a storyboard and selects corresponding visual elements (images, animations, text, etc.), then uses a video generation engine to render the promotional video.
[0484] Server: Creates storyboards, selects visual elements, and generates footage.
[0485] 5. Video data transmission
[0486] The server encodes the generated promotional video data and transmits it to the designated display device, also using a secure communication protocol.
[0487] Server: Encodes video data and sends it to the display device.
[0488] 6. Video display
[0489] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video.
[0490] Terminal: Displays the received video data.
[0491] Specific examples
[0492] A user uses an interactive terminal installed in a bookstore to input the opening part of the novel "Adventure in the Mysterious Forest."
[0493] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0494] The terminal transmits this text data to the server.
[0495] The server analyzes the text and extracts the keywords "boy," "mysterious forest," "animals," and "adventure."
[0496] The server creates a storyboard based on these keywords and selects appropriate visual elements to generate the promotional video.
[0497] The generated promotional video is sent to the bookstore's digital signage and displayed.
[0498] This system allows bookstore customers to visually get a feel for the atmosphere of a novel in advance and use it as a reference when making a purchase.
[0499] The processing flow will be explained below.
[0500] Step 1:
[0501] The user accesses a dedicated web form or application using a terminal, enters text data to be used in the promotional video, such as the opening of a novel, and clicks the submit button.
[0502] Step 2:
[0503] The device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[0504] Step 3:
[0505] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.).
[0506] Step 4:
[0507] The server creates a storyboard based on the analysis results, which includes information on the scenes and characters required for the promotional video.
[0508] Step 5:
[0509] The server selects corresponding visual elements (images, animations, text, etc.) according to the storyboard. These visual elements are integrated by the video generation engine and rendered as a promotional video.
[0510] Step 6:
[0511] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or pop-up screen). This transmission also uses a secure communication protocol.
[0512] Step 7:
[0513] The display device decodes the received video data and plays it as a promotional video, allowing bookstore visitors to visually experience the opening part of the novel.
[0514] Example 1
[0515] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0516] Traditionally, video generation for promoting novels and stories has been done manually, requiring a great deal of time and effort. Furthermore, specialized knowledge is required to select appropriate visual elements based on the content, making it difficult for anyone to easily create promotional videos. There has also been a lack of technology to instantly display generated videos on digital signage or pop-up screens. There is a need to solve these issues and automate the effective promotion of novels and stories quickly.
[0517] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0518] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for generating a promotional video by integrating specified visual elements based on the analyzed text data, means for transmitting the generated promotional video to a predetermined display device, means for playing the transmitted promotional video, and means for displaying the promotional video on a digital signage or pop-up screen. This makes it possible to automatically generate a promotional video based on the content of a novel or story and display it on a digital signage or pop-up screen quickly and effectively.
[0519] The "means for receiving input text data" is a function for transmitting text data input by a user using a terminal to a server and receiving the data.
[0520] "Means for natural language analysis" refers to algorithms or programs that analyze input text data, understand its content, and extract key elements and keywords.
[0521] The "means for generating promotional videos" is a function that combines the necessary visual elements based on the analyzed text data to create videos.
[0522] The "means for transmitting to a display device" is a communication function for encoding the generated promotional video data into an appropriate format and sending it to a display device such as a digital signage or a pop-up screen.
[0523] The "means for playing promotional video" is a function for the display device to decode the video data received, convert it into a playable format, and display it.
[0524] "Means for displaying on digital signage or a pop-up screen" refers to a function for actually displaying the generated promotional video on a digital signage or a pop-up screen.
[0525] "Major elements" are important items such as characters, settings, and themes that are extracted from the input text data.
[0526] "Visual elements" are visual elements such as images, animations, and text that make up the promotional video.
[0527] This invention is a system that inputs text data for a novel or story, analyzes its content, automatically generates a promotional video, and displays it on a display device. This system is mainly composed of three elements: a server, a terminal, and a display device. Each element works together to automate the promotional video generation process.
[0528] Hardware and software used
[0529] server:
[0530] Natural language processing is performed using Python's "spaCy" library and "Hugging Face"'s "Transformers" library.
[0531] Generate promotional videos using Unity and Blender.
[0532] Device:
[0533] A device such as a computer, tablet, or smartphone that accepts user input.
[0534] The ability to enter text data using a dedicated web form or application and click a submit button.
[0535] Display device:
[0536] Display devices such as digital signage and pop-up screens.
[0537] A decoding library and video player application for decoding and displaying received promotional videos.
[0538] System Operation
[0539] User Input
[0540] A user uses a terminal to input text data, such as the opening of a novel, which is received through a dedicated web form or application.
[0541] example:
[0542] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0543]
[0544] User: Enter the opening line of the novel and click the submit button.
[0545] Server processing
[0546] The server analyzes the received text data using natural language processing algorithms. Through this analysis, key elements such as characters, setting, and theme are extracted. A storyboard is then created based on the analysis results, and appropriate visual elements are selected. A promotional video is then generated based on the storyboard using Unity or Blender.
[0547] Transmission to a display device and display
[0548] The generated promotional video is encoded and transmitted from the server to the display device, which can then decode the received video data and play it back as the promotional video.
[0549] Specific examples
[0550] As an example, we will explain the creation of a promotional video for the novel "Adventure in the Mysterious Forest" using the following prompt sentence.
[0551] Example prompt sentence:
[0552] Prompt: Generate a promotional video for the novel "Adventure in the Mysterious Forest." Analyze the text below to extract characters, setting, and themes, and create a video with appropriate visual elements.
[0553] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0554] This system allows bookstore customers to visually get a feel for the atmosphere of a novel beforehand, which they can use as a reference when making a purchase. It also makes it possible to significantly reduce the effort required to create promotional materials.
[0555] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0556] Step 1:
[0557] The user uses a terminal to input the opening part of a novel. The input text data is received by a dedicated web form or application. Specifically, the user inputs the data in Unicode text format and clicks the submit button. The input data format is encoded in UTF-8.
[0558] input:
[0559] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0560] output:
[0561] UTF-8 encoded text data.
[0562] Step 2:
[0563] The device sends the entered text data to the server as an HTTP POST request. This transmission uses the HTTPS protocol to maintain data integrity. Specifically, the request body contains the text data, and the URL is the server's API endpoint.
[0564] input:
[0565] UTF-8 encoded text data.
[0566] output:
[0567] The HTTP POST request sent to the server.
[0568] Step 3:
[0569] The server analyzes the received text data using natural language processing (NLP) algorithms, such as the Python "spaCy" library and the "Transformers" library from "Hugging Face," to analyze the text content and extract key elements (e.g., characters, setting, and theme).
[0570] input:
[0571] The text data sent in the HTTP POST request.
[0572] output:
[0573] A list of the main elements (characters, setting, theme).
[0574] Step 4:
[0575] The server creates a storyboard based on the analyzed text data, determining the order and structure of scenes based on the extracted key elements, and selecting appropriate visual elements (images, animations, text, etc.).
[0576] input:
[0577] A list of the main elements (characters, setting, theme).
[0578] output:
[0579] Storyboard and selected visual elements.
[0580] Step 5:
[0581] The server generates promotional videos using Unity, Blender, etc. It combines storyboards and visual elements, renders each scene, and creates the final video file.
[0582] input:
[0583] Storyboards and visual elements.
[0584] output:
[0585] Generated promotional video file.
[0586] Step 6:
[0587] The server encodes and converts the generated promotional video data into a suitable format, and transmits the encoded video data to the appropriate display device using a secure communication protocol (HTTPS).
[0588] input:
[0589] Promotional video file.
[0590] output:
[0591] Encoded video data.
[0592] Step 7:
[0593] The display device decodes the received video data and plays it as a promotional video. It uses a specific decoding library to convert the received data into a playable format.
[0594] input:
[0595] Encoded video data.
[0596] output:
[0597] Decoded promotional video.
[0598] Through this series of processing steps, a promotional video can be automatically generated from the text of a novel entered by the user and quickly played back on a display device.
[0599] (Application example 1)
[0600] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0601] In the past, self-publishing authors and bloggers had to spend time and money creating promotional videos to effectively promote their work. It was also difficult for individuals and small publishers to produce high-quality advertising videos without large-scale production. This left many people without a means to widely publicize their work. The present invention solves this problem.
[0602] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0603] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for integrating specified visual elements based on the analyzed text data to generate a promotional video, means for generating a URL for the generated promotional video and transmitting it to a predetermined display device or user terminal, and means for playing and sharing the transmitted promotional video on the user terminal or on a social networking platform. This allows authors and bloggers to easily generate high-quality promotional videos for their works and promote them widely.
[0604] The "means for receiving input text data" is a function for transmitting text data input by a user to a server via a terminal, and for the server to receive the data.
[0605] "Means for natural language analysis" refers to a function that uses natural language processing technology to analyze received text data and understand the content and components of the text.
[0606] "Means for integrating visual elements to generate promotional videos" refers to a function that creates promotional videos by combining visual elements such as images, animations, and text based on information obtained through natural language analysis.
[0607] "Means for generating a URL for the generated promotional video and transmitting it to a specified display device or user terminal" refers to a function that generates an internet address (URL) for the created promotional video and transmits it to a display device or a terminal used by the user.
[0608] "Means for playing and sharing the transmitted promotional video on user devices or SNS platforms" refers to the function of making the promotional video available for viewing on user devices or SNS (social networking service) platforms and sharing it with other users.
[0609] The system for implementing this invention mainly comprises three main components: a user terminal, a server, and a display device. The role and specific processing of each component will be explained below.
[0610] User terminal
[0611] A user uses a device such as a smartphone, PC, or tablet to input text data for a novel or article. The input is done through a dedicated web form or application. For example, a user launches an application on their smartphone, enters the opening part of a novel in the text box, and clicks the submit button. This sends the input text data to the server as an HTTP POST request.
[0612] server
[0613] The server receives the input text data and analyzes it using natural language processing (NLP) algorithms. This analysis extracts key elements of the text (such as characters, setting, and theme). Specifically, an NLP engine (e.g., spaCy or Transformers) is used to understand the content of the text and identify important keywords and phrases. Next, visual elements (images, animations, text, etc.) are integrated to create a storyboard based on the analysis results. A video generation engine (e.g., FFmpeg) is used to generate a promotional video based on this storyboard.
[0614] After generating the video, the server generates a URL for the promotional video and sends this URL to the designated display device or user terminal. This process uses a secure communication protocol (e.g., HTTPS) to ensure data integrity, ensuring the secure transfer of video data.
[0615] Display device and playback
[0616] The video is played on a designated display device (digital signage, tablet, etc.) or user device using the URL of the received promotional video. Furthermore, the user device can share the generated promotional video on social media platforms (e.g., Facebook, Twitter), allowing the work to be disseminated to a wider audience.
[0617] Specific examples
[0618] For example, suppose a user uses a smartphone to input the opening part of the novel "Adventure in the Mysterious Forest."
[0619] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0620] When this text is entered, the server analyzes it and extracts keywords such as "boy," "mysterious forest," "animals," and "adventure." Based on the analysis results, the server creates a storyboard, selects appropriate visual elements, and generates a promotional video. The URL of the generated promotional video is sent to the user's device. Users can use the URL to share it on social media.
[0621] Prompt Sentence Examples
[0622] "Analyze the following text and extract its main elements. These elements can be characters, places, themes, etc."
[0623] Text: "One day, a young boy named Taro gets lost in a mysterious forest. There he meets various animals and embarks on a new adventure."
[0624] By following the above steps, this invention makes it possible to automatically generate high-quality promotional videos based on text data entered by the user and share them widely.
[0625] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0626] Step 1:
[0627] Users use devices such as smartphones, tablets, and PCs to enter the opening of a novel or the text they want to advertise. Input is done through an app or a web form and begins by clicking a submit button. The entered text data is sent to the server as an HTTP POST request.
[0628] Input: Text data (e.g., "One day, a young boy named Taro gets lost in a mysterious forest. There, he meets various animals and embarks on a new adventure.")
[0629] Output: A request with text data sent to the server
[0630] Step 2:
[0631] The server receives the text data sent by the user, stores it in a database, and passes it on to the next analysis step.
[0632] Input: Text data sent by the user
[0633] Output: Text data stored on the server
[0634] Step 3:
[0635] The server uses a natural language processing (NLP) engine (e.g., spaCy or Transformers) to analyze the received text data and extract key elements (e.g., characters, setting, theme, etc.).
[0636] Input: Text data stored on the server
[0637] Output: Extracted key elements (e.g., "boy," "mysterious forest," "animal," "adventure")
[0638] Step 4:
[0639] The server creates a storyboard based on the analysis results. The storyboard indicates which visual elements (images, animations, text, etc.) are used in which scenes. The visual elements are selected from a pre-asset library.
[0640] Input: Extracted key elements
[0641] Output: Storyboard
[0642] Step 5:
[0643] The server uses a video generation engine (e.g., FFmpeg) to generate promotional videos based on the storyboard, and the generated videos are stored on the server.
[0644] Input: Storyboard
[0645] Output: Promotional video
[0646] Step 6:
[0647] The server generates a URL for the generated promotional video and sends the URL to the user terminal. The communication uses a secure protocol (e.g., HTTPS).
[0648] Input: Promotional video
[0649] Output: Promotional video URL
[0650] Step 7:
[0651] The user's device will use the received URL to play the promotional video, and the user can also share the video on social media platforms or other media.
[0652] Input: Promotional video URL
[0653] Output: Promotional videos played and shared videos
[0654] The above steps enable the automatic generation and widespread sharing of promotional videos based on text data entered by the user.
[0655] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0656] The system for implementing this invention receives text data of a novel entered by a user, analyzes that text, generates a promotional video, and finally displays it on a digital signage or pop-up screen. It also incorporates an emotion engine that recognizes the user's emotions, and generates the promotional video based on the analysis results. This system is primarily composed of three elements: a server, a terminal, and a display device.
[0657] System configuration
[0658] 1. User Input
[0659] A user uses a device (e.g., a PC, tablet, or smartphone) to input text data, such as the opening of a novel, which is then entered into the device through a dedicated web form or application.
[0660] User: Enter the opening line of a novel and click the submit button.
[0661] 2. Sending text data
[0662] The user's device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[0663] Terminal: Sends text data to the server.
[0664] 3. Text Analysis
[0665] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.).
[0666] Server: Analyzes the text data and extracts key elements.
[0667] 4. Emotion analysis
[0668] The server uses an emotion engine to analyze the user's emotions from the input text data, and this analysis identifies the emotional tone and mood of the text.
[0669] Server: Analyzes user sentiment from text data.
[0670] 5. Image Generation
[0671] The server creates a storyboard based on the analysis results (text content and user emotions), which includes information on scenes and characters required for the promotional video.
[0672] The server selects corresponding visual elements (images, animations, text, etc.) according to the storyboard, and incorporates music and narration that match the user's emotions into the promotional video.Then, it uses a video generation engine to render the promotional video.
[0673] Server: Creates storyboards and generates footage by selecting music and narration that match the visual elements and emotions.
[0674] 6. Video data transmission
[0675] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or pop-up screen). This transmission also uses a secure communication protocol.
[0676] Server: Encodes video data and sends it to the display device.
[0677] 7. Video display
[0678] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video.
[0679] Terminal: Displays the received video data.
[0680] Specific examples
[0681] A user uses an interactive terminal installed in a bookstore to input the opening part of the novel "Adventure in the Mysterious Forest."
[0682] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0683] The terminal transmits this text data to the server.
[0684] The server analyzes the text and extracts the main elements: "boy," "mysterious forest," "animals," and "adventure."
[0685] The emotion engine analyzes the emotional tone of this text as "adventure" and "excitement."
[0686] The server creates a storyboard based on the analysis results, selecting the appropriate visual elements (forest, boy, animals) as well as adventurous music and narration.
[0687] The server combines these elements to generate a promotional video and transmits it to the bookstore's digital signage.
[0688] Digital signage plays this promotional video, visually conveying the atmosphere of the novel to customers.
[0689] This system allows bookstore visitors to get a visual and emotional feel for the novel beforehand, helping them make a purchasing decision.
[0690] The processing flow will be explained below.
[0691] Step 1:
[0692] The user accesses a dedicated web form or application using a terminal, enters text data to be used in the promotional video, such as the opening of a novel, and clicks the submit button.
[0693] Step 2:
[0694] The device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[0695] Step 3:
[0696] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.).
[0697] Step 4:
[0698] The server uses an emotion engine to analyze the user's emotions from the input text data, and this analysis identifies the emotional tone and mood of the text.
[0699] Step 5:
[0700] The server creates a storyboard based on the analysis results (text content and user emotions), which includes information on scenes and characters required for the promotional video.
[0701] Step 6:
[0702] The server selects corresponding visual elements (images, animations, text, etc.) according to the storyboard. It also integrates music and narration that match the user's emotions into the promotional video. These visual elements and emotional elements are integrated by the video generation engine and rendered as a promotional video.
[0703] Step 7:
[0704] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or pop-up screen). This transmission also uses a secure communication protocol.
[0705] Step 8:
[0706] The display device decodes the received video data and plays it as a promotional video, allowing bookstore visitors to visually and emotionally experience the atmosphere of the novel.
[0707] Specific examples
[0708] A user uses an interactive terminal installed in a bookstore to input the opening part of the novel "Adventure in the Mysterious Forest."
[0709] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0710] Step 1:
[0711] A user enters text into a web form and clicks the submit button.
[0712] Step 2:
[0713] The terminal sends this text data to the server as an HTTP POST request.
[0714] Step 3:
[0715] The server receives the text data and uses a natural language processing algorithm to extract key elements such as "boy," "mysterious forest," "animals," and "adventure."
[0716] Step 4:
[0717] The server uses an emotion engine to analyze the emotional tones contained in the text, such as "adventure" or "excitement."
[0718] Step 5:
[0719] Based on the analysis results, the server creates a storyboard that includes scenes of the boy getting lost in the forest and meeting animals.
[0720] Step 6:
[0721] The server selects the visual elements, adds adventurous music and narration, and renders the promotional video.
[0722] Step 7:
[0723] The server encodes the promotional video and transmits it to the bookstore's digital signage.
[0724] Step 8:
[0725] The digital signage plays the received video, visually conveying the atmosphere of the novel to customers.
[0726] The system allows shoppers to visually and emotionally experience the opening pages of the novel, helping them make purchasing decisions.
[0727] Example 2
[0728] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0729] Conventional promotional video generation systems simply convert text data entered by users into visual elements, making it difficult to generate videos that reflect the user's emotions and intentions. Furthermore, the accuracy of text analysis was insufficient, resulting in low-quality generated videos. Furthermore, even after the video was generated, it was often not properly transmitted to the display device in real time, resulting in display timing discrepancies.
[0730] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0731] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for analyzing emotions based on the analyzed text data, means for integrating specified visual elements based on the analyzed text data and the emotion analysis results to generate a promotional video, means for transmitting the generated promotional video to a predetermined display device, and means for playing the transmitted promotional video, thereby enabling the generation of high-quality promotional videos that reflect the user's input data and emotions.
[0732] "Input text data" refers to character information provided by a user to the system through a terminal.
[0733] "Natural language analysis" is the process by which a computer understands and analyzes human language.
[0734] "Sentiment analysis" is the process of identifying a user's emotional tone or mood from text data.
[0735] "Visual elements" refer to the visual elements such as images, animations, and text that make up the promotional video.
[0736] A "promotional video" is a short video content created for advertising or marketing purposes.
[0737] A "storyboard" is a blueprint that shows the scene composition and character placement when creating a video.
[0738] "Display device" refers to a device for playing promotional videos, such as digital signage or pop-up screens.
[0739] The system for implementing this invention receives text data of a novel entered by a user, analyzes the text, generates a promotional video, and finally displays it on a display device. This system is mainly composed of three elements: a server, a terminal, and a display device. The specific configuration and operation of this system are described below.
[0740] User Input
[0741] Users use their own devices (e.g., PCs, tablets, smartphones) to input text data, such as the opening of a novel, through a dedicated web form or application.
[0742] For example, suppose a user enters the opening line of the novel "Adventure in the Mysterious Forest" as follows:
[0743] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0744] Sending text data
[0745] The user's device sends the entered text data to the server as an HTTP POST request, using a secure communication protocol such as HTTPS.
[0746] Text analytics
[0747] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.) using tools such as the Google Cloud Natural Language API.
[0748] Emotion analysis
[0749] The server uses an emotion engine to analyze the user's emotions from the input text data. This analysis identifies the emotional tone and mood of the text, for example, using emotion recognition tools such as IBM Watson NLU.
[0750] Image Generation
[0751] The server creates a storyboard based on the text content and the results of user sentiment analysis. The storyboard includes information on the scenes and characters required for the promotional video. The server then selects visual elements (images, animations, text, etc.) according to the storyboard and incorporates music and narration that match the emotions into the promotional video. During this process, the promotional video is rendered using a video generation engine such as Adobe After Effects.
[0752] Video data transmission
[0753] The server then encodes the generated promotional video and transmits it to the designated display device, again using a secure communication protocol such as HTTPS.
[0754] Video display
[0755] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video, allowing customers to visually experience the atmosphere of the novel.
[0756] As a concrete example, we present a process for generating a promotional video for the opening of the novel "Adventure in the Mysterious Forest." This system allows bookstore visitors to visually and emotionally grasp the atmosphere of the novel beforehand, and can use this information to make a purchase decision.
[0757] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0758] Step 1:
[0759] User Input
[0760] Specific description:
[0761] Users use their own devices to input text data, such as the opening of a novel, through a dedicated web form or application.
[0762] input:
[0763] The text data to be input by the user (e.g., the opening part of a novel) is entered into the input field.
[0764] output:
[0765] Text data entered through the terminal is stored in a buffer.
[0766] Specific behavior:
[0767] A user visits a web form, enters text into an input field, and then clicks a "Submit" button when finished.
[0768] Step 2:
[0769] Sending text data
[0770] Specific description:
[0771] The user's device sends the entered text data to the server as an HTTP POST request, using a secure communication protocol such as HTTPS.
[0772] input:
[0773] Text data stored on the device.
[0774] output:
[0775] The text data is sent to the server as an HTTP POST request.
[0776] Specific behavior:
[0777] The device creates an HTTP POST request and sends the text data as a payload to the server.
[0778] Step 3:
[0779] Text analytics
[0780] Specific description:
[0781] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (such as characters, setting, and theme).
[0782] input:
[0783] The text data received by the server.
[0784] output:
[0785] Data from which key elements (characters, setting, theme, etc.) have been extracted.
[0786] Specific behavior:
[0787] The server stores the text data in a database, then uses NLP algorithms (e.g., Google Cloud Natural Language API) to analyze the text and extract key elements.
[0788] Step 4:
[0789] Emotion analysis
[0790] Specific description:
[0791] The server uses an emotion engine to analyze the user's emotions from the input text data, and this analysis identifies the emotional tone and mood of the text.
[0792] input:
[0793] The parsed text data.
[0794] output:
[0795] Data that identifies emotional tone and mood.
[0796] Specific behavior:
[0797] The server invokes an emotion recognition tool (e.g., IBM Watson NLU) to analyze the text data and identify the emotional tone.
[0798] Step 5:
[0799] Image Generation
[0800] Specific description:
[0801] The server creates a storyboard based on the analysis results (text content and user emotions), selects visual elements (images, animations, text, etc.) according to the storyboard, and incorporates music and narration that match the emotions into the promotional video. Finally, it renders the promotional video using a video generation engine.
[0802] input:
[0803] Data identified key elements and emotional tones.
[0804] output:
[0805] Promotional video data.
[0806] Specific behavior:
[0807] The server runs a script that generates a storyboard, selects appropriate images and animations based on the storyboard, adds music and narration, and renders the promotional video using a video generation engine (e.g., Adobe After Effects).
[0808] Step 6:
[0809] Video data transmission
[0810] Specific description:
[0811] The server then encodes the generated promotional video and transmits it to the designated display device, again using a secure communication protocol such as HTTPS.
[0812] input:
[0813] Promotional video data.
[0814] output:
[0815] The encoded video data is transmitted to a display device.
[0816] Specific behavior:
[0817] The server encodes the video file into the appropriate format and sends it to the IP address of the specified display device.
[0818] Step 7:
[0819] Video display
[0820] Specific description:
[0821] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video, allowing customers to visually experience the atmosphere of the novel.
[0822] input:
[0823] Video data sent to the display device.
[0824] output:
[0825] Promotional video displayed on display device.
[0826] Specific behavior:
[0827] The display device receives the video data sent from the server, decodes it, and plays it back using playback software.
[0828] (Application example 2)
[0829] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0830] In modern brick-and-mortar stores such as bookstores and convenience stores, there are limited means to effectively communicate the contents of a book to customers. In particular, it is difficult to visually and emotionally convey the appeal and emotional tone of the story, resulting in low book promotion effectiveness. Conventional methods only convey the contents of a book through text and still images, which means that customers do not receive sufficient information and it is difficult to stimulate their desire to purchase.
[0831] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0832] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for integrating designated visual elements and music and narration that match emotions based on the analyzed text data to generate a promotional video, means for transmitting the generated promotional video to a predetermined display device, and means for playing the transmitted promotional video, thereby making it possible to visually and emotionally convey the content and emotional tone of a book.
[0833] "Input text data" refers to novels or any other text data that a user inputs using a terminal.
[0834] The "receiving means" refers to a communication means for transmitting text data input by a user to a server and receiving the data.
[0835] "Means for natural language analysis" refers to means for analyzing received text data using natural language processing (NLP) algorithms to understand its content and components.
[0836] "Designated visual elements" refer to images or animations that are selected based on the content and emotional tone of the text data.
[0837] "Music and narration that matches the emotion" refers to music and narration selected to match the emotional tone analyzed from the text data.
[0838] The "means for generating a promotional video" refers to a means for generating a promotional video by integrating specified visual elements and music based on the analyzed text data and emotional tone.
[0839] The "means for transmitting to a display device" refers to a means for transmitting the generated promotional video data to a display device such as a digital signage or a smartphone.
[0840] The "means for playing" refers to a means for decoding the transmitted promotional video data and visually playing it back on a display device.
[0841] "Major elements" refer to important components of text data, such as characters, settings, and themes.
[0842] "Emotional tone" refers to the emotional tone or mood analyzed from text data.
[0843] A "storyboard" is a blueprint that visually organizes the scene and character information required to create a promotional video.
[0844] This invention provides a system for effectively generating and displaying promotional videos for books in brick-and-mortar stores such as bookstores, convenience stores, etc. This system is capable of processing user input, text data transmission, text analysis, emotion analysis, video generation, video data transmission, and video display.
[0845] Specifically, a user uses a smartphone application to input a novel or any other text data. This input data is securely sent to a server using the HTTPS protocol. The server then analyzes the received text data using natural language processing (NLP) algorithms. This analysis extracts the content and key elements of the text (such as characters, setting, and theme).
[0846] Next, the emotional engine analyzes the text for emotional tone, identifying the emotional tone and mood of the text. Based on this analysis, the server creates a storyboard and selects the specified visual elements (images and animations) as well as music and narration that match the emotion. This generates a promotional video.
[0847] The generated promotional video is encoded as video data and sent to a designated display device, such as a bookstore. This transmission is also done using a secure communication protocol. The display device decodes the received video data and plays it as a promotional video. This allows the appeal of the book to be conveyed visually and emotionally to store visitors.
[0848] A specific example is shown below.
[0849] Examples:
[0850] A user uses an interactive terminal installed in a bookstore to input the opening part of a novel. For example, "One day, a young boy named Taro gets lost in a mysterious forest. There, he meets various animals and embarks on a new adventure." The terminal then sends this text data to a server. The server analyzes the text and extracts the key elements—"boy," "mysterious forest," "animals," and "adventure." The emotion engine then interprets this text as "adventure" and "excitement." The server then creates a storyboard based on the analysis results and selects the appropriate visual elements (forest, boy, animals) as well as adventurous music and narration. Finally, the server integrates these elements to generate a promotional video, which is then sent to the bookstore's digital signage. The digital signage plays this promotional video, visually conveying the atmosphere of the novel to customers.
[0851] Example prompts to input to a generative AI model:
[0852] The opening of "Adventure in the Mysterious Forest":
[0853] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0854] Extract key elements from this text, analyze the emotional tone and generate a promotional video.
[0855] In this way, the present invention can enhance the effectiveness of book promotion in physical stores and stimulate customers' desire to purchase.
[0856] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0857] Step 1:
[0858] A user starts the smartphone application and enters the opening part of a novel or any other text data into the interface. The entered text data is temporarily saved within the application.
[0859] Input: Text data of the novel entered by the user.
[0860] Output: Text data that is temporarily stored on the device.
[0861] Step 2:
[0862] The terminal securely transmits the input text data to the server using the HTTPS protocol, where the data is sent to the server using an HTTP POST request.
[0863] Input: Text data stored in the device.
[0864] Output: The text data sent to the server.
[0865] Step 3:
[0866] The server applies natural language processing (NLP) algorithms to analyze the received text data, extracting key elements (characters, setting, theme, etc.) as a result of the analysis.
[0867] Input: The text data sent to the server.
[0868] Output: Key elements extracted by natural language processing.
[0869] Specific behavior: Performs text analysis using NLP libraries (e.g. spaCy, NLTK).
[0870] Step 4:
[0871] The server uses an emotion engine to analyze the emotional tone from the text data, and determines what emotion the input text evokes based on the emotional tone.
[0872] Input: The text data sent to the server.
[0873] Output: The emotional tone identified by the emotion engine (e.g., excitement, sadness, joy).
[0874] What it does: Identifies emotional tone using a sentiment analysis model (e.g., TextBlob, VADER).
[0875] Step 5:
[0876] The server creates a storyboard based on the analysis results, and selects corresponding visual elements (images, animations) as well as music and narration that match the emotion, thereby generating a promotional video.
[0877] Input: Key elements and emotional tone.
[0878] Output: The generated promo video.
[0879] Specific operation: Use a video generation engine (e.g. FFmpeg) to generate video according to the storyboard.
[0880] Step 6:
[0881] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or smartphone). This transmission also uses the HTTPS protocol.
[0882] Input: Generated promo video.
[0883] Output: The encoded video data sent to a display device.
[0884] Specific operation: Encodes video data and sends it via an HTTP POST request.
[0885] Step 7:
[0886] The display device decodes the received video data and visually reproduces it as a promotional video.
[0887] Input: Transmitted video data.
[0888] Output: The promotional video that will be played.
[0889] Specific operation: Decode and play video data using a decoding library (e.g., VLC, FFmpeg).
[0890] Through this series of steps, a promotional video generated based on the text data entered by the user is played on the display device, making it possible to provide visitors with a visually and emotionally effective promotion.
[0891] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0892] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0893] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0894] [Third embodiment]
[0895] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0896] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0897] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0898] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0899] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0900] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0901] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0902] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0903] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0904] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0905] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0906] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0907] The system for implementing this invention receives text data of a novel entered by a user, analyzes the text, generates a promotional video, and finally displays it on a digital signage or pop-up screen. This system is mainly composed of three elements: a server, a terminal, and a display device.
[0908] System configuration
[0909] 1. User Input
[0910] A user uses a device (e.g., a PC, tablet, or smartphone) to input text data, such as the opening of a novel, which is then entered into the device through a dedicated web form or application.
[0911] User: Enter the opening line of a novel and click the submit button.
[0912] 2. Sending text data
[0913] The user's device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[0914] Terminal: Sends text data to the server.
[0915] 3. Text Analysis
[0916] The server uses natural language processing (NLP) algorithms to analyze the received text data. NLP analyzes the content of the text and extracts key elements (e.g., characters, setting, theme, etc.).
[0917] Server: Analyzes the text data and extracts key elements.
[0918] 4. Image Generation
[0919] Based on the analysis results, the server creates a storyboard and selects corresponding visual elements (images, animations, text, etc.), then uses a video generation engine to render the promotional video.
[0920] Server: Creates storyboards, selects visual elements, and generates footage.
[0921] 5. Video data transmission
[0922] The server encodes the generated promotional video data and transmits it to the designated display device, also using a secure communication protocol.
[0923] Server: Encodes video data and sends it to the display device.
[0924] 6. Video display
[0925] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video.
[0926] Terminal: Displays the received video data.
[0927] Specific examples
[0928] A user uses an interactive terminal installed in a bookstore to input the opening part of the novel "Adventure in the Mysterious Forest."
[0929] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0930] The terminal transmits this text data to the server.
[0931] The server analyzes the text and extracts the keywords "boy," "mysterious forest," "animals," and "adventure."
[0932] The server creates a storyboard based on these keywords and selects appropriate visual elements to generate the promotional video.
[0933] The generated promotional video is sent to the bookstore's digital signage and displayed.
[0934] This system allows bookstore customers to visually get a feel for the atmosphere of a novel in advance and use it as a reference when making a purchase.
[0935] The processing flow will be explained below.
[0936] Step 1:
[0937] The user accesses a dedicated web form or application using a terminal, enters text data to be used in the promotional video, such as the opening of a novel, and clicks the submit button.
[0938] Step 2:
[0939] The device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[0940] Step 3:
[0941] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.).
[0942] Step 4:
[0943] The server creates a storyboard based on the analysis results, which includes information on the scenes and characters required for the promotional video.
[0944] Step 5:
[0945] The server selects corresponding visual elements (images, animations, text, etc.) according to the storyboard. These visual elements are integrated by the video generation engine and rendered as a promotional video.
[0946] Step 6:
[0947] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or pop-up screen). This transmission also uses a secure communication protocol.
[0948] Step 7:
[0949] The display device decodes the received video data and plays it as a promotional video, allowing bookstore visitors to visually experience the opening part of the novel.
[0950] Example 1
[0951] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0952] Traditionally, video generation for promoting novels and stories has been done manually, requiring a great deal of time and effort. Furthermore, specialized knowledge is required to select appropriate visual elements based on the content, making it difficult for anyone to easily create promotional videos. There has also been a lack of technology to instantly display generated videos on digital signage or pop-up screens. There is a need to solve these issues and automate the effective promotion of novels and stories quickly.
[0953] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0954] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for generating a promotional video by integrating specified visual elements based on the analyzed text data, means for transmitting the generated promotional video to a predetermined display device, means for playing the transmitted promotional video, and means for displaying the promotional video on a digital signage or pop-up screen. This makes it possible to automatically generate a promotional video based on the content of a novel or story and display it on a digital signage or pop-up screen quickly and effectively.
[0955] The "means for receiving input text data" is a function for transmitting text data input by a user using a terminal to a server and receiving the data.
[0956] "Means for natural language analysis" refers to algorithms or programs that analyze input text data, understand its content, and extract key elements and keywords.
[0957] The "means for generating promotional videos" is a function that combines the necessary visual elements based on the analyzed text data to create videos.
[0958] The "means for transmitting to a display device" is a communication function for encoding the generated promotional video data into an appropriate format and sending it to a display device such as a digital signage or a pop-up screen.
[0959] The "means for playing promotional video" is a function for the display device to decode the video data received, convert it into a playable format, and display it.
[0960] "Means for displaying on digital signage or a pop-up screen" refers to a function for actually displaying the generated promotional video on a digital signage or a pop-up screen.
[0961] "Major elements" are important items such as characters, settings, and themes that are extracted from the input text data.
[0962] "Visual elements" are visual elements such as images, animations, and text that make up the promotional video.
[0963] This invention is a system that inputs text data for a novel or story, analyzes its content, automatically generates a promotional video, and displays it on a display device. This system is mainly composed of three elements: a server, a terminal, and a display device. Each element works together to automate the promotional video generation process.
[0964] Hardware and software used
[0965] server:
[0966] Natural language processing is performed using Python's "spaCy" library and "Hugging Face"'s "Transformers" library.
[0967] Generate promotional videos using Unity and Blender.
[0968] Device:
[0969] A device such as a computer, tablet, or smartphone that accepts user input.
[0970] The ability to enter text data using a dedicated web form or application and click a submit button.
[0971] Display device:
[0972] Display devices such as digital signage and pop-up screens.
[0973] A decoding library and video player application for decoding and displaying received promotional videos.
[0974] System Operation
[0975] User Input
[0976] A user uses a terminal to input text data, such as the opening of a novel, which is received through a dedicated web form or application.
[0977] example:
[0978] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0979]
[0980] User: Enter the opening line of the novel and click the submit button.
[0981] Server processing
[0982] The server analyzes the received text data using natural language processing algorithms. Through this analysis, key elements such as characters, setting, and theme are extracted. A storyboard is then created based on the analysis results, and appropriate visual elements are selected. A promotional video is then generated based on the storyboard using Unity or Blender.
[0983] Transmission to a display device and display
[0984] The generated promotional video is encoded and transmitted from the server to the display device, which can then decode the received video data and play it back as the promotional video.
[0985] Specific examples
[0986] As an example, we will explain the creation of a promotional video for the novel "Adventure in the Mysterious Forest" using the following prompt sentence.
[0987] Example prompt sentence:
[0988] Prompt: Generate a promotional video for the novel "Adventure in the Mysterious Forest." Analyze the text below to extract characters, setting, and themes, and create a video with appropriate visual elements.
[0989] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0990] This system allows bookstore customers to visually get a feel for the atmosphere of a novel beforehand, which they can use as a reference when making a purchase. It also makes it possible to significantly reduce the effort required to create promotional materials.
[0991] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0992] Step 1:
[0993] The user uses a terminal to input the opening part of a novel. The input text data is received by a dedicated web form or application. Specifically, the user inputs the data in Unicode text format and clicks the submit button. The input data format is encoded in UTF-8.
[0994] input:
[0995] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[0996] output:
[0997] UTF-8 encoded text data.
[0998] Step 2:
[0999] The device sends the entered text data to the server as an HTTP POST request. This transmission uses the HTTPS protocol to maintain data integrity. Specifically, the request body contains the text data, and the URL is the server's API endpoint.
[1000] input:
[1001] UTF-8 encoded text data.
[1002] output:
[1003] The HTTP POST request sent to the server.
[1004] Step 3:
[1005] The server analyzes the received text data using natural language processing (NLP) algorithms, such as the Python "spaCy" library and the "Transformers" library from "Hugging Face," to analyze the text content and extract key elements (e.g., characters, setting, and theme).
[1006] input:
[1007] The text data sent in the HTTP POST request.
[1008] output:
[1009] A list of the main elements (characters, setting, theme).
[1010] Step 4:
[1011] The server creates a storyboard based on the analyzed text data, determining the order and structure of scenes based on the extracted key elements, and selecting appropriate visual elements (images, animations, text, etc.).
[1012] input:
[1013] A list of the main elements (characters, setting, theme).
[1014] output:
[1015] Storyboard and selected visual elements.
[1016] Step 5:
[1017] The server generates promotional videos using Unity, Blender, etc. It combines storyboards and visual elements, renders each scene, and creates the final video file.
[1018] input:
[1019] Storyboards and visual elements.
[1020] output:
[1021] Generated promotional video file.
[1022] Step 6:
[1023] The server encodes and converts the generated promotional video data into a suitable format, and transmits the encoded video data to the appropriate display device using a secure communication protocol (HTTPS).
[1024] input:
[1025] Promotional video file.
[1026] output:
[1027] Encoded video data.
[1028] Step 7:
[1029] The display device decodes the received video data and plays it as a promotional video. It uses a specific decoding library to convert the received data into a playable format.
[1030] input:
[1031] Encoded video data.
[1032] output:
[1033] Decoded promotional video.
[1034] Through this series of processing steps, a promotional video can be automatically generated from the text of a novel entered by the user and quickly played back on a display device.
[1035] (Application example 1)
[1036] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1037] In the past, self-publishing authors and bloggers had to spend time and money creating promotional videos to effectively promote their work. It was also difficult for individuals and small publishers to produce high-quality advertising videos without large-scale production. This left many people without a means to widely publicize their work. The present invention solves this problem.
[1038] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1039] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for integrating specified visual elements based on the analyzed text data to generate a promotional video, means for generating a URL for the generated promotional video and transmitting it to a predetermined display device or user terminal, and means for playing and sharing the transmitted promotional video on the user terminal or on a social networking platform. This allows authors and bloggers to easily generate high-quality promotional videos for their works and promote them widely.
[1040] The "means for receiving input text data" is a function for transmitting text data input by a user to a server via a terminal, and for the server to receive the data.
[1041] "Means for natural language analysis" refers to a function that uses natural language processing technology to analyze received text data and understand the content and components of the text.
[1042] "Means for integrating visual elements to generate promotional videos" refers to a function that creates promotional videos by combining visual elements such as images, animations, and text based on information obtained through natural language analysis.
[1043] "Means for generating a URL for the generated promotional video and transmitting it to a specified display device or user terminal" refers to a function that generates an internet address (URL) for the created promotional video and transmits it to a display device or a terminal used by the user.
[1044] "Means for playing and sharing the transmitted promotional video on user devices or SNS platforms" refers to the function of making the promotional video available for viewing on user devices or SNS (social networking service) platforms and sharing it with other users.
[1045] The system for implementing this invention mainly comprises three main components: a user terminal, a server, and a display device. The role and specific processing of each component will be explained below.
[1046] User terminal
[1047] A user uses a device such as a smartphone, PC, or tablet to input text data for a novel or article. The input is done through a dedicated web form or application. For example, a user launches an application on their smartphone, enters the opening part of a novel in the text box, and clicks the submit button. This sends the input text data to the server as an HTTP POST request.
[1048] server
[1049] The server receives the input text data and analyzes it using natural language processing (NLP) algorithms. This analysis extracts key elements of the text (such as characters, setting, and theme). Specifically, an NLP engine (e.g., spaCy or Transformers) is used to understand the content of the text and identify important keywords and phrases. Next, visual elements (images, animations, text, etc.) are integrated to create a storyboard based on the analysis results. A video generation engine (e.g., FFmpeg) is used to generate a promotional video based on this storyboard.
[1050] After generating the video, the server generates a URL for the promotional video and sends this URL to the designated display device or user terminal. This process uses a secure communication protocol (e.g., HTTPS) to ensure data integrity, ensuring the secure transfer of video data.
[1051] Display device and playback
[1052] The video is played on a designated display device (digital signage, tablet, etc.) or user device using the URL of the received promotional video. Furthermore, the user device can share the generated promotional video on social media platforms (e.g., Facebook, Twitter), allowing the work to be disseminated to a wider audience.
[1053] Specific examples
[1054] For example, suppose a user uses a smartphone to input the opening part of the novel "Adventure in the Mysterious Forest."
[1055] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1056] When this text is entered, the server analyzes it and extracts keywords such as "boy," "mysterious forest," "animals," and "adventure." Based on the analysis results, the server creates a storyboard, selects appropriate visual elements, and generates a promotional video. The URL of the generated promotional video is sent to the user's device. Users can use the URL to share it on social media.
[1057] Prompt Sentence Examples
[1058] "Analyze the following text and extract its main elements. These elements can be characters, places, themes, etc."
[1059] Text: "One day, a young boy named Taro gets lost in a mysterious forest. There he meets various animals and embarks on a new adventure."
[1060] By following the above steps, this invention makes it possible to automatically generate high-quality promotional videos based on text data entered by the user and share them widely.
[1061] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1062] Step 1:
[1063] Users use devices such as smartphones, tablets, and PCs to enter the opening of a novel or the text they want to advertise. Input is done through an app or a web form and begins by clicking a submit button. The entered text data is sent to the server as an HTTP POST request.
[1064] Input: Text data (e.g., "One day, a young boy named Taro gets lost in a mysterious forest. There, he meets various animals and embarks on a new adventure.")
[1065] Output: A request with text data sent to the server
[1066] Step 2:
[1067] The server receives the text data sent by the user, stores it in a database, and passes it on to the next analysis step.
[1068] Input: Text data sent by the user
[1069] Output: Text data stored on the server
[1070] Step 3:
[1071] The server uses a natural language processing (NLP) engine (e.g., spaCy or Transformers) to analyze the received text data and extract key elements (e.g., characters, setting, theme, etc.).
[1072] Input: Text data stored on the server
[1073] Output: Extracted key elements (e.g., "boy," "mysterious forest," "animal," "adventure")
[1074] Step 4:
[1075] The server creates a storyboard based on the analysis results. The storyboard indicates which visual elements (images, animations, text, etc.) are used in which scenes. The visual elements are selected from a pre-asset library.
[1076] Input: Extracted key elements
[1077] Output: Storyboard
[1078] Step 5:
[1079] The server uses a video generation engine (e.g., FFmpeg) to generate promotional videos based on the storyboard, and the generated videos are stored on the server.
[1080] Input: Storyboard
[1081] Output: Promotional video
[1082] Step 6:
[1083] The server generates a URL for the generated promotional video and sends the URL to the user terminal. The communication uses a secure protocol (e.g., HTTPS).
[1084] Input: Promotional video
[1085] Output: Promotional video URL
[1086] Step 7:
[1087] The user's device will use the received URL to play the promotional video, and the user can also share the video on social media platforms or other media.
[1088] Input: Promotional video URL
[1089] Output: Promotional videos played and shared videos
[1090] The above steps enable the automatic generation and widespread sharing of promotional videos based on text data entered by the user.
[1091] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1092] The system for implementing this invention receives text data of a novel entered by a user, analyzes that text, generates a promotional video, and finally displays it on a digital signage or pop-up screen. It also incorporates an emotion engine that recognizes the user's emotions, and generates the promotional video based on the analysis results. This system is primarily composed of three elements: a server, a terminal, and a display device.
[1093] System configuration
[1094] 1. User Input
[1095] A user uses a device (e.g., a PC, tablet, or smartphone) to input text data, such as the opening of a novel, which is then entered into the device through a dedicated web form or application.
[1096] User: Enter the opening line of a novel and click the submit button.
[1097] 2. Sending text data
[1098] The user's device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[1099] Terminal: Sends text data to the server.
[1100] 3. Text Analysis
[1101] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.).
[1102] Server: Analyzes the text data and extracts key elements.
[1103] 4. Emotion analysis
[1104] The server uses an emotion engine to analyze the user's emotions from the input text data, and this analysis identifies the emotional tone and mood of the text.
[1105] Server: Analyzes user sentiment from text data.
[1106] 5. Image Generation
[1107] The server creates a storyboard based on the analysis results (text content and user emotions), which includes information on scenes and characters required for the promotional video.
[1108] The server selects corresponding visual elements (images, animations, text, etc.) according to the storyboard, and incorporates music and narration that match the user's emotions into the promotional video.Then, it uses a video generation engine to render the promotional video.
[1109] Server: Creates storyboards and generates footage by selecting music and narration that match the visual elements and emotions.
[1110] 6. Video data transmission
[1111] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or pop-up screen). This transmission also uses a secure communication protocol.
[1112] Server: Encodes video data and sends it to the display device.
[1113] 7. Video display
[1114] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video.
[1115] Terminal: Displays the received video data.
[1116] Specific examples
[1117] A user uses an interactive terminal installed in a bookstore to input the opening part of the novel "Adventure in the Mysterious Forest."
[1118] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1119] The terminal transmits this text data to the server.
[1120] The server analyzes the text and extracts the main elements: "boy," "mysterious forest," "animals," and "adventure."
[1121] The emotion engine analyzes the emotional tone of this text as "adventure" and "excitement."
[1122] The server creates a storyboard based on the analysis results, selecting the appropriate visual elements (forest, boy, animals) as well as adventurous music and narration.
[1123] The server combines these elements to generate a promotional video and transmits it to the bookstore's digital signage.
[1124] Digital signage plays this promotional video, visually conveying the atmosphere of the novel to customers.
[1125] This system allows bookstore visitors to get a visual and emotional feel for the novel beforehand, helping them make a purchasing decision.
[1126] The processing flow will be explained below.
[1127] Step 1:
[1128] The user accesses a dedicated web form or application using a terminal, enters text data to be used in the promotional video, such as the opening of a novel, and clicks the submit button.
[1129] Step 2:
[1130] The device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[1131] Step 3:
[1132] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.).
[1133] Step 4:
[1134] The server uses an emotion engine to analyze the user's emotions from the input text data, and this analysis identifies the emotional tone and mood of the text.
[1135] Step 5:
[1136] The server creates a storyboard based on the analysis results (text content and user emotions), which includes information on scenes and characters required for the promotional video.
[1137] Step 6:
[1138] The server selects corresponding visual elements (images, animations, text, etc.) according to the storyboard. It also integrates music and narration that match the user's emotions into the promotional video. These visual elements and emotional elements are integrated by the video generation engine and rendered as a promotional video.
[1139] Step 7:
[1140] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or pop-up screen). This transmission also uses a secure communication protocol.
[1141] Step 8:
[1142] The display device decodes the received video data and plays it as a promotional video, allowing bookstore visitors to visually and emotionally experience the atmosphere of the novel.
[1143] Specific examples
[1144] A user uses an interactive terminal installed in a bookstore to input the opening part of the novel "Adventure in the Mysterious Forest."
[1145] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1146] Step 1:
[1147] A user enters text into a web form and clicks the submit button.
[1148] Step 2:
[1149] The terminal sends this text data to the server as an HTTP POST request.
[1150] Step 3:
[1151] The server receives the text data and uses a natural language processing algorithm to extract key elements such as "boy," "mysterious forest," "animals," and "adventure."
[1152] Step 4:
[1153] The server uses an emotion engine to analyze the emotional tones contained in the text, such as "adventure" or "excitement."
[1154] Step 5:
[1155] Based on the analysis results, the server creates a storyboard that includes scenes of the boy getting lost in the forest and meeting animals.
[1156] Step 6:
[1157] The server selects the visual elements, adds adventurous music and narration, and renders the promotional video.
[1158] Step 7:
[1159] The server encodes the promotional video and transmits it to the bookstore's digital signage.
[1160] Step 8:
[1161] The digital signage plays the received video, visually conveying the atmosphere of the novel to customers.
[1162] The system allows shoppers to visually and emotionally experience the opening pages of the novel, helping them make purchasing decisions.
[1163] Example 2
[1164] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1165] Conventional promotional video generation systems simply convert text data entered by users into visual elements, making it difficult to generate videos that reflect the user's emotions and intentions. Furthermore, the accuracy of text analysis was insufficient, resulting in low-quality generated videos. Furthermore, even after the video was generated, it was often not properly transmitted to the display device in real time, resulting in display timing discrepancies.
[1166] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1167] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for analyzing emotions based on the analyzed text data, means for integrating specified visual elements based on the analyzed text data and the emotion analysis results to generate a promotional video, means for transmitting the generated promotional video to a predetermined display device, and means for playing the transmitted promotional video, thereby enabling the generation of high-quality promotional videos that reflect the user's input data and emotions.
[1168] "Input text data" refers to character information provided by a user to the system through a terminal.
[1169] "Natural language analysis" is the process by which a computer understands and analyzes human language.
[1170] "Sentiment analysis" is the process of identifying a user's emotional tone or mood from text data.
[1171] "Visual elements" refer to the visual elements such as images, animations, and text that make up the promotional video.
[1172] A "promotional video" is a short video content created for advertising or marketing purposes.
[1173] A "storyboard" is a blueprint that shows the scene composition and character placement when creating a video.
[1174] "Display device" refers to a device for playing promotional videos, such as digital signage or pop-up screens.
[1175] The system for implementing this invention receives text data of a novel entered by a user, analyzes the text, generates a promotional video, and finally displays it on a display device. This system is mainly composed of three elements: a server, a terminal, and a display device. The specific configuration and operation of this system are described below.
[1176] User Input
[1177] Users use their own devices (e.g., PCs, tablets, smartphones) to input text data, such as the opening of a novel, through a dedicated web form or application.
[1178] For example, suppose a user enters the opening line of the novel "Adventure in the Mysterious Forest" as follows:
[1179] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1180] Sending text data
[1181] The user's device sends the entered text data to the server as an HTTP POST request, using a secure communication protocol such as HTTPS.
[1182] Text analytics
[1183] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.) using tools such as the Google Cloud Natural Language API.
[1184] Emotion analysis
[1185] The server uses an emotion engine to analyze the user's emotions from the input text data. This analysis identifies the emotional tone and mood of the text, for example, using emotion recognition tools such as IBM Watson NLU.
[1186] Image Generation
[1187] The server creates a storyboard based on the text content and the results of user sentiment analysis. The storyboard includes information on the scenes and characters required for the promotional video. The server then selects visual elements (images, animations, text, etc.) according to the storyboard and incorporates music and narration that match the emotions into the promotional video. During this process, the promotional video is rendered using a video generation engine such as Adobe After Effects.
[1188] Video data transmission
[1189] The server then encodes the generated promotional video and transmits it to the designated display device, again using a secure communication protocol such as HTTPS.
[1190] Video display
[1191] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video, allowing customers to visually experience the atmosphere of the novel.
[1192] As a concrete example, we present a process for generating a promotional video for the opening of the novel "Adventure in the Mysterious Forest." This system allows bookstore visitors to visually and emotionally grasp the atmosphere of the novel beforehand, and can use this information to make a purchase decision.
[1193] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1194] Step 1:
[1195] User Input
[1196] Specific description:
[1197] Users use their own devices to input text data, such as the opening of a novel, through a dedicated web form or application.
[1198] input:
[1199] The text data to be input by the user (e.g., the opening part of a novel) is entered into the input field.
[1200] output:
[1201] Text data entered through the terminal is stored in a buffer.
[1202] Specific behavior:
[1203] A user visits a web form, enters text into an input field, and then clicks a "Submit" button when finished.
[1204] Step 2:
[1205] Sending text data
[1206] Specific description:
[1207] The user's device sends the entered text data to the server as an HTTP POST request, using a secure communication protocol such as HTTPS.
[1208] input:
[1209] Text data stored on the device.
[1210] output:
[1211] The text data is sent to the server as an HTTP POST request.
[1212] Specific behavior:
[1213] The device creates an HTTP POST request and sends the text data as a payload to the server.
[1214] Step 3:
[1215] Text analytics
[1216] Specific description:
[1217] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (such as characters, setting, and theme).
[1218] input:
[1219] The text data received by the server.
[1220] output:
[1221] Data from which key elements (characters, setting, theme, etc.) have been extracted.
[1222] Specific behavior:
[1223] The server stores the text data in a database, then uses NLP algorithms (e.g., Google Cloud Natural Language API) to analyze the text and extract key elements.
[1224] Step 4:
[1225] Emotion analysis
[1226] Specific description:
[1227] The server uses an emotion engine to analyze the user's emotions from the input text data, and this analysis identifies the emotional tone and mood of the text.
[1228] input:
[1229] The parsed text data.
[1230] output:
[1231] Data that identifies emotional tone and mood.
[1232] Specific behavior:
[1233] The server invokes an emotion recognition tool (e.g., IBM Watson NLU) to analyze the text data and identify the emotional tone.
[1234] Step 5:
[1235] Image Generation
[1236] Specific description:
[1237] The server creates a storyboard based on the analysis results (text content and user emotions), selects visual elements (images, animations, text, etc.) according to the storyboard, and incorporates music and narration that match the emotions into the promotional video. Finally, it renders the promotional video using a video generation engine.
[1238] input:
[1239] Data identified key elements and emotional tones.
[1240] output:
[1241] Promotional video data.
[1242] Specific behavior:
[1243] The server runs a script that generates a storyboard, selects appropriate images and animations based on the storyboard, adds music and narration, and renders the promotional video using a video generation engine (e.g., Adobe After Effects).
[1244] Step 6:
[1245] Video data transmission
[1246] Specific description:
[1247] The server then encodes the generated promotional video and transmits it to the designated display device, again using a secure communication protocol such as HTTPS.
[1248] input:
[1249] Promotional video data.
[1250] output:
[1251] The encoded video data is transmitted to a display device.
[1252] Specific behavior:
[1253] The server encodes the video file into the appropriate format and sends it to the IP address of the specified display device.
[1254] Step 7:
[1255] Video display
[1256] Specific description:
[1257] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video, allowing customers to visually experience the atmosphere of the novel.
[1258] input:
[1259] Video data sent to the display device.
[1260] output:
[1261] Promotional video displayed on display device.
[1262] Specific behavior:
[1263] The display device receives the video data sent from the server, decodes it, and plays it back using playback software.
[1264] (Application example 2)
[1265] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1266] In modern brick-and-mortar stores such as bookstores and convenience stores, there are limited means to effectively communicate the contents of a book to customers. In particular, it is difficult to visually and emotionally convey the appeal and emotional tone of the story, resulting in low book promotion effectiveness. Conventional methods only convey the contents of a book through text and still images, which means that customers do not receive sufficient information and it is difficult to stimulate their desire to purchase.
[1267] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1268] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for integrating designated visual elements and music and narration that match emotions based on the analyzed text data to generate a promotional video, means for transmitting the generated promotional video to a predetermined display device, and means for playing the transmitted promotional video, thereby making it possible to visually and emotionally convey the content and emotional tone of a book.
[1269] "Input text data" refers to novels or any other text data that a user inputs using a terminal.
[1270] The "receiving means" refers to a communication means for transmitting text data input by a user to a server and receiving the data.
[1271] "Means for natural language analysis" refers to means for analyzing received text data using natural language processing (NLP) algorithms to understand its content and components.
[1272] "Designated visual elements" refer to images or animations that are selected based on the content and emotional tone of the text data.
[1273] "Music and narration that matches the emotion" refers to music and narration selected to match the emotional tone analyzed from the text data.
[1274] The "means for generating a promotional video" refers to a means for generating a promotional video by integrating specified visual elements and music based on the analyzed text data and emotional tone.
[1275] The "means for transmitting to a display device" refers to a means for transmitting the generated promotional video data to a display device such as a digital signage or a smartphone.
[1276] The "means for playing" refers to a means for decoding the transmitted promotional video data and visually playing it back on a display device.
[1277] "Major elements" refer to important components of text data, such as characters, settings, and themes.
[1278] "Emotional tone" refers to the emotional tone or mood analyzed from text data.
[1279] A "storyboard" is a blueprint that visually organizes the scene and character information required to create a promotional video.
[1280] This invention provides a system for effectively generating and displaying promotional videos for books in brick-and-mortar stores such as bookstores, convenience stores, etc. This system is capable of processing user input, text data transmission, text analysis, emotion analysis, video generation, video data transmission, and video display.
[1281] Specifically, a user uses a smartphone application to input a novel or any other text data. This input data is securely sent to a server using the HTTPS protocol. The server then analyzes the received text data using natural language processing (NLP) algorithms. This analysis extracts the content and key elements of the text (such as characters, setting, and theme).
[1282] Next, the emotional engine analyzes the text for emotional tone, identifying the emotional tone and mood of the text. Based on this analysis, the server creates a storyboard and selects the specified visual elements (images and animations) as well as music and narration that match the emotion. This generates a promotional video.
[1283] The generated promotional video is encoded as video data and sent to a designated display device, such as a bookstore. This transmission is also done using a secure communication protocol. The display device decodes the received video data and plays it as a promotional video. This allows the appeal of the book to be conveyed visually and emotionally to store visitors.
[1284] A specific example is shown below.
[1285] Examples:
[1286] A user uses an interactive terminal installed in a bookstore to input the opening part of a novel. For example, "One day, a young boy named Taro gets lost in a mysterious forest. There, he meets various animals and embarks on a new adventure." The terminal then sends this text data to a server. The server analyzes the text and extracts the key elements—"boy," "mysterious forest," "animals," and "adventure." The emotion engine then interprets this text as "adventure" and "excitement." The server then creates a storyboard based on the analysis results and selects the appropriate visual elements (forest, boy, animals) as well as adventurous music and narration. Finally, the server integrates these elements to generate a promotional video, which is then sent to the bookstore's digital signage. The digital signage plays this promotional video, visually conveying the atmosphere of the novel to customers.
[1287] Example prompts to input to a generative AI model:
[1288] The opening of "Adventure in the Mysterious Forest":
[1289] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1290] Extract key elements from this text, analyze the emotional tone and generate a promotional video.
[1291] In this way, the present invention can enhance the effectiveness of book promotion in physical stores and stimulate customers' desire to purchase.
[1292] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1293] Step 1:
[1294] A user starts the smartphone application and enters the opening part of a novel or any other text data into the interface. The entered text data is temporarily saved within the application.
[1295] Input: Text data of the novel entered by the user.
[1296] Output: Text data that is temporarily stored on the device.
[1297] Step 2:
[1298] The terminal securely transmits the input text data to the server using the HTTPS protocol, where the data is sent to the server using an HTTP POST request.
[1299] Input: Text data stored in the device.
[1300] Output: The text data sent to the server.
[1301] Step 3:
[1302] The server applies natural language processing (NLP) algorithms to analyze the received text data, extracting key elements (characters, setting, theme, etc.) as a result of the analysis.
[1303] Input: The text data sent to the server.
[1304] Output: Key elements extracted by natural language processing.
[1305] Specific behavior: Performs text analysis using NLP libraries (e.g. spaCy, NLTK).
[1306] Step 4:
[1307] The server uses an emotion engine to analyze the emotional tone from the text data, and determines what emotion the input text evokes based on the emotional tone.
[1308] Input: The text data sent to the server.
[1309] Output: The emotional tone identified by the emotion engine (e.g., excitement, sadness, joy).
[1310] What it does: Identifies emotional tone using a sentiment analysis model (e.g., TextBlob, VADER).
[1311] Step 5:
[1312] The server creates a storyboard based on the analysis results, and selects corresponding visual elements (images, animations) as well as music and narration that match the emotion, thereby generating a promotional video.
[1313] Input: Key elements and emotional tone.
[1314] Output: The generated promo video.
[1315] Specific operation: Use a video generation engine (e.g. FFmpeg) to generate video according to the storyboard.
[1316] Step 6:
[1317] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or smartphone). This transmission also uses the HTTPS protocol.
[1318] Input: Generated promo video.
[1319] Output: The encoded video data sent to a display device.
[1320] Specific operation: Encodes video data and sends it via an HTTP POST request.
[1321] Step 7:
[1322] The display device decodes the received video data and visually reproduces it as a promotional video.
[1323] Input: Transmitted video data.
[1324] Output: The promotional video that will be played.
[1325] Specific operation: Decode and play video data using a decoding library (e.g., VLC, FFmpeg).
[1326] Through this series of steps, a promotional video generated based on the text data entered by the user is played on the display device, making it possible to provide visitors with a visually and emotionally effective promotion.
[1327] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1328] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1329] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1330] [Fourth embodiment]
[1331] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1332] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1333] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1334] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1335] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1336] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1337] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1338] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1339] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1340] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1341] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1342] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1343] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1344] The system for implementing this invention receives text data of a novel entered by a user, analyzes the text, generates a promotional video, and finally displays it on a digital signage or pop-up screen. This system is mainly composed of three elements: a server, a terminal, and a display device.
[1345] System configuration
[1346] 1. User Input
[1347] A user uses a device (e.g., a PC, tablet, or smartphone) to input text data, such as the opening of a novel, which is then entered into the device through a dedicated web form or application.
[1348] User: Enter the opening line of a novel and click the submit button.
[1349] 2. Sending text data
[1350] The user's device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[1351] Terminal: Sends text data to the server.
[1352] 3. Text Analysis
[1353] The server uses natural language processing (NLP) algorithms to analyze the received text data. NLP analyzes the content of the text and extracts key elements (e.g., characters, setting, theme, etc.).
[1354] Server: Analyzes the text data and extracts key elements.
[1355] 4. Image Generation
[1356] Based on the analysis results, the server creates a storyboard and selects corresponding visual elements (images, animations, text, etc.), then uses a video generation engine to render the promotional video.
[1357] Server: Creates storyboards, selects visual elements, and generates footage.
[1358] 5. Video data transmission
[1359] The server encodes the generated promotional video data and transmits it to the designated display device, also using a secure communication protocol.
[1360] Server: Encodes video data and sends it to the display device.
[1361] 6. Video display
[1362] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video.
[1363] Terminal: Displays the received video data.
[1364] Specific examples
[1365] A user uses an interactive terminal installed in a bookstore to input the opening part of the novel "Adventure in the Mysterious Forest."
[1366] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1367] The terminal transmits this text data to the server.
[1368] The server analyzes the text and extracts the keywords "boy," "mysterious forest," "animals," and "adventure."
[1369] The server creates a storyboard based on these keywords and selects appropriate visual elements to generate the promotional video.
[1370] The generated promotional video is sent to the bookstore's digital signage and displayed.
[1371] This system allows bookstore customers to visually get a feel for the atmosphere of a novel in advance and use it as a reference when making a purchase.
[1372] The processing flow will be explained below.
[1373] Step 1:
[1374] The user accesses a dedicated web form or application using a terminal, enters text data to be used in the promotional video, such as the opening of a novel, and clicks the submit button.
[1375] Step 2:
[1376] The device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[1377] Step 3:
[1378] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.).
[1379] Step 4:
[1380] The server creates a storyboard based on the analysis results, which includes information on the scenes and characters required for the promotional video.
[1381] Step 5:
[1382] The server selects corresponding visual elements (images, animations, text, etc.) according to the storyboard. These visual elements are integrated by the video generation engine and rendered as a promotional video.
[1383] Step 6:
[1384] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or pop-up screen). This transmission also uses a secure communication protocol.
[1385] Step 7:
[1386] The display device decodes the received video data and plays it as a promotional video, allowing bookstore visitors to visually experience the opening part of the novel.
[1387] Example 1
[1388] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1389] Traditionally, video generation for promoting novels and stories has been done manually, requiring a great deal of time and effort. Furthermore, specialized knowledge is required to select appropriate visual elements based on the content, making it difficult for anyone to easily create promotional videos. There has also been a lack of technology to instantly display generated videos on digital signage or pop-up screens. There is a need to solve these issues and automate the effective promotion of novels and stories quickly.
[1390] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1391] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for generating a promotional video by integrating specified visual elements based on the analyzed text data, means for transmitting the generated promotional video to a predetermined display device, means for playing the transmitted promotional video, and means for displaying the promotional video on a digital signage or pop-up screen. This makes it possible to automatically generate a promotional video based on the content of a novel or story and display it on a digital signage or pop-up screen quickly and effectively.
[1392] The "means for receiving input text data" is a function for transmitting text data input by a user using a terminal to a server and receiving the data.
[1393] "Means for natural language analysis" refers to algorithms or programs that analyze input text data, understand its content, and extract key elements and keywords.
[1394] The "means for generating promotional videos" is a function that combines the necessary visual elements based on the analyzed text data to create videos.
[1395] The "means for transmitting to a display device" is a communication function for encoding the generated promotional video data into an appropriate format and sending it to a display device such as a digital signage or a pop-up screen.
[1396] The "means for playing promotional video" is a function for the display device to decode the video data received, convert it into a playable format, and display it.
[1397] "Means for displaying on digital signage or a pop-up screen" refers to a function for actually displaying the generated promotional video on a digital signage or a pop-up screen.
[1398] "Major elements" are important items such as characters, settings, and themes that are extracted from the input text data.
[1399] "Visual elements" are visual elements such as images, animations, and text that make up the promotional video.
[1400] This invention is a system that inputs text data for a novel or story, analyzes its content, automatically generates a promotional video, and displays it on a display device. This system is mainly composed of three elements: a server, a terminal, and a display device. Each element works together to automate the promotional video generation process.
[1401] Hardware and software used
[1402] server:
[1403] Natural language processing is performed using Python's "spaCy" library and "Hugging Face"'s "Transformers" library.
[1404] Generate promotional videos using Unity and Blender.
[1405] Device:
[1406] A device such as a computer, tablet, or smartphone that accepts user input.
[1407] The ability to enter text data using a dedicated web form or application and click a submit button.
[1408] Display device:
[1409] Display devices such as digital signage and pop-up screens.
[1410] A decoding library and video player application for decoding and displaying received promotional videos.
[1411] System Operation
[1412] User Input
[1413] A user uses a terminal to input text data, such as the opening of a novel, which is received through a dedicated web form or application.
[1414] example:
[1415] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1416]
[1417] User: Enter the opening line of the novel and click the submit button.
[1418] Server processing
[1419] The server analyzes the received text data using natural language processing algorithms. Through this analysis, key elements such as characters, setting, and theme are extracted. A storyboard is then created based on the analysis results, and appropriate visual elements are selected. A promotional video is then generated based on the storyboard using Unity or Blender.
[1420] Transmission to a display device and display
[1421] The generated promotional video is encoded and transmitted from the server to the display device, which can then decode the received video data and play it back as the promotional video.
[1422] Specific examples
[1423] As an example, we will explain the creation of a promotional video for the novel "Adventure in the Mysterious Forest" using the following prompt sentence.
[1424] Example prompt sentence:
[1425] Prompt: Generate a promotional video for the novel "Adventure in the Mysterious Forest." Analyze the text below to extract characters, setting, and themes, and create a video with appropriate visual elements.
[1426] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1427] This system allows bookstore customers to visually get a feel for the atmosphere of a novel beforehand, which they can use as a reference when making a purchase. It also makes it possible to significantly reduce the effort required to create promotional materials.
[1428] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1429] Step 1:
[1430] The user uses a terminal to input the opening part of a novel. The input text data is received by a dedicated web form or application. Specifically, the user inputs the data in Unicode text format and clicks the submit button. The input data format is encoded in UTF-8.
[1431] input:
[1432] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1433] output:
[1434] UTF-8 encoded text data.
[1435] Step 2:
[1436] The device sends the entered text data to the server as an HTTP POST request. This transmission uses the HTTPS protocol to maintain data integrity. Specifically, the request body contains the text data, and the URL is the server's API endpoint.
[1437] input:
[1438] UTF-8 encoded text data.
[1439] output:
[1440] The HTTP POST request sent to the server.
[1441] Step 3:
[1442] The server analyzes the received text data using natural language processing (NLP) algorithms, such as the Python "spaCy" library and the "Transformers" library from "Hugging Face," to analyze the text content and extract key elements (e.g., characters, setting, and theme).
[1443] input:
[1444] The text data sent in the HTTP POST request.
[1445] output:
[1446] A list of the main elements (characters, setting, theme).
[1447] Step 4:
[1448] The server creates a storyboard based on the analyzed text data, determining the order and structure of scenes based on the extracted key elements, and selecting appropriate visual elements (images, animations, text, etc.).
[1449] input:
[1450] A list of the main elements (characters, setting, theme).
[1451] output:
[1452] Storyboard and selected visual elements.
[1453] Step 5:
[1454] The server generates promotional videos using Unity, Blender, etc. It combines storyboards and visual elements, renders each scene, and creates the final video file.
[1455] input:
[1456] Storyboards and visual elements.
[1457] output:
[1458] Generated promotional video file.
[1459] Step 6:
[1460] The server encodes and converts the generated promotional video data into a suitable format, and transmits the encoded video data to the appropriate display device using a secure communication protocol (HTTPS).
[1461] input:
[1462] Promotional video file.
[1463] output:
[1464] Encoded video data.
[1465] Step 7:
[1466] The display device decodes the received video data and plays it as a promotional video. It uses a specific decoding library to convert the received data into a playable format.
[1467] input:
[1468] Encoded video data.
[1469] output:
[1470] Decoded promotional video.
[1471] Through this series of processing steps, a promotional video can be automatically generated from the text of a novel entered by the user and quickly played back on a display device.
[1472] (Application example 1)
[1473] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1474] In the past, self-publishing authors and bloggers had to spend time and money creating promotional videos to effectively promote their work. It was also difficult for individuals and small publishers to produce high-quality advertising videos without large-scale production. This left many people without a means to widely publicize their work. The present invention solves this problem.
[1475] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1476] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for integrating specified visual elements based on the analyzed text data to generate a promotional video, means for generating a URL for the generated promotional video and transmitting it to a predetermined display device or user terminal, and means for playing and sharing the transmitted promotional video on the user terminal or on a social networking platform. This allows authors and bloggers to easily generate high-quality promotional videos for their works and promote them widely.
[1477] The "means for receiving input text data" is a function for transmitting text data input by a user to a server via a terminal, and for the server to receive the data.
[1478] "Means for natural language analysis" refers to a function that uses natural language processing technology to analyze received text data and understand the content and components of the text.
[1479] "Means for integrating visual elements to generate promotional videos" refers to a function that creates promotional videos by combining visual elements such as images, animations, and text based on information obtained through natural language analysis.
[1480] "Means for generating a URL for the generated promotional video and transmitting it to a specified display device or user terminal" refers to a function that generates an internet address (URL) for the created promotional video and transmits it to a display device or a terminal used by the user.
[1481] "Means for playing and sharing the transmitted promotional video on user devices or SNS platforms" refers to the function of making the promotional video available for viewing on user devices or SNS (social networking service) platforms and sharing it with other users.
[1482] The system for implementing this invention mainly comprises three main components: a user terminal, a server, and a display device. The role and specific processing of each component will be explained below.
[1483] User terminal
[1484] A user uses a device such as a smartphone, PC, or tablet to input text data for a novel or article. The input is done through a dedicated web form or application. For example, a user launches an application on their smartphone, enters the opening part of a novel in the text box, and clicks the submit button. This sends the input text data to the server as an HTTP POST request.
[1485] server
[1486] The server receives the input text data and analyzes it using natural language processing (NLP) algorithms. This analysis extracts key elements of the text (such as characters, setting, and theme). Specifically, an NLP engine (e.g., spaCy or Transformers) is used to understand the content of the text and identify important keywords and phrases. Next, visual elements (images, animations, text, etc.) are integrated to create a storyboard based on the analysis results. A video generation engine (e.g., FFmpeg) is used to generate a promotional video based on this storyboard.
[1487] After generating the video, the server generates a URL for the promotional video and sends this URL to the designated display device or user terminal. This process uses a secure communication protocol (e.g., HTTPS) to ensure data integrity, ensuring the secure transfer of video data.
[1488] Display device and playback
[1489] The video is played on a designated display device (digital signage, tablet, etc.) or user device using the URL of the received promotional video. Furthermore, the user device can share the generated promotional video on social media platforms (e.g., Facebook, Twitter), allowing the work to be disseminated to a wider audience.
[1490] Specific examples
[1491] For example, suppose a user uses a smartphone to input the opening part of the novel "Adventure in the Mysterious Forest."
[1492] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1493] When this text is entered, the server analyzes it and extracts keywords such as "boy," "mysterious forest," "animals," and "adventure." Based on the analysis results, the server creates a storyboard, selects appropriate visual elements, and generates a promotional video. The URL of the generated promotional video is sent to the user's device. Users can use the URL to share it on social media.
[1494] Prompt Sentence Examples
[1495] "Analyze the following text and extract its main elements. These elements can be characters, places, themes, etc."
[1496] Text: "One day, a young boy named Taro gets lost in a mysterious forest. There he meets various animals and embarks on a new adventure."
[1497] By following the above steps, this invention makes it possible to automatically generate high-quality promotional videos based on text data entered by the user and share them widely.
[1498] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1499] Step 1:
[1500] Users use devices such as smartphones, tablets, and PCs to enter the opening of a novel or the text they want to advertise. Input is done through an app or a web form and begins by clicking a submit button. The entered text data is sent to the server as an HTTP POST request.
[1501] Input: Text data (e.g., "One day, a young boy named Taro gets lost in a mysterious forest. There, he meets various animals and embarks on a new adventure.")
[1502] Output: A request with text data sent to the server
[1503] Step 2:
[1504] The server receives the text data sent by the user, stores it in a database, and passes it on to the next analysis step.
[1505] Input: Text data sent by the user
[1506] Output: Text data stored on the server
[1507] Step 3:
[1508] The server uses a natural language processing (NLP) engine (e.g., spaCy or Transformers) to analyze the received text data and extract key elements (e.g., characters, setting, theme, etc.).
[1509] Input: Text data stored on the server
[1510] Output: Extracted key elements (e.g., "boy," "mysterious forest," "animal," "adventure")
[1511] Step 4:
[1512] The server creates a storyboard based on the analysis results. The storyboard indicates which visual elements (images, animations, text, etc.) are used in which scenes. The visual elements are selected from a pre-asset library.
[1513] Input: Extracted key elements
[1514] Output: Storyboard
[1515] Step 5:
[1516] The server uses a video generation engine (e.g., FFmpeg) to generate promotional videos based on the storyboard, and the generated videos are stored on the server.
[1517] Input: Storyboard
[1518] Output: Promotional video
[1519] Step 6:
[1520] The server generates a URL for the generated promotional video and sends the URL to the user terminal. The communication uses a secure protocol (e.g., HTTPS).
[1521] Input: Promotional video
[1522] Output: Promotional video URL
[1523] Step 7:
[1524] The user's device will use the received URL to play the promotional video, and the user can also share the video on social media platforms or other media.
[1525] Input: Promotional video URL
[1526] Output: Promotional videos played and shared videos
[1527] The above steps enable the automatic generation and widespread sharing of promotional videos based on text data entered by the user.
[1528] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1529] The system for implementing this invention receives text data of a novel entered by a user, analyzes that text, generates a promotional video, and finally displays it on a digital signage or pop-up screen. It also incorporates an emotion engine that recognizes the user's emotions, and generates the promotional video based on the analysis results. This system is primarily composed of three elements: a server, a terminal, and a display device.
[1530] System configuration
[1531] 1. User Input
[1532] A user uses a device (e.g., a PC, tablet, or smartphone) to input text data, such as the opening of a novel, which is then entered into the device through a dedicated web form or application.
[1533] User: Enter the opening line of a novel and click the submit button.
[1534] 2. Sending text data
[1535] The user's device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[1536] Terminal: Sends text data to the server.
[1537] 3. Text Analysis
[1538] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.).
[1539] Server: Analyzes the text data and extracts key elements.
[1540] 4. Emotion analysis
[1541] The server uses an emotion engine to analyze the user's emotions from the input text data, and this analysis identifies the emotional tone and mood of the text.
[1542] Server: Analyzes user sentiment from text data.
[1543] 5. Image Generation
[1544] The server creates a storyboard based on the analysis results (text content and user emotions), which includes information on scenes and characters required for the promotional video.
[1545] The server selects corresponding visual elements (images, animations, text, etc.) according to the storyboard, and incorporates music and narration that match the user's emotions into the promotional video.Then, it uses a video generation engine to render the promotional video.
[1546] Server: Creates storyboards and generates footage by selecting music and narration that match the visual elements and emotions.
[1547] 6. Video data transmission
[1548] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or pop-up screen). This transmission also uses a secure communication protocol.
[1549] Server: Encodes video data and sends it to the display device.
[1550] 7. Video display
[1551] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video.
[1552] Terminal: Displays the received video data.
[1553] Specific examples
[1554] A user uses an interactive terminal installed in a bookstore to input the opening part of the novel "Adventure in the Mysterious Forest."
[1555] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1556] The terminal transmits this text data to the server.
[1557] The server analyzes the text and extracts the main elements: "boy," "mysterious forest," "animals," and "adventure."
[1558] The emotion engine analyzes the emotional tone of this text as "adventure" and "excitement."
[1559] The server creates a storyboard based on the analysis results, selecting the appropriate visual elements (forest, boy, animals) as well as adventurous music and narration.
[1560] The server combines these elements to generate a promotional video and transmits it to the bookstore's digital signage.
[1561] Digital signage plays this promotional video, visually conveying the atmosphere of the novel to customers.
[1562] This system allows bookstore visitors to get a visual and emotional feel for the novel beforehand, helping them make a purchasing decision.
[1563] The processing flow will be explained below.
[1564] Step 1:
[1565] The user accesses a dedicated web form or application using a terminal, enters text data to be used in the promotional video, such as the opening of a novel, and clicks the submit button.
[1566] Step 2:
[1567] The device sends the entered text data to the server as an HTTP POST request, maintaining data integrity and using a secure communication protocol.
[1568] Step 3:
[1569] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.).
[1570] Step 4:
[1571] The server uses an emotion engine to analyze the user's emotions from the input text data, and this analysis identifies the emotional tone and mood of the text.
[1572] Step 5:
[1573] The server creates a storyboard based on the analysis results (text content and user emotions), which includes information on scenes and characters required for the promotional video.
[1574] Step 6:
[1575] The server selects corresponding visual elements (images, animations, text, etc.) according to the storyboard. It also integrates music and narration that match the user's emotions into the promotional video. These visual elements and emotional elements are integrated by the video generation engine and rendered as a promotional video.
[1576] Step 7:
[1577] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or pop-up screen). This transmission also uses a secure communication protocol.
[1578] Step 8:
[1579] The display device decodes the received video data and plays it as a promotional video, allowing bookstore visitors to visually and emotionally experience the atmosphere of the novel.
[1580] Specific examples
[1581] A user uses an interactive terminal installed in a bookstore to input the opening part of the novel "Adventure in the Mysterious Forest."
[1582] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1583] Step 1:
[1584] A user enters text into a web form and clicks the submit button.
[1585] Step 2:
[1586] The terminal sends this text data to the server as an HTTP POST request.
[1587] Step 3:
[1588] The server receives the text data and uses a natural language processing algorithm to extract key elements such as "boy," "mysterious forest," "animals," and "adventure."
[1589] Step 4:
[1590] The server uses an emotion engine to analyze the emotional tones contained in the text, such as "adventure" or "excitement."
[1591] Step 5:
[1592] Based on the analysis results, the server creates a storyboard that includes scenes of the boy getting lost in the forest and meeting animals.
[1593] Step 6:
[1594] The server selects the visual elements, adds adventurous music and narration, and renders the promotional video.
[1595] Step 7:
[1596] The server encodes the promotional video and transmits it to the bookstore's digital signage.
[1597] Step 8:
[1598] The digital signage plays the received video, visually conveying the atmosphere of the novel to customers.
[1599] The system allows shoppers to visually and emotionally experience the opening pages of the novel, helping them make purchasing decisions.
[1600] Example 2
[1601] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1602] Conventional promotional video generation systems simply convert text data entered by users into visual elements, making it difficult to generate videos that reflect the user's emotions and intentions. Furthermore, the accuracy of text analysis was insufficient, resulting in low-quality generated videos. Furthermore, even after the video was generated, it was often not properly transmitted to the display device in real time, resulting in display timing discrepancies.
[1603] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1604] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for analyzing emotions based on the analyzed text data, means for integrating specified visual elements based on the analyzed text data and the emotion analysis results to generate a promotional video, means for transmitting the generated promotional video to a predetermined display device, and means for playing the transmitted promotional video, thereby enabling the generation of high-quality promotional videos that reflect the user's input data and emotions.
[1605] "Input text data" refers to character information provided by a user to the system through a terminal.
[1606] "Natural language analysis" is the process by which a computer understands and analyzes human language.
[1607] "Sentiment analysis" is the process of identifying a user's emotional tone or mood from text data.
[1608] "Visual elements" refer to the visual elements such as images, animations, and text that make up the promotional video.
[1609] A "promotional video" is a short video content created for advertising or marketing purposes.
[1610] A "storyboard" is a blueprint that shows the scene composition and character placement when creating a video.
[1611] "Display device" refers to a device for playing promotional videos, such as digital signage or pop-up screens.
[1612] The system for implementing this invention receives text data of a novel entered by a user, analyzes the text, generates a promotional video, and finally displays it on a display device. This system is mainly composed of three elements: a server, a terminal, and a display device. The specific configuration and operation of this system are described below.
[1613] User Input
[1614] Users use their own devices (e.g., PCs, tablets, smartphones) to input text data, such as the opening of a novel, through a dedicated web form or application.
[1615] For example, suppose a user enters the opening line of the novel "Adventure in the Mysterious Forest" as follows:
[1616] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1617] Sending text data
[1618] The user's device sends the entered text data to the server as an HTTP POST request, using a secure communication protocol such as HTTPS.
[1619] Text analytics
[1620] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (e.g., characters, setting, theme, etc.) using tools such as the Google Cloud Natural Language API.
[1621] Emotion analysis
[1622] The server uses an emotion engine to analyze the user's emotions from the input text data. This analysis identifies the emotional tone and mood of the text, for example, using emotion recognition tools such as IBM Watson NLU.
[1623] Image Generation
[1624] The server creates a storyboard based on the text content and the results of user sentiment analysis. The storyboard includes information on the scenes and characters required for the promotional video. The server then selects visual elements (images, animations, text, etc.) according to the storyboard and incorporates music and narration that match the emotions into the promotional video. During this process, the promotional video is rendered using a video generation engine such as Adobe After Effects.
[1625] Video data transmission
[1626] The server then encodes the generated promotional video and transmits it to the designated display device, again using a secure communication protocol such as HTTPS.
[1627] Video display
[1628] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video, allowing customers to visually experience the atmosphere of the novel.
[1629] As a concrete example, we present a process for generating a promotional video for the opening of the novel "Adventure in the Mysterious Forest." This system allows bookstore visitors to visually and emotionally grasp the atmosphere of the novel beforehand, and can use this information to make a purchase decision.
[1630] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1631] Step 1:
[1632] User Input
[1633] Specific description:
[1634] Users use their own devices to input text data, such as the opening of a novel, through a dedicated web form or application.
[1635] input:
[1636] The text data to be input by the user (e.g., the opening part of a novel) is entered into the input field.
[1637] output:
[1638] Text data entered through the terminal is stored in a buffer.
[1639] Specific behavior:
[1640] A user visits a web form, enters text into an input field, and then clicks a "Submit" button when finished.
[1641] Step 2:
[1642] Sending text data
[1643] Specific description:
[1644] The user's device sends the entered text data to the server as an HTTP POST request, using a secure communication protocol such as HTTPS.
[1645] input:
[1646] Text data stored on the device.
[1647] output:
[1648] The text data is sent to the server as an HTTP POST request.
[1649] Specific behavior:
[1650] The device creates an HTTP POST request and sends the text data as a payload to the server.
[1651] Step 3:
[1652] Text analytics
[1653] Specific description:
[1654] The server stores the received text data and analyzes it using natural language processing (NLP) algorithms to understand the content of the text and extract key elements (such as characters, setting, and theme).
[1655] input:
[1656] The text data received by the server.
[1657] output:
[1658] Data from which key elements (characters, setting, theme, etc.) have been extracted.
[1659] Specific behavior:
[1660] The server stores the text data in a database, then uses NLP algorithms (e.g., Google Cloud Natural Language API) to analyze the text and extract key elements.
[1661] Step 4:
[1662] Emotion analysis
[1663] Specific description:
[1664] The server uses an emotion engine to analyze the user's emotions from the input text data, and this analysis identifies the emotional tone and mood of the text.
[1665] input:
[1666] The parsed text data.
[1667] output:
[1668] Data that identifies emotional tone and mood.
[1669] Specific behavior:
[1670] The server invokes an emotion recognition tool (e.g., IBM Watson NLU) to analyze the text data and identify the emotional tone.
[1671] Step 5:
[1672] Image Generation
[1673] Specific description:
[1674] The server creates a storyboard based on the analysis results (text content and user emotions), selects visual elements (images, animations, text, etc.) according to the storyboard, and incorporates music and narration that match the emotions into the promotional video. Finally, it renders the promotional video using a video generation engine.
[1675] input:
[1676] Data identified key elements and emotional tones.
[1677] output:
[1678] Promotional video data.
[1679] Specific behavior:
[1680] The server runs a script that generates a storyboard, selects appropriate images and animations based on the storyboard, adds music and narration, and renders the promotional video using a video generation engine (e.g., Adobe After Effects).
[1681] Step 6:
[1682] Video data transmission
[1683] Specific description:
[1684] The server then encodes the generated promotional video and transmits it to the designated display device, again using a secure communication protocol such as HTTPS.
[1685] input:
[1686] Promotional video data.
[1687] output:
[1688] The encoded video data is transmitted to a display device.
[1689] Specific behavior:
[1690] The server encodes the video file into the appropriate format and sends it to the IP address of the specified display device.
[1691] Step 7:
[1692] Video display
[1693] Specific description:
[1694] The display device (digital signage or pop-up screen) decodes the received video data and plays it as a promotional video, allowing customers to visually experience the atmosphere of the novel.
[1695] input:
[1696] Video data sent to the display device.
[1697] output:
[1698] Promotional video displayed on display device.
[1699] Specific behavior:
[1700] The display device receives the video data sent from the server, decodes it, and plays it back using playback software.
[1701] (Application example 2)
[1702] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1703] In modern brick-and-mortar stores such as bookstores and convenience stores, there are limited means to effectively communicate the contents of a book to customers. In particular, it is difficult to visually and emotionally convey the appeal and emotional tone of the story, resulting in low book promotion effectiveness. Conventional methods only convey the contents of a book through text and still images, which means that customers do not receive sufficient information and it is difficult to stimulate their desire to purchase.
[1704] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1705] In this invention, the server includes means for receiving input text data, means for natural language analysis of the received text data, means for integrating designated visual elements and music and narration that match emotions based on the analyzed text data to generate a promotional video, means for transmitting the generated promotional video to a predetermined display device, and means for playing the transmitted promotional video, thereby making it possible to visually and emotionally convey the content and emotional tone of a book.
[1706] "Input text data" refers to novels or any other text data that a user inputs using a terminal.
[1707] The "receiving means" refers to a communication means for transmitting text data input by a user to a server and receiving the data.
[1708] "Means for natural language analysis" refers to means for analyzing received text data using natural language processing (NLP) algorithms to understand its content and components.
[1709] "Designated visual elements" refer to images or animations that are selected based on the content and emotional tone of the text data.
[1710] "Music and narration that matches the emotion" refers to music and narration selected to match the emotional tone analyzed from the text data.
[1711] The "means for generating a promotional video" refers to a means for generating a promotional video by integrating specified visual elements and music based on the analyzed text data and emotional tone.
[1712] The "means for transmitting to a display device" refers to a means for transmitting the generated promotional video data to a display device such as a digital signage or a smartphone.
[1713] The "means for playing" refers to a means for decoding the transmitted promotional video data and visually playing it back on a display device.
[1714] "Major elements" refer to important components of text data, such as characters, settings, and themes.
[1715] "Emotional tone" refers to the emotional tone or mood analyzed from text data.
[1716] A "storyboard" is a blueprint that visually organizes the scene and character information required to create a promotional video.
[1717] This invention provides a system for effectively generating and displaying promotional videos for books in brick-and-mortar stores such as bookstores, convenience stores, etc. This system is capable of processing user input, text data transmission, text analysis, emotion analysis, video generation, video data transmission, and video display.
[1718] Specifically, a user uses a smartphone application to input a novel or any other text data. This input data is securely sent to a server using the HTTPS protocol. The server then analyzes the received text data using natural language processing (NLP) algorithms. This analysis extracts the content and key elements of the text (such as characters, setting, and theme).
[1719] Next, the emotional engine analyzes the text for emotional tone, identifying the emotional tone and mood of the text. Based on this analysis, the server creates a storyboard and selects the specified visual elements (images and animations) as well as music and narration that match the emotion. This generates a promotional video.
[1720] The generated promotional video is encoded as video data and sent to a designated display device, such as a bookstore. This transmission is also done using a secure communication protocol. The display device decodes the received video data and plays it as a promotional video. This allows the appeal of the book to be conveyed visually and emotionally to store visitors.
[1721] A specific example is shown below.
[1722] Examples:
[1723] A user uses an interactive terminal installed in a bookstore to input the opening part of a novel. For example, "One day, a young boy named Taro gets lost in a mysterious forest. There, he meets various animals and embarks on a new adventure." The terminal then sends this text data to a server. The server analyzes the text and extracts the key elements—"boy," "mysterious forest," "animals," and "adventure." The emotion engine then interprets this text as "adventure" and "excitement." The server then creates a storyboard based on the analysis results and selects the appropriate visual elements (forest, boy, animals) as well as adventurous music and narration. Finally, the server integrates these elements to generate a promotional video, which is then sent to the bookstore's digital signage. The digital signage plays this promotional video, visually conveying the atmosphere of the novel to customers.
[1724] Example prompts to input to a generative AI model:
[1725] The opening of "Adventure in the Mysterious Forest":
[1726] One day, a young boy named Taro gets lost in a mysterious forest, where he meets various animals and embarks on a new adventure.
[1727] Extract key elements from this text, analyze the emotional tone and generate a promotional video.
[1728] In this way, the present invention can enhance the effectiveness of book promotion in physical stores and stimulate customers' desire to purchase.
[1729] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1730] Step 1:
[1731] A user starts the smartphone application and enters the opening part of a novel or any other text data into the interface. The entered text data is temporarily saved within the application.
[1732] Input: Text data of the novel entered by the user.
[1733] Output: Text data that is temporarily stored on the device.
[1734] Step 2:
[1735] The terminal securely transmits the input text data to the server using the HTTPS protocol, where the data is sent to the server using an HTTP POST request.
[1736] Input: Text data stored in the device.
[1737] Output: The text data sent to the server.
[1738] Step 3:
[1739] The server applies natural language processing (NLP) algorithms to analyze the received text data, extracting key elements (characters, setting, theme, etc.) as a result of the analysis.
[1740] Input: The text data sent to the server.
[1741] Output: Key elements extracted by natural language processing.
[1742] Specific behavior: Performs text analysis using NLP libraries (e.g. spaCy, NLTK).
[1743] Step 4:
[1744] The server uses an emotion engine to analyze the emotional tone from the text data, and determines what emotion the input text evokes based on the emotional tone.
[1745] Input: The text data sent to the server.
[1746] Output: The emotional tone identified by the emotion engine (e.g., excitement, sadness, joy).
[1747] What it does: Identifies emotional tone using a sentiment analysis model (e.g., TextBlob, VADER).
[1748] Step 5:
[1749] The server creates a storyboard based on the analysis results, and selects corresponding visual elements (images, animations) as well as music and narration that match the emotion, thereby generating a promotional video.
[1750] Input: Key elements and emotional tone.
[1751] Output: The generated promo video.
[1752] Specific operation: Use a video generation engine (e.g. FFmpeg) to generate video according to the storyboard.
[1753] Step 6:
[1754] The server encodes the generated promotional video and transmits it to the designated display device (digital signage or smartphone). This transmission also uses the HTTPS protocol.
[1755] Input: Generated promo video.
[1756] Output: The encoded video data sent to a display device.
[1757] Specific operation: Encodes video data and sends it via an HTTP POST request.
[1758] Step 7:
[1759] The display device decodes the received video data and visually reproduces it as a promotional video.
[1760] Input: Transmitted video data.
[1761] Output: The promotional video that will be played.
[1762] Specific operation: Decode and play video data using a decoding library (e.g., VLC, FFmpeg).
[1763] Through this series of steps, a promotional video generated based on the text data entered by the user is played on the display device, making it possible to provide visitors with a visually and emotionally effective promotion.
[1764] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1765] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1766] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1767] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1768] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1769] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1770] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1771] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1772] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1773] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1774] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1775] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1776] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1777] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1778] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1779] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1780] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1781] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1782] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1783] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1784] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1785] The following is further disclosed regarding the above embodiment.
[1786] (Claim 1)
[1787] means for receiving input text data;
[1788] means for natural language analysis of the received text data;
[1789] means for generating a promotional video by integrating designated visual elements based on the analyzed text data;
[1790] means for transmitting the generated promotional video to a predetermined display device;
[1791] means for playing the transmitted promotional video;
[1792] A system including:
[1793] (Claim 2)
[1794] 2. The system according to claim 1, wherein the natural language analysis means extracts main elements from text data.
[1795] (Claim 3)
[1796] 2. The system according to claim 1, wherein the means for integrating visual elements and generating a promotional video creates a storyboard and generates the video based on the storyboard.
[1797] "Example 1"
[1798] (Claim 1)
[1799] means for receiving input text data;
[1800] means for natural language analysis of the received text data;
[1801] means for generating a promotional video by integrating designated visual elements based on the analyzed text data;
[1802] means for transmitting the generated promotional video to a predetermined display device;
[1803] means for playing the transmitted promotional video;
[1804] a means for displaying the information on a digital signage or pop-up screen;
[1805] A system including:
[1806] (Claim 2)
[1807] 2. The system according to claim 1, wherein the natural language analysis means extracts main elements from text data.
[1808] (Claim 3)
[1809] 2. The system according to claim 1, wherein the means for integrating visual elements and generating a promotional video creates a storyboard and generates the video based on the storyboard.
[1810] "Application Example 1"
[1811] (Claim 1)
[1812] means for receiving input text data;
[1813] means for natural language analysis of the received text data;
[1814] means for generating a promotional video by integrating designated visual elements based on the analyzed text data;
[1815] a means for generating a URL for the generated promotional video and transmitting the URL to a predetermined display device or user terminal;
[1816] A means for playing and sharing the transmitted promotional video on a user terminal or an SNS platform;
[1817] A system including:
[1818] (Claim 2)
[1819] 2. The system according to claim 1, wherein the natural language analysis means extracts main elements from text data.
[1820] (Claim 3)
[1821] 2. The system according to claim 1, wherein the means for integrating visual elements and generating a promotional video creates a storyboard and generates the video based on the storyboard.
[1822] "Example 2: Combining Emotion Engines"
[1823] (Claim 1)
[1824] means for receiving input text data;
[1825] means for natural language analysis of the received text data;
[1826] means for analyzing emotions based on the analyzed text data;
[1827] a means for generating a promotional video by integrating designated visual elements based on the analyzed text data and the emotion analysis result;
[1828] means for transmitting the generated promotional video to a predetermined display device;
[1829] means for playing the transmitted promotional video;
[1830] A system including:
[1831] (Claim 2)
[1832] 2. The system according to claim 1, wherein the natural language analysis means extracts main elements from text data.
[1833] (Claim 3)
[1834] 2. The system according to claim 1, wherein the means for integrating visual elements and generating a promotional video creates a storyboard and generates the video based on the storyboard.
[1835] "Application example 2 when combining emotion engines"
[1836] (Claim 1)
[1837] means for receiving input text data;
[1838] means for natural language analysis of the received text data;
[1839] a means for generating a promotional video by integrating music and narration that match designated visual elements and emotions based on the analyzed text data;
[1840] means for transmitting the generated promotional video to a predetermined display device;
[1841] means for playing the transmitted promotional video;
[1842] A system including:
[1843] (Claim 2)
[1844] 2. The system according to claim 1, wherein the natural language analysis means extracts key elements and emotional tones from text data.
[1845] (Claim 3)
[1846] 2. The system according to claim 1, wherein the means for generating a promotional video by integrating the visual elements and music and narration that match the emotions creates a storyboard and generates the video based on the storyboard. [Explanation of symbols]
[1847] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving input text data; means for natural language analysis of the received text data; means for generating a promotional video by integrating designated visual elements based on the analyzed text data; means for transmitting the generated promotional video to a predetermined display device; means for playing the transmitted promotional video; A system including:
2. 2. The system according to claim 1, wherein the natural language analysis means extracts main elements from text data.
3. 2. The system according to claim 1, wherein the means for integrating visual elements and generating a promotional video creates a storyboard and generates the video based on the storyboard.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A