System
A system using OCR and sentiment analysis generates short videos with book information and music to promote book appeal on social media, addressing the decline of bookstores and manual content creation challenges.
Patent Information
- Application Number
- JP2024124057
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
The decline of bookstores reduces opportunities for people to encounter books by chance and diminishes the unique appeal of bookstores, necessitating a more widespread method to communicate the appeal of books.
A system that uses OCR technology to extract book information from images, performs sentiment analysis on user impressions, generates POP designs, selects appropriate music, integrates the design with music to create a short video, and distributes it to social media.
Enables users to easily and efficiently generate visually appealing short videos reflecting their reading experience, which can be shared on social media, thereby promoting book appeal.
Smart Images

Figure 2026022540000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In today's world, the number of bookstores is declining, reducing the number of opportunities for people to encounter books by chance. This has limited the ways in which people who love to read and those who have new opportunities to pick up a book can find books. Furthermore, the unique and attractive appeal of bookstores, such as point-of-purchase (POP), is also disappearing, creating a need for a more widespread method of communicating the appeal of books. [Means for solving the problem]
[0005] In order to solve the above problems, the present invention provides a system that includes an image analysis means that takes a picture of the cover of a book that has been read and uses OCR technology to extract the book title, author name, and cover image; a bibliographic information acquisition means that compares this with a bibliographic database to obtain accurate bibliographic information; a sentiment analysis means that uses sentiment analysis to analyze the impressions and keywords entered by the user; an electronic design generation means that automatically generates a POP design based on the bibliographic information and the results of the sentiment analysis; a music selection means that selects appropriate music; a short video generation means that integrates the generated design with the selected music to generate a short video; and a short video distribution means that distributes this video to SNS.
[0006] "Image analysis means" refers to the ability to extract book titles, author names, and visual elements from images using optical character recognition technology.
[0007] "Means for obtaining bibliographic information" refers to the function of obtaining accurate bibliographic information by comparing data extracted using optical character recognition technology with a book database.
[0008] "Emotion analysis means" refers to technology that analyzes the impressions and keywords entered by users and extracts the essence of their emotions.
[0009] "Electronic design generation means" refers to technology that automatically generates designed POP images based on bibliographic information and sentiment analysis results.
[0010] "Music selection means" refers to the function of selecting appropriate music from a database by referring to the results of emotion analysis.
[0011] "Short video generation means" refers to the technology that generates a short video by integrating a POP image with selected music.
[0012] "Short video distribution means" refers to the function of distributing the generated short videos to platforms such as social media. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] This invention is a system that takes a photo of the cover of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, compares this with a bibliographic database to obtain accurate bibliographic information, and then uses sentiment analysis to analyze the impressions and keywords entered by the user.It then automatically generates a POP design based on the bibliographic information and the results of the sentiment analysis, selects an appropriate song, and combines the generated design with the selected song to generate a short video, which is then distributed to social media.
[0035] Program processing
[0036] 1. Taking and uploading images
[0037] User: Take a photo of the cover of a book you have just read using your smartphone or other device.
[0038] Device: Upload the captured image to the server via the device app.
[0039] 2. OCR processing
[0040] Server: Receives the uploaded image and inputs it into the OCR engine, which extracts the book title, author name, and visual elements from the image.
[0041] 3. Obtaining bibliographic information
[0042] Server: Based on the extracted book title and author name, it checks the book database to get the correct bibliographic information. If no bibliographic information is found, it generates feedback to the user requesting manual input.
[0043] 4. Enter your thoughts
[0044] User: Enter their thoughts and keywords about the book they read into the input form in the device app.
[0045] 5. Emotion analysis
[0046] Server: Analyzes the user-entered comments and extracts the main essence using a sentiment analysis algorithm.
[0047] 6. POP Text Generation
[0048] Server: Generates POP text based on the results of sentiment analysis and bibliographic information.
[0049] 7. Design Generation
[0050] Server: Runs the automatic design engine and generates POP images based on bibliographic information, sentiment analysis results, and cover images.
[0051] 8. Music Selection
[0052] Server: Refers to the results of the sentiment analysis and selects appropriate songs from music databases such as LINE MUSIC.
[0053] 9. Short video generation
[0054] Server: Integrates POP images with selected music to generate short videos.
[0055] Device: Receives short video data sent from the server and converts it into a format that can be posted to social media.
[0056] 10. Social Media Distribution
[0057] User: Press the "Post" button on the device app and select the generated short video.
[0058] Device: Calls an API to upload the video to the social networking site of the user's choice, and notifies the user when the post is complete.
[0059] Specific examples
[0060] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the author's name from the image. It then compares the image with a bibliographic database to obtain accurate bibliographic information. If the user inputs a sentiment such as "This is a work filled with adventure and emotion," the sentiment analysis algorithm analyzes this and generates POP text. Using the keyword "adventure," a song that evokes the image of adventure is selected. Finally, the cover image, POP text, and selected song are integrated to generate a short video, which users can post on social media to share the appeal of the book with many people.
[0061] In this way, the system of the present invention allows users to easily create POPs like those found in bookstores at home and share them on social media.
[0062] The processing flow will be explained below.
[0063] Step 1:
[0064] Users take a photo of the cover of a book they have finished reading with their smartphone, and the image is saved in the device app.
[0065] Step 2:
[0066] The device compresses the captured image and uploads it to the server, along with the user ID and session information.
[0067] Step 3:
[0068] The server inputs the received image into an OCR engine, which extracts the book title, author name, and visual elements from the image.
[0069] Step 4:
[0070] The server compares the OCR-extracted book title and author name with a book database to obtain accurate bibliographic information. If no bibliographic information is found, it presents alternative title suggestions and asks the user for confirmation.
[0071] Step 5:
[0072] Users enter their impressions of the book they have read and keywords into the input form of the terminal app, and the input is automatically sent to the server.
[0073] Step 6:
[0074] The server analyzes the user's input and uses a sentiment analysis algorithm to extract the main essence, thereby determining which sentiment prevails.
[0075] Step 7:
[0076] The server generates POP text based on the results of sentiment analysis and bibliographic information, and the generated text is structured to reflect the emotional emphasis.
[0077] Step 8:
[0078] The server runs an automated design engine that generates a POP image based on bibliographic information, sentiment analysis results, and the cover image, incorporating key design elements from the cover.
[0079] Step 9:
[0080] The server then refers to the results of the emotion analysis and selects appropriate songs from a music database, which are chosen to match the emotional tone.
[0081] Step 10:
[0082] The server combines the POP image with the selected music and generates a short video using a video generation tool such as FFmpeg, which is then sent to the device.
[0083] Step 11:
[0084] The device receives the short video sent from the server and displays a confirmation screen to the user, who then confirms the video and it is ready to post.
[0085] Step 12:
[0086] The user presses the "Post" button on the device app and selects the social media platform (e.g., Instagram, Twitter) to post the generated short video to.
[0087] Step 13:
[0088] The device will call the upload API of the selected social networking site to upload the short video, and the user will receive a notification when the post is complete.
[0089] This completes the entire process, allowing users to easily generate a POP for the book and share it on social media.
[0090] Example 1
[0091] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0092] Conventional systems lack a way for users to efficiently and accurately record information about books they have read and share that content in an engaging way on social media. This requires manual input and editing, which takes a lot of time and effort. Furthermore, analyzing impressions and keywords and selecting design and music based on those must also be done manually, resulting in an inconsistent user experience. Furthermore, distributing the generated content to social media can be difficult.
[0093] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0094] In this invention, the server includes image analysis means, bibliographic information acquisition means, emotion analysis means, electronic design generation means, music selection means, short video generation means, short video distribution means, a terminal for uploading images to the server via a user interface, a terminal for accepting user feedback, means for comparing extracted book titles and author names with a database to acquire bibliographic information, text generation means for generating POP text based on the emotion analysis results, and means for the short video generation means to integrate the images, generated text, and selected music. This allows users to easily and efficiently record information about books they have read, create attractive short videos that reflect the emotion analysis results, and share them on social media.
[0095] An "image analysis tool" is a device or software that uses machine learning algorithms or optical character recognition techniques to extract textual information and visual elements from an image.
[0096] The "bibliographic information acquisition means" refers to a device or software that compares the extracted book title and author name with a database to acquire accurate bibliographic information.
[0097] "Emotion analysis means" refers to a device or software that analyzes text data entered by a user and extracts emotions and key essences.
[0098] An "electronic design generator" is a device or software that automatically creates an attractive design based on POP text and visual elements.
[0099] A "music selection means" is a device or software that selects appropriate music based on the results of emotion analysis.
[0100] The "short video generation means" refers to a device or software that integrates the generated design and selected music to create a short video.
[0101] "Short video distribution means" refers to devices or software used to upload and distribute completed short videos to platforms such as social media.
[0102] A "terminal that uploads images to a server via a user interface" is a device that has hardware or software for sending images taken by a user to a server.
[0103] A "terminal that accepts user feedback input" is a device that provides an input form in which users can input their feedback or keywords, and that is equipped with hardware or software that accepts that data.
[0104] "Means for obtaining bibliographic information by comparing extracted book titles and author names with a database" refers to devices or software that executes the process of using book titles and author names extracted using OCR technology to compare them with an external database and obtain accurate bibliographic information.
[0105] "Text generation means for generating POP text based on the results of sentiment analysis" refers to a device or software that uses the results of sentiment analysis to automatically generate text that succinctly conveys the features and appeal of a book.
[0106] This system photographs the cover image of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, compares this with a bibliographic database to obtain accurate bibliographic information, and then analyzes the impressions and keywords entered by the user using sentiment analysis.The system then automatically generates a POP design based on the bibliographic information and the results of the sentiment analysis, selects an appropriate song, and combines the generated design with the selected song to create a short video, which is then distributed to social media.
[0107] Hardware and software used
[0108] Image analysis methods
[0109] The server stores the cover image uploaded by the user in Google Cloud Storage, then uses the Google Cloud Vision OCR API to extract the book title, author name, and cover image. The OCR engine takes image data as input and outputs text information.
[0110] Bibliographic information acquisition method
[0111] The server sends the extracted book title and author name to the Google Books API to get the exact bibliographic information, which is then returned from the database and stored in an internal database.
[0112] Emotion analysis means
[0113] The server inputs the user's feedback text into the IBM Watson Natural Language Understanding API to extract sentiment and key points, and the analysis results are also stored in a database.
[0114] Electronic Design Generator
[0115] The server uses the Canva API to generate POP images based on bibliographic information, sentiment analysis results, and cover images. The resulting designs are visually appealing and have a compelling layout.
[0116] Music selection method
[0117] Based on the results of the emotion analysis, the server accesses the LINE MUSIC API to select the appropriate song, and records the selected song data in a database.
[0118] Short video creation method
[0119] The server then combines the POP image and the selected music using video editing software such as Adobe Premiere API to create a short video, which is then stored in Google Cloud Storage.
[0120] Short video distribution methods
[0121] The device converts the short video data received from the server into a format that can be posted to SNS. When the user taps the "Post" button in the app, the device calls the SNS's API and uploads the video.
[0122] Examples of concrete examples and prompts
[0123] For example, a user takes a photo of the cover of a "fantasy novel" with their smartphone. The server sends this image to Google Cloud Vision OCR, which extracts "fantasy novel" and the author's name from the image. Next, it uses the Google Books API to obtain accurate bibliographic information. If the user enters their impression that the work is "full of adventure and emotion," the server performs sentiment analysis using the IBM Watson Natural Language Understanding API and extracts the key essence, "adventure." The server selects music related to "adventure" from the LINE MUSIC API and generates a pop-up image using the Canva API. Finally, the pop-up image and music are integrated using the Adobe Premiere API to generate a short video, which the user can post on social media.
[0124] An example of a prompt is as follows:
[0125] "Please build a system that takes a photo of the cover of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, and compares it with a bibliographic database. Analyzes the user's impressions using sentiment analysis, automatically generates POPs, and distributes short videos with appropriate music on social media."
[0126] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0127] Step 1:
[0128] User: Uses smartphone camera to take a picture of the cover of the book they have just read. Checks that the image is clear. The input is the image of the book cover, and the output is the image data.
[0129] Step 2:
[0130] Terminal: The captured image data is received within a dedicated app, compressed, and uploaded to the server. The input is the image data, and the output is the compressed image data sent to the server.
[0131] Step 3:
[0132] Server: Receives uploaded image data and stores it in Google Cloud Storage. Sends the stored image data to the Google Cloud Vision OCR API to extract the book title, author name, and cover image. The input is compressed image data, and the output is extracted text information.
[0133] Step 4:
[0134] Server: Sends the book title and author name obtained through OCR processing to the Google Books API to obtain bibliographic information. Stores the bibliographic information in an internal database and verifies the accuracy of the book title and author name. The input is the extracted text information, and the output is the bibliographic information stored in the database.
[0135] Step 5:
[0136] User: Follow the prompts from the system and enter your thoughts and keywords about the book you read into the input form within the app. Once you've finished entering your thoughts and keywords, tap the "Submit" button. The input is text data of your thoughts and keywords, and the output is text data sent to the server.
[0137] Step 6:
[0138] Server: Sends the sentiment text data to the IBM Watson Natural Language Understanding API, extracts sentiment and key essence, and stores the resulting analysis data in a database. The input is sentiment text data, and the output is the sentiment analysis results.
[0139] Step 7:
[0140] Server: Automatically generates POP text based on the sentiment analysis results and bibliographic information. The generated text succinctly summarizes the "features of this book" and includes visually appealing phrases. The input is the sentiment analysis results and bibliographic information, and the output is the generated POP text.
[0141] Step 8:
[0142] Server: Automatically generates digital designs using the Canva API. Generates POP images based on bibliographic information, sentiment analysis results, and cover images, creating visually appealing layouts. The inputs are bibliographic information, sentiment analysis results, and cover images, and the output is a POP image.
[0143] Step 9:
[0144] Server: Based on the results of the sentiment analysis, the server uses the LINE MUSIC API to select appropriate songs. The selected song data is recorded in a database. The input is the sentiment analysis results, and the output is the selected song data.
[0145] Step 10:
[0146] Server: The POP images and selected music are integrated using video editing software such as Adobe Premiere API to generate a short video. The generated video data is stored in Google Cloud Storage. The input is the POP images and music data, and the output is the generated short video.
[0147] Step 11:
[0148] User: Press the "Post" button in the app and select the generated short video. Select the social media platform to post to (e.g. Instagram, Twitter) and press the "Send" button. The input is the short video to be uploaded, and the output is the video posted to the social media platform.
[0149] Step 12:
[0150] Terminal: Calls the API of the selected SNS and uploads the video. When the upload is complete, a completion notification is displayed to the user. The input is the short video data, and the output is the video uploaded to the SNS.
[0151] (Application example 1)
[0152] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0153] To effectively share their impressions and reviews of a book after reading it, users need a simple and efficient process for creating a video with appropriate visual content and music, and then sharing it on social networking services. However, current technology requires these steps to be performed individually, which is tedious and can lack accuracy and consistency. An integrated system is needed to solve this problem.
[0154] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0155] In this invention, the server includes an image analysis means, a bibliographic information acquisition means, a sentiment analysis means, an electronic design generation means, a music selection means, a short video generation means, a short video distribution means, a means for taking an image from a user terminal and inputting the data, and a means for sharing and distributing the generated short video to a social networking service. This makes it possible to automatically generate a short video containing visual content that combines bibliographic information and sentiment analysis results and appropriate music based on the user's impressions of a book they have finished reading, and to easily share the generated short video on a social networking service.
[0156] "Image analysis means" refers to technology for extracting information from images taken by a user.
[0157] "Bibliographic information acquisition means" refers to technology for acquiring accurate bibliographic information about a book based on the extracted information.
[0158] "Emotion analysis means" refers to technology that analyzes the impressions and keywords entered by users and understands their emotions.
[0159] "Electronic design generation means" refers to technology for automatically generating POP text and graphics based on sentiment analysis results and bibliographic information.
[0160] "Music selection means" refers to technology for selecting appropriate music based on the results of emotion analysis.
[0161] "Short video generation means" refers to the technology for generating short videos by combining POP images with selected music.
[0162] "Short video distribution means" refers to technology for distributing the generated short videos to social networking services, etc.
[0163] "Means for taking pictures and inputting data from a user device" refers to technology that allows users to take pictures of book covers using devices such as smartphones and input that data into the system.
[0164] "Means for sharing and distributing the generated short video on a social networking service" refers to technology that enables the generated short video to be easily posted and shared on a social networking service.
[0165] A system for implementing this invention utilizes book information and impressions based on a user's reading experience to automatically generate short videos containing visual content and music, and share them on social networking services.
[0166] First, the user takes a photo of the cover of the book they have just finished reading using their smartphone. The user device takes the image and uploads the data to a server via an application. This is where the user device, including the smartphone, comes into play.
[0167] The server then receives the uploaded image and performs optical character recognition (OCR) using image analysis techniques to extract the book title, author name, and visual elements from the image. The OCR engine used is Tesseract or similar.
[0168] The server then uses the extracted title and author name to search the book database to find the correct bibliographic information, or if the bibliographic information is not found, prompt the user to enter it manually.
[0169] The user then enters their thoughts about the book and keywords into the input form of the terminal application. The server receives the thoughts and performs sentiment analysis using a sentiment analysis engine such as the Google Cloud Natural Language API.
[0170] The server generates POP text based on the results of sentiment analysis and bibliographic information. Next, the electronic design generation means operates and automatically generates POP images based on the bibliographic information, sentiment analysis results, and cover image. An automatic design engine is used.
[0171] Furthermore, the results of the sentiment analysis are used to select appropriate songs from a music database. For example, songs that evoke adventures are selected based on impressions that evoke adventures. A suitable music database can be found on a general music distribution service.
[0172] Finally, the server combines the POP image with the selected music to generate a short video. This method of generating short videos may utilize a video editing library. The generated short video is sent from the server to the user's device, and the user can easily share it on social networking services by pressing the "post" button on their device. By utilizing the API of the SNS platform, users can seamlessly share content.
[0173] Specific examples
[0174] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the author's name from the image. It then compares the image with a bibliographic database to obtain accurate bibliographic information. If the user inputs their impression that the book is "full of adventure and emotion," the sentiment analysis algorithm analyzes this and generates a pop-up message with an "adventure" theme. A song that evokes the image of adventure is selected based on the keyword "adventure." Finally, a short video is generated by integrating the cover image, pop-up message, and selected song, which users can post on social media to share the appeal of the book with many people.
[0175] Example prompt sentence:
[0176] "You can take a photo of the cover of a fantasy novel and input your impression that it's 'a work filled with adventure and emotion.' The system will then use OCR to extract information about the book, perform a sentiment analysis, select an adventure-themed POP design and music, and create a short video to post on social media."
[0177] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0178] Step 1:
[0179] Users take a photo of the cover of a book they have finished reading using a device such as a smartphone, and the image is uploaded to a server via a device application.
[0180] Input: An image file taken on the user's device
[0181] Output: Image data uploaded to the server
[0182] Step 2:
[0183] The server receives the uploaded image and performs optical character recognition (OCR) using image analysis techniques. The OCR engine extracts the book title, author name, and visual elements from the image.
[0184] Input: Uploaded image data
[0185] Output: Text data of book title, author name, and visual elements
[0186] Specific operation: The OCR engine running on the server extracts text from the image and generates analysis results.
[0187] Step 3:
[0188] The server then searches the bibliographic database based on the extracted title and author name to retrieve the correct bibliographic information, and if no bibliographic information is found, generates feedback to the user requesting manual input.
[0189] Input: Text data of book title and author name
[0190] Output: Bibliographic information (e.g., publisher name, publication date, ISBN, etc.)
[0191] Specific operation: The server communicates with the bibliographic database to search and retrieve the corresponding bibliographic information.
[0192] Step 4:
[0193] Users enter their thoughts and keywords about the book they have just read into the input form in the terminal application.
[0194] Input: User-entered comments and keywords
[0195] Output: Text data of impressions and keywords
[0196] Specific operation: The user enters their thoughts and keywords in text format into the input form on the device.
[0197] Step 5:
[0198] The server receives the inputted impressions and performs sentiment analysis using the sentiment analysis engine, which extracts the main essence and classifies the emotions into numerical values and categories.
[0199] Input: Text data of impressions and keywords
[0200] Output: Sentiment analysis results (e.g., positive, negative, main essence)
[0201] What it does: The sentiment analysis engine analyzes the sentiment text and extracts sentiment and key topics.
[0202] Step 6:
[0203] The server generates POP text based on the results of sentiment analysis and bibliographic information. The generated POP text will highlight the book's appeal.
[0204] Input: Sentiment analysis results, bibliographic information
[0205] Output: POP text
[0206] Specific operation: The server uses templates and generation algorithms to automatically generate text that combines sentiment analysis results and bibliographic information.
[0207] Step 7:
[0208] Using an electronic design generation tool, a POP image is generated based on bibliographic information, sentiment analysis results, and cover image.
[0209] Input: Bibliographic information, sentiment analysis results, cover image
[0210] Output: POP image
[0211] How it works: The online design tool automatically generates visually appealing POP images based on bibliographic information and sentiment analysis results.
[0212] Step 8:
[0213] The results of the sentiment analysis are used to select appropriate songs from a music database.
[0214] Input: Sentiment analysis results
[0215] Output: Selected music files
[0216] Specific operation: The server searches the music database and selects music files that match the emotion analysis results.
[0217] Step 9:
[0218] The server combines the POP image with the selected music to generate a short video.
[0219] Input: POP image, selected music file
[0220] Output: Short video file
[0221] What it does: The video editing library merges the POP image with the selected music to generate a short video in a specific format.
[0222] Step 10:
[0223] The user's device receives the short video data sent from the server and converts it into a format that can be posted to SNS. When the user presses the "Post" button, the generated short video is shared and distributed to the selected social networking service.
[0224] Input: Short video file
[0225] Output: Content posted to social media
[0226] Specific operation: The terminal application converts the short video and calls the SNS API to post it.
[0227] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0228] This invention is a system that photographs the cover image of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, recognizes the user's emotions using an emotion engine, and based on this, compares it with a bibliographic database to obtain accurate bibliographic information.It then uses emotion analysis to analyze the impressions and keywords entered by the user, automatically generates a designed POP based on the bibliographic information and the results of the emotion analysis, selects an appropriate song, and combines the generated design with the selected song to generate a short video, which is then distributed to social media.
[0229] Program processing
[0230] 1. Taking and uploading images
[0231] Users take a photo of the cover of a book they have finished reading with their smartphone, and the image is saved in the device app.
[0232] The device compresses the captured image and uploads it to the server, along with the user ID and session information.
[0233] 2. OCR processing
[0234] The server inputs the received image into an OCR engine, which extracts the book title, author name, and visual elements from the image.
[0235] 3. Obtaining bibliographic information
[0236] The server compares the OCR-extracted book title and author name with a book database to obtain accurate bibliographic information. If no bibliographic information is found, it presents alternative title suggestions and asks the user for confirmation.
[0237] 4. Emotion recognition
[0238] Users enter their impressions of the book they have read and keywords into the input form of the terminal app, and the input is automatically sent to the server.
[0239] The server uses an emotion engine to analyze the impressions and keywords entered by the user and recognize emotions. The emotion engine extracts emotions using natural language processing technology. Furthermore, if the user's facial image or voice is used, emotions can also be recognized from this data.
[0240] 5. Emotion analysis
[0241] The server analyzes the user's impressions, including the emotion recognition results, and extracts the main essence using an emotion analysis algorithm.
[0242] 6. POP Text Generation
[0243] The server generates POP text based on the results of sentiment analysis and bibliographic information, and the generated text is structured to reflect the emotional emphasis.
[0244] 7. Design Generation
[0245] The server runs an automated design engine that generates a POP image based on bibliographic information, sentiment analysis results, and the cover image, incorporating key design elements from the cover.
[0246] 8. Music Selection
[0247] The server then refers to the results of the emotion analysis and selects appropriate songs from a music database, which are chosen to match the emotional tone.
[0248] 9. Short video generation
[0249] The server combines the POP image with the selected music and generates a short video using a video generation tool such as FFmpeg, which is then sent to the device.
[0250] 10. Social Media Distribution
[0251] The device receives the short video sent from the server and displays a confirmation screen to the user, who then confirms the video and it is ready to post.
[0252] The user presses the "Post" button on the device app and selects the social media platform (e.g., Instagram, Twitter) to post the generated short video to.
[0253] The device will call the upload API of the selected social networking site to upload the short video, and the user will receive a notification when the post is complete.
[0254] Specific examples
[0255] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the "author's name" from the image. It then compares the image with a bibliographic database to obtain accurate bibliographic information. If a user writes their impression of a book as "full of adventure and emotion," the emotion engine analyzes it and uses a sentiment analysis algorithm to extract the key essences of "adventure" and "emotion." From this result, a POP text is generated, such as "an inspiring story with adventurous elements."
[0256] The automated design engine then generates a POP image based on the text, the acquired bibliographic information, and the cover image. Furthermore, it selects an appropriate song from a music database based on the results of sentiment analysis. Finally, a short video is generated by integrating the POP image and the selected song, which users can easily post to social media.
[0257] In this way, the system of the present invention allows users to easily create POPs like those found in bookstores at home and share them on social media.By incorporating an emotion engine, it is possible to provide more personalized POPs that reflect the user's emotions.
[0258] The processing flow will be explained below.
[0259] Step 1:
[0260] Users take a photo of the cover of a book they have finished reading with their smartphone, and the image is saved in the device's app storage.
[0261] Step 2:
[0262] The device compresses the captured image and uploads it to the server, which includes the user ID and session information.
[0263] Step 3:
[0264] The server inputs the received image into an OCR engine, which extracts the book title, author name, and visual elements from the image.
[0265] Step 4:
[0266] The server compares the book title and author name extracted by OCR with a book database to obtain accurate bibliographic information. If the bibliographic information is not found as a result of the comparison with the book database, it presents several alternative title candidates and asks the user for confirmation.
[0267] Step 5:
[0268] The user enters their thoughts about the book they have read and keywords into the input form of the terminal app. The input information is converted into JSON format and sent to the server.
[0269] Step 6:
[0270] The server uses an emotion engine to analyze the impressions and keywords entered by the user and recognize emotions. The emotion engine uses natural language processing technology to extract emotions from text. If the user also provides facial images or voice data, the engine can also recognize emotions from that data.
[0271] Step 7:
[0272] The server generates POP text based on the recognized sentiment and bibliographic information. The generated text reflects the essence obtained from the sentiment analysis results.
[0273] Step 8:
[0274] The server runs an automatic design engine to generate POP images based on bibliographic information, sentiment analysis results, and cover images. The automatic design engine creates POP images using design templates.
[0275] Step 9:
[0276] The server then refers to the emotion analysis results and selects appropriate songs from a music database that match the emotional tone.
[0277] Step 10:
[0278] The server combines the POP image with the selected music and generates a short video using a video generation tool such as FFmpeg, which is then sent to the device.
[0279] Step 11:
[0280] The device receives the short video sent from the server and displays a confirmation screen to the user, who can then confirm the video and indicate that it is ready to be posted.
[0281] Step 12:
[0282] The user presses the "Post" button on the device app and selects the social media platform (e.g., Instagram, Twitter) to post the generated short video to.
[0283] Step 13:
[0284] The device will call the upload API of the selected social networking service to upload the short video, and a notification will be displayed to the user once the post is successful.
[0285] This completes the entire process, allowing users to easily generate a POP for the book and share it on social media.
[0286] Example 2
[0287] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0288] While much information is shared on social media these days, there are limited ways for users to easily and effectively share their reading experiences. There is a particular need for sharing book reviews and ratings in a visually appealing format, but existing technologies require manual input and processing of information, which is time-consuming. Furthermore, it is difficult to generate personalized content that reflects user sentiment using sentiment analysis technology.
[0289] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an image analysis means, a bibliographic information acquisition means, an emotion analysis means, an electronic design generation means, a music selection means, a short video generation means, a short video distribution means, an image compression and upload means, an OCR analysis means, a user interaction means, a voice and face image analysis means, a display confirmation means, and an SNS posting means. This enables users to easily and attractively visualize their impressions and reviews of books they have finished reading and share them on SNS.
[0290] "Image analysis means" refers to means having the function of analyzing an image and extracting information.
[0291] The "bibliographic information acquisition means" is a means having a function for acquiring accurate book information related to a book category.
[0292] An "emotion analysis means" is a means that has the function of analyzing and recognizing emotions from impressions and keywords entered by the user.
[0293] "Electronic design generation means" means a means having a function for generating a design electronically.
[0294] The "music selection means" is a means having a function for selecting appropriate music based on the emotion analysis results.
[0295] The "short video generation means" is a means having a function for generating a short video by integrating a POP image with a selected piece of music.
[0296] "Short video distribution means" refers to a means that has the function of distributing the generated short video to social media and other platforms.
[0297] The "image compression and uploading means" is a means having a function for compressing a captured image and uploading it to a server.
[0298] "OCR Analysis Means" means a means capable of extracting book titles, author names, and visual elements from an image using optical character recognition technology.
[0299] "User interaction means" refers to means that has the function of providing an interface for users to input their thoughts and keywords into the system.
[0300] The "voice and facial image analysis means" is a means having a function for analyzing emotions from the voice and facial image data provided by the user.
[0301] The "display confirmation means" is a means having a function of allowing the user to confirm the generated content.
[0302] "SNS posting means" refers to a means that has the function of allowing users to easily post content they have viewed to SNS.
[0303] This invention is a system that takes a photo of the cover of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, recognizes the user's emotions using an emotion engine, and obtains accurate bibliographic information by matching it with a book database. Furthermore, it analyzes the impressions and keywords entered by the user using emotion analysis, automatically generates a POP design based on the bibliographic information and the emotion analysis results, selects an appropriate song, and generates a short video by integrating the generated design with the selected song, which is then distributed to social media.
[0304] The system is programmed as follows:
[0305] A user uses their smartphone camera to take a photo of the cover of a book they have just read. The captured image is saved in the device app. The device compresses the image and uploads it to the server, along with the user ID and session information. The server then inputs the uploaded image into an OCR engine (e.g., Google Cloud Vision API) to extract the book title, author name, and visual elements from the image.
[0306] The server compares the book title and author name extracted by OCR with a book database (e.g., Google Books API, Open Library) to obtain accurate bibliographic information. If bibliographic information is not found, alternative title suggestions are presented and the user is asked for confirmation. The user then enters their thoughts and keywords about the book they read into the input form of the device app. The input is automatically sent to the server. The server uses an emotion engine (e.g., IBM Watson NLU) to analyze the thoughts and keywords entered by the user and recognize emotions. If the user provides facial images or voice, emotions can also be recognized from this data, if necessary.
[0307] The server generates POP text based on the results of sentiment analysis and bibliographic information. The generated text is structured to reflect the emotional emphasis. The server then runs an automated design engine (e.g., Adobe Creative Cloud API) to generate a POP image based on the bibliographic information, sentiment analysis results, and cover image. The POP image also incorporates key design elements from the cover.
[0308] The server then refers to the emotion analysis results and selects an appropriate song from a music database (e.g., Spotify API, Apple Music API). The song is selected to match the emotional tone. The server then combines the POP image with the selected song and generates a short video using a video generation tool such as FFmpeg. The generated video data is then sent to the device.
[0309] The user can view the short video sent from the server on their device. After viewing, the user presses the "Post" button on the device app to post the video to a social networking platform (e.g., Instagram, Twitter). The device then calls the upload API of the selected social networking platform and uploads the short video. Once posting is complete, a notification is displayed to the user.
[0310] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the author's name from the image. It then compares the image with a book database to obtain accurate bibliographic information. If the user writes their impression, "This is a work filled with adventure and emotion," the emotion engine analyzes it and extracts the key essences of "adventure" and "emotion" using a sentiment analysis algorithm. From these results, the server creates a POP text such as "An inspiring story with adventurous elements."
[0311] The automated design engine then generates a POP image based on the text, the acquired bibliographic information, and the cover image. Furthermore, it selects an appropriate song from a music database based on the results of sentiment analysis. Finally, a short video is generated by integrating the POP image and the selected song, which users can easily post to social media.
[0312] The following are examples of prompt sentences:
[0313] "Please explain in natural language the process of a system that takes a photo of a book cover, extracts the book title and author name using OCR technology, analyzes the emotions felt after reading using an emotion engine, compares it with a book database to obtain bibliographic information, generates POP text and images based on the emotion analysis results, selects music, creates a video using a short video creation tool (e.g., FFmpeg), and posts it to social media."
[0314] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0315] Step 1: Capture and upload images
[0316] The user takes a photo of the book cover using the smartphone camera. The captured image is automatically saved in the device app. The device compresses the saved image and uploads it to the server. When uploading, the user ID and session information are also sent. This sends the compressed image data and user information to the server.
[0317] Step 2: OCR
[0318] The server inputs the uploaded image into an OCR engine (e.g., optical character recognition software). The OCR engine extracts the book title, author name, and visual elements from the image. It receives image data as input and obtains text data of the book title, author name, and visual elements as output. Specifically, the OCR engine identifies character regions in the image and reads character data from those regions.
[0319] Step 3: Obtain bibliographic information
[0320] The server uses the book title and author name extracted by OCR to check against a book database (e.g., a book information service). This allows accurate bibliographic information to be obtained. It receives text data of the book title and author name as input, and obtains detailed book information (e.g., publication year, genre, summary) as output. Specifically, it sends a request to the book database via an API and obtains the corresponding book information.
[0321] Step 4: Emotion Recognition
[0322] The user enters their thoughts and keywords about the book they have read into an input form in the device app. The input is automatically sent to the server. The server uses an emotion engine (e.g., natural language processing software) to analyze the thoughts and keywords entered by the user and recognize the emotion. It receives the text data of the thoughts and keywords as input and obtains the type of emotion (e.g., joy, sadness) as output. Specifically, the emotion engine analyzes the text and runs an algorithm to classify the emotion.
[0323] Step 5: Sentiment Analysis
[0324] The server analyzes the emotion results recognized by the emotion engine and the user's impressions, and extracts the main essence using an emotion analysis algorithm. It receives emotion data and impression text as input, and obtains the extracted essence (e.g., adventure, emotion) as output. Specifically, it uses a text analysis algorithm to extract important keywords and themes from the impressions.
[0325] Step 6: POP Text Generation
[0326] The server generates POP text based on the results of sentiment analysis and bibliographic information. It receives the sentiment essence and bibliographic information as input and obtains a catchy slogan and description as output. Specifically, it uses a template engine to compose the text and adjust it to reflect the emotional emphasis.
[0327] Step 7: Design Generation
[0328] The server runs an automatic design engine (e.g., design software API) to generate POP images based on bibliographic information, sentiment analysis results, and cover images. It receives design elements (e.g., book title, author name, emotional essence) as input and obtains POP image data as output. Specifically, it automatically adjusts the layout and coloring to generate visually appealing images.
[0329] Step 8: Music Selection
[0330] The server refers to the emotion analysis results and selects appropriate songs from a music database (e.g., a music streaming service). It receives the emotion essence as input and obtains song information as output. Specifically, it searches for songs that match the emotion and executes an algorithm to select the most suitable one.
[0331] Step 9: Short video generation
[0332] The server combines the POP image with the selected music and generates a short video using a video generation tool (e.g., FFmpeg). It receives POP image data and music information as input and obtains a short video file as output. Specifically, it adjusts the timing of the image and music and encodes them as continuous visual content.
[0333] Step 10: Social Media Distribution
[0334] The device receives the short video sent from the server and displays a confirmation screen to the user. The user checks the video and presses the "Post" button to post the video to the SNS platform (e.g., SNS service API). The device receives the short video file as input and receives a notification that posting to the SNS has been completed as output. Specifically, the device calls the SNS upload API and displays a notification to the user when posting is successful.
[0335] (Application example 2)
[0336] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0337] Traditionally, creating recommended book POPs in bookstores required a lot of time and effort, and the content of the POPs was not emotionally personalized, limiting their appeal to customers. Furthermore, updating promotional materials displayed on in-store displays was done manually, resulting in a lack of immediacy. For these reasons, a method was needed for bookstore staff to quickly and effectively promote books.
[0338] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an image analysis means, a bibliographic information acquisition means, a sentiment analysis means, an electronic design generation means, a music selection means, a short video generation means, a short video distribution means, and an automatic display means. This allows bookstore staff to simply take a photo of a book cover with their smartphone, and automatically generate an individual recommended POP based on related bibliographic information, user reviews, and sentiment analysis, and display it in real time on an electronic display in the store or post it to social media.
[0339] "Image analysis means" refers to means for extracting information from captured images.
[0340] "Bibliographic information acquisition means" refers to a means for acquiring information related to a book (e.g., book title, author name) from a database.
[0341] "Sentiment analysis means" means means for analyzing emotions from user input or other data.
[0342] "Electronic design generation means" refers to a means for automatically generating designs based on book information and sentiment analysis results.
[0343] The "music selection means" is a means for selecting appropriate music based on the emotion analysis results.
[0344] The "short video generation means" is a means for integrating a POP image with selected music to generate a short video.
[0345] "Short video distribution means" refers to a means for distributing the generated short video to an SNS platform.
[0346] "Automatic display means" refers to a means for displaying the generated POP images and short videos on electronic displays in the store in real time.
[0347] This system allows bookstore staff to take a photo of a book cover with their smartphone, automatically generating recommended POPs based on related bibliographic information, user reviews, and sentiment analysis, and posting them on in-store electronic displays and social media. The details of this system are described below.
[0348] The system starts by having the user take a photo of the book cover with their smartphone. The image is compressed and uploaded to a server, where it is analyzed using image analysis tools (e.g., pytesseract) to extract the book title, author, and visual elements.
[0349] Next, the server uses the bibliographic information acquisition means to acquire accurate bibliographic information from the book database, including the book title, author name, publication year, genre, etc.
[0350] The server then uses sentiment analysis tools to analyze the reviews and keywords entered by the user and recognize emotions. This process includes an emotion engine using natural language processing techniques. Facial images and voice data may also be used for emotion recognition.
[0351] Based on the results of the sentiment analysis and the bibliographic information, the server automatically generates POP text using an electronic design generation tool, and then creates a POP image using a design engine. This POP image reflects the bibliographic information and the user's sentiment.
[0352] Furthermore, the server uses a music selection means to select appropriate music from a music database based on the result of the emotion analysis, and the music is selected to match the emotional tone.
[0353] The server combines these POP images with the selected music and generates a short video using a short video generation tool (e.g., FFmpeg). This video is automatically generated and delivered to the user's smartphone.
[0354] Finally, the generated short video is posted to a social networking site using a short video distribution method. Users can select a social networking site (e.g., Instagram or Twitter) on their smartphone to complete the posting.
[0355] To accommodate in-store promotions, the system is equipped with an automatic display means, which displays the generated POP images on electronic displays in the store in real time.
[0356] Hardware and Software Use
[0357] The main hardware used in this invention is a smartphone and an electronic display in a bookstore, and the main software is an image analysis engine (pytesseract), a sentiment analysis engine, a design generation API, a video generation tool (FFmpeg), and a SNS upload API.
[0358] Specific examples
[0359] For example, if a user takes a photo of the cover of a "fantasy novel" and writes in their review that it is "a work filled with adventure and emotion," the system will analyze this information and generate a pop-up image that emphasizes "adventure" and "emotion." This image is then combined with a selected song to create a short video. The video can then be displayed on in-store displays and posted to social media.
[0360] Prompt Sentence Examples
[0361] "Users take a photo of the cover of a book they have read and enter their thoughts on the book. The system analyzes the information and generates a short video that combines recommended pop music and posts it to social media."
[0362] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0363] Step 1:
[0364] Image capture and upload
[0365] The user takes a photo of the book cover with their smartphone. This image is saved in the smartphone app. The device compresses the image file and uploads it to the server. The user ID and session information are also sent at the time of uploading.
[0366] Step 2:
[0367] OCR processing
[0368] The server inputs the received image into an OCR engine (pytesseract), which extracts text from the image and obtains information such as the book title and author name. The extracted text data is sent to the next processing step.
[0369] Step 3:
[0370] Bibliographic information acquisition
[0371] The server uses the book title and author name extracted by OCR to retrieve detailed bibliographic information from a book database. Using the book title and author name as input data, the server compares the book title, author name, publication year, genre, and other bibliographic information to output.
[0372] Step 4:
[0373] emotion recognition
[0374] The user enters their thoughts and keywords about the book they have read into an input form on a smartphone app. This text data is sent to a server. The server uses an emotion analysis tool (emotion engine) to analyze the thoughts and keywords entered by the user and recognize the emotion. The emotion analysis engine takes this data as input and outputs the type of emotion (e.g., joy, sadness).
[0375] Step 5:
[0376] Emotion analysis
[0377] The server analyzes the sentiment data, including the emotion recognition results, and uses a sentiment analysis algorithm to extract the key essence and generate the data needed to generate POPs. The input is the emotion recognition results and user sentiment data, and the output is the essence extraction results.
[0378] Step 6:
[0379] POP Text Generation
[0380] The server automatically generates POP text based on the results of sentiment analysis and bibliographic information. The generated text reflects the emotional emphasis. The input is the sentiment analysis results and bibliographic information, and the output is POP text.
[0381] Step 7:
[0382] Design Generation
[0383] The server runs an electronic design generator to automatically generate a POP image based on the sentiment analysis results, bibliographic information, and cover image. This POP image also incorporates the main design elements of the cover. The input is the POP text, bibliographic information, and cover image, and the output is the POP image.
[0384] Step 8:
[0385] Music Selection
[0386] The server refers to the results of the emotion analysis and selects an appropriate song from a music database. The selected song matches the emotional tone. The input is the emotion analysis result, and the output is the selected song.
[0387] Step 9:
[0388] Short video generation
[0389] The server combines the POP image and the selected music and generates a short video using a short video generation tool (FFmpeg). The input is the POP image and music URL, and the output is the short video data.
[0390] Step 10:
[0391] SNS distribution
[0392] The device receives the short video sent from the server and displays a confirmation screen to the user. The user reviews the video and is ready to post. When the user selects an SNS platform (e.g., Instagram, Twitter) and presses the "Post" button, the device calls the upload API of the selected SNS and uploads the short video. The input is the short video data and the SNS information selected by the user, and the output is a notification of successful posting to the SNS.
[0393] Step 11:
[0394] Automatic display
[0395] The server uses an automatic display means to display the generated POP image on an electronic display in the store in real time. The input is the POP image, and the output is the display on the in-store display.
[0396] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0397] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0398] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0399] [Second embodiment]
[0400] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0401] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0402] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0403] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0404] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0405] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0406] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0407] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0408] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0409] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0410] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0411] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0412] This invention is a system that takes a photo of the cover of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, compares this with a bibliographic database to obtain accurate bibliographic information, and then uses sentiment analysis to analyze the impressions and keywords entered by the user.It then automatically generates a POP design based on the bibliographic information and the results of the sentiment analysis, selects an appropriate song, and combines the generated design with the selected song to generate a short video, which is then distributed to social media.
[0413] Program processing
[0414] 1. Taking and uploading images
[0415] User: Take a photo of the cover of a book you have just read using your smartphone or other device.
[0416] Device: Upload the captured image to the server via the device app.
[0417] 2. OCR processing
[0418] Server: Receives the uploaded image and inputs it into the OCR engine, which extracts the book title, author name, and visual elements from the image.
[0419] 3. Obtaining bibliographic information
[0420] Server: Based on the extracted book title and author name, it checks the book database to get the correct bibliographic information. If no bibliographic information is found, it generates feedback to the user requesting manual input.
[0421] 4. Enter your thoughts
[0422] User: Enter their thoughts and keywords about the book they read into the input form in the device app.
[0423] 5. Emotion analysis
[0424] Server: Analyzes the user-entered comments and extracts the main essence using a sentiment analysis algorithm.
[0425] 6. POP Text Generation
[0426] Server: Generates POP text based on the results of sentiment analysis and bibliographic information.
[0427] 7. Design Generation
[0428] Server: Runs the automatic design engine and generates POP images based on bibliographic information, sentiment analysis results, and cover images.
[0429] 8. Music Selection
[0430] Server: Refers to the results of the sentiment analysis and selects appropriate songs from music databases such as LINE MUSIC.
[0431] 9. Short video generation
[0432] Server: Integrates POP images with selected music to generate short videos.
[0433] Device: Receives short video data sent from the server and converts it into a format that can be posted to social media.
[0434] 10. Social Media Distribution
[0435] User: Press the "Post" button on the device app and select the generated short video.
[0436] Device: Calls an API to upload the video to the social networking site of the user's choice, and notifies the user when the post is complete.
[0437] Specific examples
[0438] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the author's name from the image. It then compares the image with a bibliographic database to obtain accurate bibliographic information. If the user inputs a sentiment such as "This is a work filled with adventure and emotion," the sentiment analysis algorithm analyzes this and generates POP text. Using the keyword "adventure," a song that evokes the image of adventure is selected. Finally, the cover image, POP text, and selected song are integrated to generate a short video, which users can post on social media to share the appeal of the book with many people.
[0439] In this way, the system of the present invention allows users to easily create POPs like those found in bookstores at home and share them on social media.
[0440] The processing flow will be explained below.
[0441] Step 1:
[0442] Users take a photo of the cover of a book they have finished reading with their smartphone, and the image is saved in the device app.
[0443] Step 2:
[0444] The device compresses the captured image and uploads it to the server, along with the user ID and session information.
[0445] Step 3:
[0446] The server inputs the received image into an OCR engine, which extracts the book title, author name, and visual elements from the image.
[0447] Step 4:
[0448] The server compares the OCR-extracted book title and author name with a book database to obtain accurate bibliographic information. If no bibliographic information is found, it presents alternative title suggestions and asks the user for confirmation.
[0449] Step 5:
[0450] Users enter their impressions of the book they have read and keywords into the input form of the terminal app, and the input is automatically sent to the server.
[0451] Step 6:
[0452] The server analyzes the user's input and uses a sentiment analysis algorithm to extract the main essence, thereby determining which sentiment prevails.
[0453] Step 7:
[0454] The server generates POP text based on the results of sentiment analysis and bibliographic information, and the generated text is structured to reflect the emotional emphasis.
[0455] Step 8:
[0456] The server runs an automated design engine that generates a POP image based on bibliographic information, sentiment analysis results, and the cover image, incorporating key design elements from the cover.
[0457] Step 9:
[0458] The server then refers to the results of the emotion analysis and selects appropriate songs from a music database, which are chosen to match the emotional tone.
[0459] Step 10:
[0460] The server combines the POP image with the selected music and generates a short video using a video generation tool such as FFmpeg, which is then sent to the device.
[0461] Step 11:
[0462] The device receives the short video sent from the server and displays a confirmation screen to the user, who then confirms the video and it is ready to post.
[0463] Step 12:
[0464] The user presses the "Post" button on the device app and selects the social media platform (e.g., Instagram, Twitter) to post the generated short video to.
[0465] Step 13:
[0466] The device will call the upload API of the selected social networking site to upload the short video, and the user will receive a notification when the post is complete.
[0467] This completes the entire process, allowing users to easily generate a POP for the book and share it on social media.
[0468] Example 1
[0469] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0470] Conventional systems lack a way for users to efficiently and accurately record information about books they have read and share that content in an engaging way on social media. This requires manual input and editing, which takes a lot of time and effort. Furthermore, analyzing impressions and keywords and selecting design and music based on those must also be done manually, resulting in an inconsistent user experience. Furthermore, distributing the generated content to social media can be difficult.
[0471] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0472] In this invention, the server includes image analysis means, bibliographic information acquisition means, emotion analysis means, electronic design generation means, music selection means, short video generation means, short video distribution means, a terminal for uploading images to the server via a user interface, a terminal for accepting user feedback, means for comparing extracted book titles and author names with a database to acquire bibliographic information, text generation means for generating POP text based on the emotion analysis results, and means for the short video generation means to integrate the images, generated text, and selected music. This allows users to easily and efficiently record information about books they have read, create attractive short videos that reflect the emotion analysis results, and share them on social media.
[0473] An "image analysis tool" is a device or software that uses machine learning algorithms or optical character recognition techniques to extract textual information and visual elements from an image.
[0474] The "bibliographic information acquisition means" refers to a device or software that compares the extracted book title and author name with a database to acquire accurate bibliographic information.
[0475] "Emotion analysis means" refers to a device or software that analyzes text data entered by a user and extracts emotions and key essences.
[0476] An "electronic design generator" is a device or software that automatically creates an attractive design based on POP text and visual elements.
[0477] A "music selection means" is a device or software that selects appropriate music based on the results of emotion analysis.
[0478] The "short video generation means" refers to a device or software that integrates the generated design and selected music to create a short video.
[0479] "Short video distribution means" refers to devices or software used to upload and distribute completed short videos to platforms such as social media.
[0480] A "terminal that uploads images to a server via a user interface" is a device that has hardware or software for sending images taken by a user to a server.
[0481] A "terminal that accepts user feedback input" is a device that provides an input form in which users can input their feedback or keywords, and that is equipped with hardware or software that accepts that data.
[0482] "Means for obtaining bibliographic information by comparing extracted book titles and author names with a database" refers to devices or software that executes the process of using book titles and author names extracted using OCR technology to compare them with an external database and obtain accurate bibliographic information.
[0483] "Text generation means for generating POP text based on the results of sentiment analysis" refers to a device or software that uses the results of sentiment analysis to automatically generate text that succinctly conveys the features and appeal of a book.
[0484] This system photographs the cover image of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, compares this with a bibliographic database to obtain accurate bibliographic information, and then analyzes the impressions and keywords entered by the user using sentiment analysis.The system then automatically generates a POP design based on the bibliographic information and the results of the sentiment analysis, selects an appropriate song, and combines the generated design with the selected song to create a short video, which is then distributed to social media.
[0485] Hardware and software used
[0486] Image analysis methods
[0487] The server stores the cover image uploaded by the user in Google Cloud Storage, then uses the Google Cloud Vision OCR API to extract the book title, author name, and cover image. The OCR engine takes image data as input and outputs text information.
[0488] Bibliographic information acquisition method
[0489] The server sends the extracted book title and author name to the Google Books API to get the exact bibliographic information, which is then returned from the database and stored in an internal database.
[0490] Emotion analysis means
[0491] The server inputs the user's feedback text into the IBM Watson Natural Language Understanding API to extract sentiment and key points, and the analysis results are also stored in a database.
[0492] Electronic Design Generator
[0493] The server uses the Canva API to generate POP images based on bibliographic information, sentiment analysis results, and cover images. The resulting designs are visually appealing and have a compelling layout.
[0494] Music selection method
[0495] Based on the results of the emotion analysis, the server accesses the LINE MUSIC API to select the appropriate song, and records the selected song data in a database.
[0496] Short video creation method
[0497] The server then combines the POP image and the selected music using video editing software such as Adobe Premiere API to create a short video, which is then stored in Google Cloud Storage.
[0498] Short video distribution methods
[0499] The device converts the short video data received from the server into a format that can be posted to SNS. When the user taps the "Post" button in the app, the device calls the SNS's API and uploads the video.
[0500] Examples of concrete examples and prompts
[0501] For example, a user takes a photo of the cover of a "fantasy novel" with their smartphone. The server sends this image to Google Cloud Vision OCR, which extracts "fantasy novel" and the author's name from the image. Next, it uses the Google Books API to obtain accurate bibliographic information. If the user enters their impression that the work is "full of adventure and emotion," the server performs sentiment analysis using the IBM Watson Natural Language Understanding API and extracts the key essence, "adventure." The server selects music related to "adventure" from the LINE MUSIC API and generates a pop-up image using the Canva API. Finally, the pop-up image and music are integrated using the Adobe Premiere API to generate a short video, which the user can post on social media.
[0502] An example of a prompt is as follows:
[0503] "Please build a system that takes a photo of the cover of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, and compares it with a bibliographic database. Analyzes the user's impressions using sentiment analysis, automatically generates POPs, and distributes short videos with appropriate music on social media."
[0504] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0505] Step 1:
[0506] User: Uses smartphone camera to take a picture of the cover of the book they have just read. Checks that the image is clear. The input is the image of the book cover, and the output is the image data.
[0507] Step 2:
[0508] Terminal: The captured image data is received within a dedicated app, compressed, and uploaded to the server. The input is the image data, and the output is the compressed image data sent to the server.
[0509] Step 3:
[0510] Server: Receives uploaded image data and stores it in Google Cloud Storage. Sends the stored image data to the Google Cloud Vision OCR API to extract the book title, author name, and cover image. The input is compressed image data, and the output is extracted text information.
[0511] Step 4:
[0512] Server: Sends the book title and author name obtained through OCR processing to the Google Books API to obtain bibliographic information. Stores the bibliographic information in an internal database and verifies the accuracy of the book title and author name. The input is the extracted text information, and the output is the bibliographic information stored in the database.
[0513] Step 5:
[0514] User: Follow the prompts from the system and enter your thoughts and keywords about the book you read into the input form within the app. Once you've finished entering your thoughts and keywords, tap the "Submit" button. The input is text data of your thoughts and keywords, and the output is text data sent to the server.
[0515] Step 6:
[0516] Server: Sends the sentiment text data to the IBM Watson Natural Language Understanding API, extracts sentiment and key essence, and stores the resulting analysis data in a database. The input is sentiment text data, and the output is the sentiment analysis results.
[0517] Step 7:
[0518] Server: Automatically generates POP text based on the sentiment analysis results and bibliographic information. The generated text succinctly summarizes the "features of this book" and includes visually appealing phrases. The input is the sentiment analysis results and bibliographic information, and the output is the generated POP text.
[0519] Step 8:
[0520] Server: Automatically generates digital designs using the Canva API. Generates POP images based on bibliographic information, sentiment analysis results, and cover images, creating visually appealing layouts. The inputs are bibliographic information, sentiment analysis results, and cover images, and the output is a POP image.
[0521] Step 9:
[0522] Server: Based on the results of the sentiment analysis, the server uses the LINE MUSIC API to select appropriate songs. The selected song data is recorded in a database. The input is the sentiment analysis results, and the output is the selected song data.
[0523] Step 10:
[0524] Server: The POP images and selected music are integrated using video editing software such as Adobe Premiere API to generate a short video. The generated video data is stored in Google Cloud Storage. The input is the POP images and music data, and the output is the generated short video.
[0525] Step 11:
[0526] User: Press the "Post" button in the app and select the generated short video. Select the social media platform to post to (e.g. Instagram, Twitter) and press the "Send" button. The input is the short video to be uploaded, and the output is the video posted to the social media platform.
[0527] Step 12:
[0528] Terminal: Calls the API of the selected SNS and uploads the video. When the upload is complete, a completion notification is displayed to the user. The input is the short video data, and the output is the video uploaded to the SNS.
[0529] (Application example 1)
[0530] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0531] To effectively share their impressions and reviews of a book after reading it, users need a simple and efficient process for creating a video with appropriate visual content and music, and then sharing it on social networking services. However, current technology requires these steps to be performed individually, which is tedious and can lack accuracy and consistency. An integrated system is needed to solve this problem.
[0532] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0533] In this invention, the server includes an image analysis means, a bibliographic information acquisition means, a sentiment analysis means, an electronic design generation means, a music selection means, a short video generation means, a short video distribution means, a means for taking an image from a user terminal and inputting the data, and a means for sharing and distributing the generated short video to a social networking service. This makes it possible to automatically generate a short video containing visual content that combines bibliographic information and sentiment analysis results and appropriate music based on the user's impressions of a book they have finished reading, and to easily share the generated short video on a social networking service.
[0534] "Image analysis means" refers to technology for extracting information from images taken by a user.
[0535] "Bibliographic information acquisition means" refers to technology for acquiring accurate bibliographic information about a book based on the extracted information.
[0536] "Emotion analysis means" refers to technology that analyzes the impressions and keywords entered by users and understands their emotions.
[0537] "Electronic design generation means" refers to technology for automatically generating POP text and graphics based on sentiment analysis results and bibliographic information.
[0538] "Music selection means" refers to technology for selecting appropriate music based on the results of emotion analysis.
[0539] "Short video generation means" refers to the technology for generating short videos by combining POP images with selected music.
[0540] "Short video distribution means" refers to technology for distributing the generated short videos to social networking services, etc.
[0541] "Means for taking pictures and inputting data from a user device" refers to technology that allows users to take pictures of book covers using devices such as smartphones and input that data into the system.
[0542] "Means for sharing and distributing the generated short video on a social networking service" refers to technology that enables the generated short video to be easily posted and shared on a social networking service.
[0543] A system for implementing this invention utilizes book information and impressions based on a user's reading experience to automatically generate short videos containing visual content and music, and share them on social networking services.
[0544] First, the user takes a photo of the cover of the book they have just finished reading using their smartphone. The user device takes the image and uploads the data to a server via an application. This is where the user device, including the smartphone, comes into play.
[0545] The server then receives the uploaded image and performs optical character recognition (OCR) using image analysis techniques to extract the book title, author name, and visual elements from the image. The OCR engine used is Tesseract or similar.
[0546] The server then uses the extracted title and author name to search the book database to find the correct bibliographic information, or if the bibliographic information is not found, prompt the user to enter it manually.
[0547] The user then enters their thoughts about the book and keywords into the input form of the terminal application. The server receives the thoughts and performs sentiment analysis using a sentiment analysis engine such as the Google Cloud Natural Language API.
[0548] The server generates POP text based on the results of sentiment analysis and bibliographic information. Next, the electronic design generation means operates and automatically generates POP images based on the bibliographic information, sentiment analysis results, and cover image. An automatic design engine is used.
[0549] Furthermore, the results of the sentiment analysis are used to select appropriate songs from a music database. For example, songs that evoke adventures are selected based on impressions that evoke adventures. A suitable music database can be found on a general music distribution service.
[0550] Finally, the server combines the POP image with the selected music to generate a short video. This method of generating short videos may utilize a video editing library. The generated short video is sent from the server to the user's device, and the user can easily share it on social networking services by pressing the "post" button on their device. By utilizing the API of the SNS platform, users can seamlessly share content.
[0551] Specific examples
[0552] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the author's name from the image. It then compares the image with a bibliographic database to obtain accurate bibliographic information. If the user inputs their impression that the book is "full of adventure and emotion," the sentiment analysis algorithm analyzes this and generates a pop-up message with an "adventure" theme. A song that evokes the image of adventure is selected based on the keyword "adventure." Finally, a short video is generated by integrating the cover image, pop-up message, and selected song, which users can post on social media to share the appeal of the book with many people.
[0553] Example prompt sentence:
[0554] "You can take a photo of the cover of a fantasy novel and input your impression that it's 'a work filled with adventure and emotion.' The system will then use OCR to extract information about the book, perform a sentiment analysis, select an adventure-themed POP design and music, and create a short video to post on social media."
[0555] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0556] Step 1:
[0557] Users take a photo of the cover of a book they have finished reading using a device such as a smartphone, and the image is uploaded to a server via a device application.
[0558] Input: An image file taken on the user's device
[0559] Output: Image data uploaded to the server
[0560] Step 2:
[0561] The server receives the uploaded image and performs optical character recognition (OCR) using image analysis techniques. The OCR engine extracts the book title, author name, and visual elements from the image.
[0562] Input: Uploaded image data
[0563] Output: Text data of book title, author name, and visual elements
[0564] Specific operation: The OCR engine running on the server extracts text from the image and generates analysis results.
[0565] Step 3:
[0566] The server then searches the bibliographic database based on the extracted title and author name to retrieve the correct bibliographic information, and if no bibliographic information is found, generates feedback to the user requesting manual input.
[0567] Input: Text data of book title and author name
[0568] Output: Bibliographic information (e.g., publisher name, publication date, ISBN, etc.)
[0569] Specific operation: The server communicates with the bibliographic database to search and retrieve the corresponding bibliographic information.
[0570] Step 4:
[0571] Users enter their thoughts and keywords about the book they have just read into the input form in the terminal application.
[0572] Input: User-entered comments and keywords
[0573] Output: Text data of impressions and keywords
[0574] Specific operation: The user enters their thoughts and keywords in text format into the input form on the device.
[0575] Step 5:
[0576] The server receives the inputted impressions and performs sentiment analysis using the sentiment analysis engine, which extracts the main essence and classifies the emotions into numerical values and categories.
[0577] Input: Text data of impressions and keywords
[0578] Output: Sentiment analysis results (e.g., positive, negative, main essence)
[0579] What it does: The sentiment analysis engine analyzes the sentiment text and extracts sentiment and key topics.
[0580] Step 6:
[0581] The server generates POP text based on the results of sentiment analysis and bibliographic information. The generated POP text will highlight the book's appeal.
[0582] Input: Sentiment analysis results, bibliographic information
[0583] Output: POP text
[0584] Specific operation: The server uses templates and generation algorithms to automatically generate text that combines sentiment analysis results and bibliographic information.
[0585] Step 7:
[0586] Using an electronic design generation tool, a POP image is generated based on bibliographic information, sentiment analysis results, and cover image.
[0587] Input: Bibliographic information, sentiment analysis results, cover image
[0588] Output: POP image
[0589] How it works: The online design tool automatically generates visually appealing POP images based on bibliographic information and sentiment analysis results.
[0590] Step 8:
[0591] The results of the sentiment analysis are used to select appropriate songs from a music database.
[0592] Input: Sentiment analysis results
[0593] Output: Selected music files
[0594] Specific operation: The server searches the music database and selects music files that match the emotion analysis results.
[0595] Step 9:
[0596] The server combines the POP image with the selected music to generate a short video.
[0597] Input: POP image, selected music file
[0598] Output: Short video file
[0599] What it does: The video editing library merges the POP image with the selected music to generate a short video in a specific format.
[0600] Step 10:
[0601] The user's device receives the short video data sent from the server and converts it into a format that can be posted to SNS. When the user presses the "Post" button, the generated short video is shared and distributed to the selected social networking service.
[0602] Input: Short video file
[0603] Output: Content posted to social media
[0604] Specific operation: The terminal application converts the short video and calls the SNS API to post it.
[0605] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0606] This invention is a system that photographs the cover image of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, recognizes the user's emotions using an emotion engine, and based on this, compares it with a bibliographic database to obtain accurate bibliographic information.It then uses emotion analysis to analyze the impressions and keywords entered by the user, automatically generates a designed POP based on the bibliographic information and the results of the emotion analysis, selects an appropriate song, and combines the generated design with the selected song to generate a short video, which is then distributed to social media.
[0607] Program processing
[0608] 1. Taking and uploading images
[0609] Users take a photo of the cover of a book they have finished reading with their smartphone, and the image is saved in the device app.
[0610] The device compresses the captured image and uploads it to the server, along with the user ID and session information.
[0611] 2. OCR processing
[0612] The server inputs the received image into an OCR engine, which extracts the book title, author name, and visual elements from the image.
[0613] 3. Obtaining bibliographic information
[0614] The server compares the OCR-extracted book title and author name with a book database to obtain accurate bibliographic information. If no bibliographic information is found, it presents alternative title suggestions and asks the user for confirmation.
[0615] 4. Emotion recognition
[0616] Users enter their impressions of the book they have read and keywords into the input form of the terminal app, and the input is automatically sent to the server.
[0617] The server uses an emotion engine to analyze the impressions and keywords entered by the user and recognize emotions. The emotion engine extracts emotions using natural language processing technology. Furthermore, if the user's facial image or voice is used, emotions can also be recognized from this data.
[0618] 5. Emotion analysis
[0619] The server analyzes the user's impressions, including the emotion recognition results, and extracts the main essence using an emotion analysis algorithm.
[0620] 6. POP Text Generation
[0621] The server generates POP text based on the results of sentiment analysis and bibliographic information, and the generated text is structured to reflect the emotional emphasis.
[0622] 7. Design Generation
[0623] The server runs an automated design engine that generates a POP image based on bibliographic information, sentiment analysis results, and the cover image, incorporating key design elements from the cover.
[0624] 8. Music Selection
[0625] The server then refers to the results of the emotion analysis and selects appropriate songs from a music database, which are chosen to match the emotional tone.
[0626] 9. Short video generation
[0627] The server combines the POP image with the selected music and generates a short video using a video generation tool such as FFmpeg, which is then sent to the device.
[0628] 10. Social Media Distribution
[0629] The device receives the short video sent from the server and displays a confirmation screen to the user, who then confirms the video and it is ready to post.
[0630] The user presses the "Post" button on the device app and selects the social media platform (e.g., Instagram, Twitter) to post the generated short video to.
[0631] The device will call the upload API of the selected social networking site to upload the short video, and the user will receive a notification when the post is complete.
[0632] Specific examples
[0633] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the "author's name" from the image. It then compares the image with a bibliographic database to obtain accurate bibliographic information. If a user writes their impression of a book as "full of adventure and emotion," the emotion engine analyzes it and uses a sentiment analysis algorithm to extract the key essences of "adventure" and "emotion." From this result, a POP text is generated, such as "an inspiring story with adventurous elements."
[0634] The automated design engine then generates a POP image based on the text, the acquired bibliographic information, and the cover image. Furthermore, it selects an appropriate song from a music database based on the results of sentiment analysis. Finally, a short video is generated by integrating the POP image and the selected song, which users can easily post to social media.
[0635] In this way, the system of the present invention allows users to easily create POPs like those found in bookstores at home and share them on social media.By incorporating an emotion engine, it is possible to provide more personalized POPs that reflect the user's emotions.
[0636] The processing flow will be explained below.
[0637] Step 1:
[0638] Users take a photo of the cover of a book they have finished reading with their smartphone, and the image is saved in the device's app storage.
[0639] Step 2:
[0640] The device compresses the captured image and uploads it to the server, which includes the user ID and session information.
[0641] Step 3:
[0642] The server inputs the received image into an OCR engine, which extracts the book title, author name, and visual elements from the image.
[0643] Step 4:
[0644] The server compares the book title and author name extracted by OCR with a book database to obtain accurate bibliographic information. If the bibliographic information is not found as a result of the comparison with the book database, it presents several alternative title candidates and asks the user for confirmation.
[0645] Step 5:
[0646] The user enters their thoughts about the book they have read and keywords into the input form of the terminal app. The input information is converted into JSON format and sent to the server.
[0647] Step 6:
[0648] The server uses an emotion engine to analyze the impressions and keywords entered by the user and recognize emotions. The emotion engine uses natural language processing technology to extract emotions from text. If the user also provides facial images or voice data, the engine can also recognize emotions from that data.
[0649] Step 7:
[0650] The server generates POP text based on the recognized sentiment and bibliographic information. The generated text reflects the essence obtained from the sentiment analysis results.
[0651] Step 8:
[0652] The server runs an automatic design engine to generate POP images based on bibliographic information, sentiment analysis results, and cover images. The automatic design engine creates POP images using design templates.
[0653] Step 9:
[0654] The server then refers to the emotion analysis results and selects appropriate songs from a music database that match the emotional tone.
[0655] Step 10:
[0656] The server combines the POP image with the selected music and generates a short video using a video generation tool such as FFmpeg, which is then sent to the device.
[0657] Step 11:
[0658] The device receives the short video sent from the server and displays a confirmation screen to the user, who can then confirm the video and indicate that it is ready to be posted.
[0659] Step 12:
[0660] The user presses the "Post" button on the device app and selects the social media platform (e.g., Instagram, Twitter) to post the generated short video to.
[0661] Step 13:
[0662] The device will call the upload API of the selected social networking service to upload the short video, and a notification will be displayed to the user once the post is successful.
[0663] This completes the entire process, allowing users to easily generate a POP for the book and share it on social media.
[0664] Example 2
[0665] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0666] While much information is shared on social media these days, there are limited ways for users to easily and effectively share their reading experiences. There is a particular need for sharing book reviews and ratings in a visually appealing format, but existing technologies require manual input and processing of information, which is time-consuming. Furthermore, it is difficult to generate personalized content that reflects user sentiment using sentiment analysis technology.
[0667] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an image analysis means, a bibliographic information acquisition means, an emotion analysis means, an electronic design generation means, a music selection means, a short video generation means, a short video distribution means, an image compression and upload means, an OCR analysis means, a user interaction means, a voice and face image analysis means, a display confirmation means, and an SNS posting means. This enables users to easily and attractively visualize their impressions and reviews of books they have finished reading and share them on SNS.
[0668] "Image analysis means" refers to means having the function of analyzing an image and extracting information.
[0669] The "bibliographic information acquisition means" is a means having a function for acquiring accurate book information related to a book category.
[0670] An "emotion analysis means" is a means that has the function of analyzing and recognizing emotions from impressions and keywords entered by the user.
[0671] "Electronic design generation means" means a means having a function for generating a design electronically.
[0672] The "music selection means" is a means having a function for selecting appropriate music based on the emotion analysis results.
[0673] The "short video generation means" is a means having a function for generating a short video by integrating a POP image with a selected piece of music.
[0674] "Short video distribution means" refers to a means that has the function of distributing the generated short video to social media and other platforms.
[0675] The "image compression and uploading means" is a means having a function for compressing a captured image and uploading it to a server.
[0676] "OCR Analysis Means" means a means capable of extracting book titles, author names, and visual elements from an image using optical character recognition technology.
[0677] "User interaction means" refers to means that has the function of providing an interface for users to input their thoughts and keywords into the system.
[0678] The "voice and facial image analysis means" is a means having a function for analyzing emotions from the voice and facial image data provided by the user.
[0679] The "display confirmation means" is a means having a function of allowing the user to confirm the generated content.
[0680] "SNS posting means" refers to a means that has the function of allowing users to easily post content they have viewed to SNS.
[0681] This invention is a system that takes a photo of the cover of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, recognizes the user's emotions using an emotion engine, and obtains accurate bibliographic information by matching it with a book database. Furthermore, it analyzes the impressions and keywords entered by the user using emotion analysis, automatically generates a POP design based on the bibliographic information and the emotion analysis results, selects an appropriate song, and generates a short video by integrating the generated design with the selected song, which is then distributed to social media.
[0682] The system is programmed as follows:
[0683] A user uses their smartphone camera to take a photo of the cover of a book they have just read. The captured image is saved in the device app. The device compresses the image and uploads it to the server, along with the user ID and session information. The server then inputs the uploaded image into an OCR engine (e.g., Google Cloud Vision API) to extract the book title, author name, and visual elements from the image.
[0684] The server compares the book title and author name extracted by OCR with a book database (e.g., Google Books API, Open Library) to obtain accurate bibliographic information. If bibliographic information is not found, alternative title suggestions are presented and the user is asked for confirmation. The user then enters their thoughts and keywords about the book they read into the input form of the device app. The input is automatically sent to the server. The server uses an emotion engine (e.g., IBM Watson NLU) to analyze the thoughts and keywords entered by the user and recognize emotions. If the user provides facial images or voice, emotions can also be recognized from this data, if necessary.
[0685] The server generates POP text based on the results of sentiment analysis and bibliographic information. The generated text is structured to reflect the emotional emphasis. The server then runs an automated design engine (e.g., Adobe Creative Cloud API) to generate a POP image based on the bibliographic information, sentiment analysis results, and cover image. The POP image also incorporates key design elements from the cover.
[0686] The server then refers to the emotion analysis results and selects an appropriate song from a music database (e.g., Spotify API, Apple Music API). The song is selected to match the emotional tone. The server then combines the POP image with the selected song and generates a short video using a video generation tool such as FFmpeg. The generated video data is then sent to the device.
[0687] The user can view the short video sent from the server on their device. After viewing, the user presses the "Post" button on the device app to post the video to a social networking platform (e.g., Instagram, Twitter). The device then calls the upload API of the selected social networking platform and uploads the short video. Once posting is complete, a notification is displayed to the user.
[0688] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the author's name from the image. It then compares the image with a book database to obtain accurate bibliographic information. If the user writes their impression, "This is a work filled with adventure and emotion," the emotion engine analyzes it and extracts the key essences of "adventure" and "emotion" using a sentiment analysis algorithm. From these results, the server creates a POP text such as "An inspiring story with adventurous elements."
[0689] The automated design engine then generates a POP image based on the text, the acquired bibliographic information, and the cover image. Furthermore, it selects an appropriate song from a music database based on the results of sentiment analysis. Finally, a short video is generated by integrating the POP image and the selected song, which users can easily post to social media.
[0690] The following are examples of prompt sentences:
[0691] "Please explain in natural language the process of a system that takes a photo of a book cover, extracts the book title and author name using OCR technology, analyzes the emotions felt after reading using an emotion engine, compares it with a book database to obtain bibliographic information, generates POP text and images based on the emotion analysis results, selects music, creates a video using a short video creation tool (e.g., FFmpeg), and posts it to social media."
[0692] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0693] Step 1: Capture and upload images
[0694] The user takes a photo of the book cover using the smartphone camera. The captured image is automatically saved in the device app. The device compresses the saved image and uploads it to the server. When uploading, the user ID and session information are also sent. This sends the compressed image data and user information to the server.
[0695] Step 2: OCR
[0696] The server inputs the uploaded image into an OCR engine (e.g., optical character recognition software). The OCR engine extracts the book title, author name, and visual elements from the image. It receives image data as input and obtains text data of the book title, author name, and visual elements as output. Specifically, the OCR engine identifies character regions in the image and reads character data from those regions.
[0697] Step 3: Obtain bibliographic information
[0698] The server uses the book title and author name extracted by OCR to check against a book database (e.g., a book information service). This allows accurate bibliographic information to be obtained. It receives text data of the book title and author name as input, and obtains detailed book information (e.g., publication year, genre, summary) as output. Specifically, it sends a request to the book database via an API and obtains the corresponding book information.
[0699] Step 4: Emotion Recognition
[0700] The user enters their thoughts and keywords about the book they have read into an input form in the device app. The input is automatically sent to the server. The server uses an emotion engine (e.g., natural language processing software) to analyze the thoughts and keywords entered by the user and recognize the emotion. It receives the text data of the thoughts and keywords as input and obtains the type of emotion (e.g., joy, sadness) as output. Specifically, the emotion engine analyzes the text and runs an algorithm to classify the emotion.
[0701] Step 5: Sentiment Analysis
[0702] The server analyzes the emotion results recognized by the emotion engine and the user's impressions, and extracts the main essence using an emotion analysis algorithm. It receives emotion data and impression text as input, and obtains the extracted essence (e.g., adventure, emotion) as output. Specifically, it uses a text analysis algorithm to extract important keywords and themes from the impressions.
[0703] Step 6: POP Text Generation
[0704] The server generates POP text based on the results of sentiment analysis and bibliographic information. It receives the sentiment essence and bibliographic information as input and obtains a catchy slogan and description as output. Specifically, it uses a template engine to compose the text and adjust it to reflect the emotional emphasis.
[0705] Step 7: Design Generation
[0706] The server runs an automatic design engine (e.g., design software API) to generate POP images based on bibliographic information, sentiment analysis results, and cover images. It receives design elements (e.g., book title, author name, emotional essence) as input and obtains POP image data as output. Specifically, it automatically adjusts the layout and coloring to generate visually appealing images.
[0707] Step 8: Music Selection
[0708] The server refers to the emotion analysis results and selects appropriate songs from a music database (e.g., a music streaming service). It receives the emotion essence as input and obtains song information as output. Specifically, it searches for songs that match the emotion and executes an algorithm to select the most suitable one.
[0709] Step 9: Short video generation
[0710] The server combines the POP image with the selected music and generates a short video using a video generation tool (e.g., FFmpeg). It receives POP image data and music information as input and obtains a short video file as output. Specifically, it adjusts the timing of the image and music and encodes them as continuous visual content.
[0711] Step 10: Social Media Distribution
[0712] The device receives the short video sent from the server and displays a confirmation screen to the user. The user checks the video and presses the "Post" button to post the video to the SNS platform (e.g., SNS service API). The device receives the short video file as input and receives a notification that posting to the SNS has been completed as output. Specifically, the device calls the SNS upload API and displays a notification to the user when posting is successful.
[0713] (Application example 2)
[0714] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0715] Traditionally, creating recommended book POPs in bookstores required a lot of time and effort, and the content of the POPs was not emotionally personalized, limiting their appeal to customers. Furthermore, updating promotional materials displayed on in-store displays was done manually, resulting in a lack of immediacy. For these reasons, a method was needed for bookstore staff to quickly and effectively promote books.
[0716] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an image analysis means, a bibliographic information acquisition means, a sentiment analysis means, an electronic design generation means, a music selection means, a short video generation means, a short video distribution means, and an automatic display means. This allows bookstore staff to simply take a photo of a book cover with their smartphone, and automatically generate an individual recommended POP based on related bibliographic information, user reviews, and sentiment analysis, and display it in real time on an electronic display in the store or post it to social media.
[0717] "Image analysis means" refers to means for extracting information from captured images.
[0718] "Bibliographic information acquisition means" refers to a means for acquiring information related to a book (e.g., book title, author name) from a database.
[0719] "Sentiment analysis means" means means for analyzing emotions from user input or other data.
[0720] "Electronic design generation means" refers to a means for automatically generating designs based on book information and sentiment analysis results.
[0721] The "music selection means" is a means for selecting appropriate music based on the emotion analysis results.
[0722] The "short video generation means" is a means for integrating a POP image with selected music to generate a short video.
[0723] "Short video distribution means" refers to a means for distributing the generated short video to an SNS platform.
[0724] "Automatic display means" refers to a means for displaying the generated POP images and short videos on electronic displays in the store in real time.
[0725] This system allows bookstore staff to take a photo of a book cover with their smartphone, automatically generating recommended POPs based on related bibliographic information, user reviews, and sentiment analysis, and posting them on in-store electronic displays and social media. The details of this system are described below.
[0726] The system starts by having the user take a photo of the book cover with their smartphone. The image is compressed and uploaded to a server, where it is analyzed using image analysis tools (e.g., pytesseract) to extract the book title, author, and visual elements.
[0727] Next, the server uses the bibliographic information acquisition means to acquire accurate bibliographic information from the book database, including the book title, author name, publication year, genre, etc.
[0728] The server then uses sentiment analysis tools to analyze the reviews and keywords entered by the user and recognize emotions. This process includes an emotion engine using natural language processing techniques. Facial images and voice data may also be used for emotion recognition.
[0729] Based on the results of the sentiment analysis and the bibliographic information, the server automatically generates POP text using an electronic design generation tool, and then creates a POP image using a design engine. This POP image reflects the bibliographic information and the user's sentiment.
[0730] Furthermore, the server uses a music selection means to select appropriate music from a music database based on the result of the emotion analysis, and the music is selected to match the emotional tone.
[0731] The server combines these POP images with the selected music and generates a short video using a short video generation tool (e.g., FFmpeg). This video is automatically generated and delivered to the user's smartphone.
[0732] Finally, the generated short video is posted to a social networking site using a short video distribution method. Users can select a social networking site (e.g., Instagram or Twitter) on their smartphone to complete the posting.
[0733] To accommodate in-store promotions, the system is equipped with an automatic display means, which displays the generated POP images on electronic displays in the store in real time.
[0734] Hardware and Software Use
[0735] The main hardware used in this invention is a smartphone and an electronic display in a bookstore, and the main software is an image analysis engine (pytesseract), a sentiment analysis engine, a design generation API, a video generation tool (FFmpeg), and a SNS upload API.
[0736] Specific examples
[0737] For example, if a user takes a photo of the cover of a "fantasy novel" and writes in their review that it is "a work filled with adventure and emotion," the system will analyze this information and generate a pop-up image that emphasizes "adventure" and "emotion." This image is then combined with a selected song to create a short video. The video can then be displayed on in-store displays and posted to social media.
[0738] Prompt Sentence Examples
[0739] "Users take a photo of the cover of a book they have read and enter their thoughts on the book. The system analyzes the information and generates a short video that combines recommended pop music and posts it to social media."
[0740] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0741] Step 1:
[0742] Image capture and upload
[0743] The user takes a photo of the book cover with their smartphone. This image is saved in the smartphone app. The device compresses the image file and uploads it to the server. The user ID and session information are also sent at the time of uploading.
[0744] Step 2:
[0745] OCR processing
[0746] The server inputs the received image into an OCR engine (pytesseract), which extracts text from the image and obtains information such as the book title and author name. The extracted text data is sent to the next processing step.
[0747] Step 3:
[0748] Bibliographic information acquisition
[0749] The server uses the book title and author name extracted by OCR to retrieve detailed bibliographic information from a book database. Using the book title and author name as input data, the server compares the book title, author name, publication year, genre, and other bibliographic information to output.
[0750] Step 4:
[0751] emotion recognition
[0752] The user enters their thoughts and keywords about the book they have read into an input form on a smartphone app. This text data is sent to a server. The server uses an emotion analysis tool (emotion engine) to analyze the thoughts and keywords entered by the user and recognize the emotion. The emotion analysis engine takes this data as input and outputs the type of emotion (e.g., joy, sadness).
[0753] Step 5:
[0754] Emotion analysis
[0755] The server analyzes the sentiment data, including the emotion recognition results, and uses a sentiment analysis algorithm to extract the key essence and generate the data needed to generate POPs. The input is the emotion recognition results and user sentiment data, and the output is the essence extraction results.
[0756] Step 6:
[0757] POP Text Generation
[0758] The server automatically generates POP text based on the results of sentiment analysis and bibliographic information. The generated text reflects the emotional emphasis. The input is the sentiment analysis results and bibliographic information, and the output is POP text.
[0759] Step 7:
[0760] Design Generation
[0761] The server runs an electronic design generator to automatically generate a POP image based on the sentiment analysis results, bibliographic information, and cover image. This POP image also incorporates the main design elements of the cover. The input is the POP text, bibliographic information, and cover image, and the output is the POP image.
[0762] Step 8:
[0763] Music Selection
[0764] The server refers to the results of the emotion analysis and selects an appropriate song from a music database. The selected song matches the emotional tone. The input is the emotion analysis result, and the output is the selected song.
[0765] Step 9:
[0766] Short video generation
[0767] The server combines the POP image and the selected music and generates a short video using a short video generation tool (FFmpeg). The input is the POP image and music URL, and the output is the short video data.
[0768] Step 10:
[0769] SNS distribution
[0770] The device receives the short video sent from the server and displays a confirmation screen to the user. The user reviews the video and is ready to post. When the user selects an SNS platform (e.g., Instagram, Twitter) and presses the "Post" button, the device calls the upload API of the selected SNS and uploads the short video. The input is the short video data and the SNS information selected by the user, and the output is a notification of successful posting to the SNS.
[0771] Step 11:
[0772] Automatic display
[0773] The server uses an automatic display means to display the generated POP image on an electronic display in the store in real time. The input is the POP image, and the output is the display on the in-store display.
[0774] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0775] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0776] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0777] [Third embodiment]
[0778] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0779] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0780] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0781] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0782] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0783] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0784] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0785] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0786] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0787] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0788] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0789] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0790] This invention is a system that takes a photo of the cover of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, compares this with a bibliographic database to obtain accurate bibliographic information, and then uses sentiment analysis to analyze the impressions and keywords entered by the user.It then automatically generates a POP design based on the bibliographic information and the results of the sentiment analysis, selects an appropriate song, and combines the generated design with the selected song to generate a short video, which is then distributed to social media.
[0791] Program processing
[0792] 1. Taking and uploading images
[0793] User: Take a photo of the cover of a book you have just read using your smartphone or other device.
[0794] Device: Upload the captured image to the server via the device app.
[0795] 2. OCR processing
[0796] Server: Receives the uploaded image and inputs it into the OCR engine, which extracts the book title, author name, and visual elements from the image.
[0797] 3. Obtaining bibliographic information
[0798] Server: Based on the extracted book title and author name, it checks the book database to get the correct bibliographic information. If no bibliographic information is found, it generates feedback to the user requesting manual input.
[0799] 4. Enter your thoughts
[0800] User: Enter their thoughts and keywords about the book they read into the input form in the device app.
[0801] 5. Emotion analysis
[0802] Server: Analyzes the user-entered comments and extracts the main essence using a sentiment analysis algorithm.
[0803] 6. POP Text Generation
[0804] Server: Generates POP text based on the results of sentiment analysis and bibliographic information.
[0805] 7. Design Generation
[0806] Server: Runs the automatic design engine and generates POP images based on bibliographic information, sentiment analysis results, and cover images.
[0807] 8. Music Selection
[0808] Server: Refers to the results of the sentiment analysis and selects appropriate songs from music databases such as LINE MUSIC.
[0809] 9. Short video generation
[0810] Server: Integrates POP images with selected music to generate short videos.
[0811] Device: Receives short video data sent from the server and converts it into a format that can be posted to social media.
[0812] 10. Social Media Distribution
[0813] User: Press the "Post" button on the device app and select the generated short video.
[0814] Device: Calls an API to upload the video to the social networking site of the user's choice, and notifies the user when the post is complete.
[0815] Specific examples
[0816] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the author's name from the image. It then compares the image with a bibliographic database to obtain accurate bibliographic information. If the user inputs a sentiment such as "This is a work filled with adventure and emotion," the sentiment analysis algorithm analyzes this and generates POP text. Using the keyword "adventure," a song that evokes the image of adventure is selected. Finally, the cover image, POP text, and selected song are integrated to generate a short video, which users can post on social media to share the appeal of the book with many people.
[0817] In this way, the system of the present invention allows users to easily create POPs like those found in bookstores at home and share them on social media.
[0818] The processing flow will be explained below.
[0819] Step 1:
[0820] Users take a photo of the cover of a book they have finished reading with their smartphone, and the image is saved in the device app.
[0821] Step 2:
[0822] The device compresses the captured image and uploads it to the server, along with the user ID and session information.
[0823] Step 3:
[0824] The server inputs the received image into an OCR engine, which extracts the book title, author name, and visual elements from the image.
[0825] Step 4:
[0826] The server compares the OCR-extracted book title and author name with a book database to obtain accurate bibliographic information. If no bibliographic information is found, it presents alternative title suggestions and asks the user for confirmation.
[0827] Step 5:
[0828] Users enter their impressions of the book they have read and keywords into the input form of the terminal app, and the input is automatically sent to the server.
[0829] Step 6:
[0830] The server analyzes the user's input and uses a sentiment analysis algorithm to extract the main essence, thereby determining which sentiment prevails.
[0831] Step 7:
[0832] The server generates POP text based on the results of sentiment analysis and bibliographic information, and the generated text is structured to reflect the emotional emphasis.
[0833] Step 8:
[0834] The server runs an automated design engine that generates a POP image based on bibliographic information, sentiment analysis results, and the cover image, incorporating key design elements from the cover.
[0835] Step 9:
[0836] The server then refers to the results of the emotion analysis and selects appropriate songs from a music database, which are chosen to match the emotional tone.
[0837] Step 10:
[0838] The server combines the POP image with the selected music and generates a short video using a video generation tool such as FFmpeg, which is then sent to the device.
[0839] Step 11:
[0840] The device receives the short video sent from the server and displays a confirmation screen to the user, who then confirms the video and it is ready to post.
[0841] Step 12:
[0842] The user presses the "Post" button on the device app and selects the social media platform (e.g., Instagram, Twitter) to post the generated short video to.
[0843] Step 13:
[0844] The device will call the upload API of the selected social networking site to upload the short video, and the user will receive a notification when the post is complete.
[0845] This completes the entire process, allowing users to easily generate a POP for the book and share it on social media.
[0846] Example 1
[0847] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0848] Conventional systems lack a way for users to efficiently and accurately record information about books they have read and share that content in an engaging way on social media. This requires manual input and editing, which takes a lot of time and effort. Furthermore, analyzing impressions and keywords and selecting design and music based on those must also be done manually, resulting in an inconsistent user experience. Furthermore, distributing the generated content to social media can be difficult.
[0849] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0850] In this invention, the server includes image analysis means, bibliographic information acquisition means, emotion analysis means, electronic design generation means, music selection means, short video generation means, short video distribution means, a terminal for uploading images to the server via a user interface, a terminal for accepting user feedback, means for comparing extracted book titles and author names with a database to acquire bibliographic information, text generation means for generating POP text based on the emotion analysis results, and means for the short video generation means to integrate the images, generated text, and selected music. This allows users to easily and efficiently record information about books they have read, create attractive short videos that reflect the emotion analysis results, and share them on social media.
[0851] An "image analysis tool" is a device or software that uses machine learning algorithms or optical character recognition techniques to extract textual information and visual elements from an image.
[0852] The "bibliographic information acquisition means" refers to a device or software that compares the extracted book title and author name with a database to acquire accurate bibliographic information.
[0853] "Emotion analysis means" refers to a device or software that analyzes text data entered by a user and extracts emotions and key essences.
[0854] An "electronic design generator" is a device or software that automatically creates an attractive design based on POP text and visual elements.
[0855] A "music selection means" is a device or software that selects appropriate music based on the results of emotion analysis.
[0856] The "short video generation means" refers to a device or software that integrates the generated design and selected music to create a short video.
[0857] "Short video distribution means" refers to devices or software used to upload and distribute completed short videos to platforms such as social media.
[0858] A "terminal that uploads images to a server via a user interface" is a device that has hardware or software for sending images taken by a user to a server.
[0859] A "terminal that accepts user feedback input" is a device that provides an input form in which users can input their feedback or keywords, and that is equipped with hardware or software that accepts that data.
[0860] "Means for obtaining bibliographic information by comparing extracted book titles and author names with a database" refers to devices or software that executes the process of using book titles and author names extracted using OCR technology to compare them with an external database and obtain accurate bibliographic information.
[0861] "Text generation means for generating POP text based on the results of sentiment analysis" refers to a device or software that uses the results of sentiment analysis to automatically generate text that succinctly conveys the features and appeal of a book.
[0862] This system photographs the cover image of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, compares this with a bibliographic database to obtain accurate bibliographic information, and then analyzes the impressions and keywords entered by the user using sentiment analysis.The system then automatically generates a POP design based on the bibliographic information and the results of the sentiment analysis, selects an appropriate song, and combines the generated design with the selected song to create a short video, which is then distributed to social media.
[0863] Hardware and software used
[0864] Image analysis methods
[0865] The server stores the cover image uploaded by the user in Google Cloud Storage, then uses the Google Cloud Vision OCR API to extract the book title, author name, and cover image. The OCR engine takes image data as input and outputs text information.
[0866] Bibliographic information acquisition method
[0867] The server sends the extracted book title and author name to the Google Books API to get the exact bibliographic information, which is then returned from the database and stored in an internal database.
[0868] Emotion analysis means
[0869] The server inputs the user's feedback text into the IBM Watson Natural Language Understanding API to extract sentiment and key points, and the analysis results are also stored in a database.
[0870] Electronic Design Generator
[0871] The server uses the Canva API to generate POP images based on bibliographic information, sentiment analysis results, and cover images. The resulting designs are visually appealing and have a compelling layout.
[0872] Music selection method
[0873] Based on the results of the emotion analysis, the server accesses the LINE MUSIC API to select the appropriate song, and records the selected song data in a database.
[0874] Short video creation method
[0875] The server then combines the POP image and the selected music using video editing software such as Adobe Premiere API to create a short video, which is then stored in Google Cloud Storage.
[0876] Short video distribution methods
[0877] The device converts the short video data received from the server into a format that can be posted to SNS. When the user taps the "Post" button in the app, the device calls the SNS's API and uploads the video.
[0878] Examples of concrete examples and prompts
[0879] For example, a user takes a photo of the cover of a "fantasy novel" with their smartphone. The server sends this image to Google Cloud Vision OCR, which extracts "fantasy novel" and the author's name from the image. Next, it uses the Google Books API to obtain accurate bibliographic information. If the user enters their impression that the work is "full of adventure and emotion," the server performs sentiment analysis using the IBM Watson Natural Language Understanding API and extracts the key essence, "adventure." The server selects music related to "adventure" from the LINE MUSIC API and generates a pop-up image using the Canva API. Finally, the pop-up image and music are integrated using the Adobe Premiere API to generate a short video, which the user can post on social media.
[0880] An example of a prompt is as follows:
[0881] "Please build a system that takes a photo of the cover of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, and compares it with a bibliographic database. Analyzes the user's impressions using sentiment analysis, automatically generates POPs, and distributes short videos with appropriate music on social media."
[0882] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0883] Step 1:
[0884] User: Uses smartphone camera to take a picture of the cover of the book they have just read. Checks that the image is clear. The input is the image of the book cover, and the output is the image data.
[0885] Step 2:
[0886] Terminal: The captured image data is received within a dedicated app, compressed, and uploaded to the server. The input is the image data, and the output is the compressed image data sent to the server.
[0887] Step 3:
[0888] Server: Receives uploaded image data and stores it in Google Cloud Storage. Sends the stored image data to the Google Cloud Vision OCR API to extract the book title, author name, and cover image. The input is compressed image data, and the output is extracted text information.
[0889] Step 4:
[0890] Server: Sends the book title and author name obtained through OCR processing to the Google Books API to obtain bibliographic information. Stores the bibliographic information in an internal database and verifies the accuracy of the book title and author name. The input is the extracted text information, and the output is the bibliographic information stored in the database.
[0891] Step 5:
[0892] User: Follow the prompts from the system and enter your thoughts and keywords about the book you read into the input form within the app. Once you've finished entering your thoughts and keywords, tap the "Submit" button. The input is text data of your thoughts and keywords, and the output is text data sent to the server.
[0893] Step 6:
[0894] Server: Sends the sentiment text data to the IBM Watson Natural Language Understanding API, extracts sentiment and key essence, and stores the resulting analysis data in a database. The input is sentiment text data, and the output is the sentiment analysis results.
[0895] Step 7:
[0896] Server: Automatically generates POP text based on the sentiment analysis results and bibliographic information. The generated text succinctly summarizes the "features of this book" and includes visually appealing phrases. The input is the sentiment analysis results and bibliographic information, and the output is the generated POP text.
[0897] Step 8:
[0898] Server: Automatically generates digital designs using the Canva API. Generates POP images based on bibliographic information, sentiment analysis results, and cover images, creating visually appealing layouts. The inputs are bibliographic information, sentiment analysis results, and cover images, and the output is a POP image.
[0899] Step 9:
[0900] Server: Based on the results of the sentiment analysis, the server uses the LINE MUSIC API to select appropriate songs. The selected song data is recorded in a database. The input is the sentiment analysis results, and the output is the selected song data.
[0901] Step 10:
[0902] Server: The POP images and selected music are integrated using video editing software such as Adobe Premiere API to generate a short video. The generated video data is stored in Google Cloud Storage. The input is the POP images and music data, and the output is the generated short video.
[0903] Step 11:
[0904] User: Press the "Post" button in the app and select the generated short video. Select the social media platform to post to (e.g. Instagram, Twitter) and press the "Send" button. The input is the short video to be uploaded, and the output is the video posted to the social media platform.
[0905] Step 12:
[0906] Terminal: Calls the API of the selected SNS and uploads the video. When the upload is complete, a completion notification is displayed to the user. The input is the short video data, and the output is the video uploaded to the SNS.
[0907] (Application example 1)
[0908] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0909] To effectively share their impressions and reviews of a book after reading it, users need a simple and efficient process for creating a video with appropriate visual content and music, and then sharing it on social networking services. However, current technology requires these steps to be performed individually, which is tedious and can lack accuracy and consistency. An integrated system is needed to solve this problem.
[0910] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0911] In this invention, the server includes an image analysis means, a bibliographic information acquisition means, a sentiment analysis means, an electronic design generation means, a music selection means, a short video generation means, a short video distribution means, a means for taking an image from a user terminal and inputting the data, and a means for sharing and distributing the generated short video to a social networking service. This makes it possible to automatically generate a short video containing visual content that combines bibliographic information and sentiment analysis results and appropriate music based on the user's impressions of a book they have finished reading, and to easily share the generated short video on a social networking service.
[0912] "Image analysis means" refers to technology for extracting information from images taken by a user.
[0913] "Bibliographic information acquisition means" refers to technology for acquiring accurate bibliographic information about a book based on the extracted information.
[0914] "Emotion analysis means" refers to technology that analyzes the impressions and keywords entered by users and understands their emotions.
[0915] "Electronic design generation means" refers to technology for automatically generating POP text and graphics based on sentiment analysis results and bibliographic information.
[0916] "Music selection means" refers to technology for selecting appropriate music based on the results of emotion analysis.
[0917] "Short video generation means" refers to the technology for generating short videos by combining POP images with selected music.
[0918] "Short video distribution means" refers to technology for distributing the generated short videos to social networking services, etc.
[0919] "Means for taking pictures and inputting data from a user device" refers to technology that allows users to take pictures of book covers using devices such as smartphones and input that data into the system.
[0920] "Means for sharing and distributing the generated short video on a social networking service" refers to technology that enables the generated short video to be easily posted and shared on a social networking service.
[0921] A system for implementing this invention utilizes book information and impressions based on a user's reading experience to automatically generate short videos containing visual content and music, and share them on social networking services.
[0922] First, the user takes a photo of the cover of the book they have just finished reading using their smartphone. The user device takes the image and uploads the data to a server via an application. This is where the user device, including the smartphone, comes into play.
[0923] The server then receives the uploaded image and performs optical character recognition (OCR) using image analysis techniques to extract the book title, author name, and visual elements from the image. The OCR engine used is Tesseract or similar.
[0924] The server then uses the extracted title and author name to search the book database to find the correct bibliographic information, or if the bibliographic information is not found, prompt the user to enter it manually.
[0925] The user then enters their thoughts about the book and keywords into the input form of the terminal application. The server receives the thoughts and performs sentiment analysis using a sentiment analysis engine such as the Google Cloud Natural Language API.
[0926] The server generates POP text based on the results of sentiment analysis and bibliographic information. Next, the electronic design generation means operates and automatically generates POP images based on the bibliographic information, sentiment analysis results, and cover image. An automatic design engine is used.
[0927] Furthermore, the results of the sentiment analysis are used to select appropriate songs from a music database. For example, songs that evoke adventures are selected based on impressions that evoke adventures. A suitable music database can be found on a general music distribution service.
[0928] Finally, the server combines the POP image with the selected music to generate a short video. This method of generating short videos may utilize a video editing library. The generated short video is sent from the server to the user's device, and the user can easily share it on social networking services by pressing the "post" button on their device. By utilizing the API of the SNS platform, users can seamlessly share content.
[0929] Specific examples
[0930] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the author's name from the image. It then compares the image with a bibliographic database to obtain accurate bibliographic information. If the user inputs their impression that the book is "full of adventure and emotion," the sentiment analysis algorithm analyzes this and generates a pop-up message with an "adventure" theme. A song that evokes the image of adventure is selected based on the keyword "adventure." Finally, a short video is generated by integrating the cover image, pop-up message, and selected song, which users can post on social media to share the appeal of the book with many people.
[0931] Example prompt sentence:
[0932] "You can take a photo of the cover of a fantasy novel and input your impression that it's 'a work filled with adventure and emotion.' The system will then use OCR to extract information about the book, perform a sentiment analysis, select an adventure-themed POP design and music, and create a short video to post on social media."
[0933] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0934] Step 1:
[0935] Users take a photo of the cover of a book they have finished reading using a device such as a smartphone, and the image is uploaded to a server via a device application.
[0936] Input: An image file taken on the user's device
[0937] Output: Image data uploaded to the server
[0938] Step 2:
[0939] The server receives the uploaded image and performs optical character recognition (OCR) using image analysis techniques. The OCR engine extracts the book title, author name, and visual elements from the image.
[0940] Input: Uploaded image data
[0941] Output: Text data of book title, author name, and visual elements
[0942] Specific operation: The OCR engine running on the server extracts text from the image and generates analysis results.
[0943] Step 3:
[0944] The server then searches the bibliographic database based on the extracted title and author name to retrieve the correct bibliographic information, and if no bibliographic information is found, generates feedback to the user requesting manual input.
[0945] Input: Text data of book title and author name
[0946] Output: Bibliographic information (e.g., publisher name, publication date, ISBN, etc.)
[0947] Specific operation: The server communicates with the bibliographic database to search and retrieve the corresponding bibliographic information.
[0948] Step 4:
[0949] Users enter their thoughts and keywords about the book they have just read into the input form in the terminal application.
[0950] Input: User-entered comments and keywords
[0951] Output: Text data of impressions and keywords
[0952] Specific operation: The user enters their thoughts and keywords in text format into the input form on the device.
[0953] Step 5:
[0954] The server receives the inputted impressions and performs sentiment analysis using the sentiment analysis engine, which extracts the main essence and classifies the emotions into numerical values and categories.
[0955] Input: Text data of impressions and keywords
[0956] Output: Sentiment analysis results (e.g., positive, negative, main essence)
[0957] What it does: The sentiment analysis engine analyzes the sentiment text and extracts sentiment and key topics.
[0958] Step 6:
[0959] The server generates POP text based on the results of sentiment analysis and bibliographic information. The generated POP text will highlight the book's appeal.
[0960] Input: Sentiment analysis results, bibliographic information
[0961] Output: POP text
[0962] Specific operation: The server uses templates and generation algorithms to automatically generate text that combines sentiment analysis results and bibliographic information.
[0963] Step 7:
[0964] Using an electronic design generation tool, a POP image is generated based on bibliographic information, sentiment analysis results, and cover image.
[0965] Input: Bibliographic information, sentiment analysis results, cover image
[0966] Output: POP image
[0967] How it works: The online design tool automatically generates visually appealing POP images based on bibliographic information and sentiment analysis results.
[0968] Step 8:
[0969] The results of the sentiment analysis are used to select appropriate songs from a music database.
[0970] Input: Sentiment analysis results
[0971] Output: Selected music files
[0972] Specific operation: The server searches the music database and selects music files that match the emotion analysis results.
[0973] Step 9:
[0974] The server combines the POP image with the selected music to generate a short video.
[0975] Input: POP image, selected music file
[0976] Output: Short video file
[0977] What it does: The video editing library merges the POP image with the selected music to generate a short video in a specific format.
[0978] Step 10:
[0979] The user's device receives the short video data sent from the server and converts it into a format that can be posted to SNS. When the user presses the "Post" button, the generated short video is shared and distributed to the selected social networking service.
[0980] Input: Short video file
[0981] Output: Content posted to social media
[0982] Specific operation: The terminal application converts the short video and calls the SNS API to post it.
[0983] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0984] This invention is a system that photographs the cover image of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, recognizes the user's emotions using an emotion engine, and based on this, compares it with a bibliographic database to obtain accurate bibliographic information.It then uses emotion analysis to analyze the impressions and keywords entered by the user, automatically generates a designed POP based on the bibliographic information and the results of the emotion analysis, selects an appropriate song, and combines the generated design with the selected song to generate a short video, which is then distributed to social media.
[0985] Program processing
[0986] 1. Taking and uploading images
[0987] Users take a photo of the cover of a book they have finished reading with their smartphone, and the image is saved in the device app.
[0988] The device compresses the captured image and uploads it to the server, along with the user ID and session information.
[0989] 2. OCR processing
[0990] The server inputs the received image into an OCR engine, which extracts the book title, author name, and visual elements from the image.
[0991] 3. Obtaining bibliographic information
[0992] The server compares the OCR-extracted book title and author name with a book database to obtain accurate bibliographic information. If no bibliographic information is found, it presents alternative title suggestions and asks the user for confirmation.
[0993] 4. Emotion recognition
[0994] Users enter their impressions of the book they have read and keywords into the input form of the terminal app, and the input is automatically sent to the server.
[0995] The server uses an emotion engine to analyze the impressions and keywords entered by the user and recognize emotions. The emotion engine extracts emotions using natural language processing technology. Furthermore, if the user's facial image or voice is used, emotions can also be recognized from this data.
[0996] 5. Emotion analysis
[0997] The server analyzes the user's impressions, including the emotion recognition results, and extracts the main essence using an emotion analysis algorithm.
[0998] 6. POP Text Generation
[0999] The server generates POP text based on the results of sentiment analysis and bibliographic information, and the generated text is structured to reflect the emotional emphasis.
[1000] 7. Design Generation
[1001] The server runs an automated design engine that generates a POP image based on bibliographic information, sentiment analysis results, and the cover image, incorporating key design elements from the cover.
[1002] 8. Music Selection
[1003] The server then refers to the results of the emotion analysis and selects appropriate songs from a music database, which are chosen to match the emotional tone.
[1004] 9. Short video generation
[1005] The server combines the POP image with the selected music and generates a short video using a video generation tool such as FFmpeg, which is then sent to the device.
[1006] 10. Social Media Distribution
[1007] The device receives the short video sent from the server and displays a confirmation screen to the user, who then confirms the video and it is ready to post.
[1008] The user presses the "Post" button on the device app and selects the social media platform (e.g., Instagram, Twitter) to post the generated short video to.
[1009] The device will call the upload API of the selected social networking site to upload the short video, and the user will receive a notification when the post is complete.
[1010] Specific examples
[1011] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the "author's name" from the image. It then compares the image with a bibliographic database to obtain accurate bibliographic information. If a user writes their impression of a book as "full of adventure and emotion," the emotion engine analyzes it and uses a sentiment analysis algorithm to extract the key essences of "adventure" and "emotion." From this result, a POP text is generated, such as "an inspiring story with adventurous elements."
[1012] The automated design engine then generates a POP image based on the text, the acquired bibliographic information, and the cover image. Furthermore, it selects an appropriate song from a music database based on the results of sentiment analysis. Finally, a short video is generated by integrating the POP image and the selected song, which users can easily post to social media.
[1013] In this way, the system of the present invention allows users to easily create POPs like those found in bookstores at home and share them on social media.By incorporating an emotion engine, it is possible to provide more personalized POPs that reflect the user's emotions.
[1014] The processing flow will be explained below.
[1015] Step 1:
[1016] Users take a photo of the cover of a book they have finished reading with their smartphone, and the image is saved in the device's app storage.
[1017] Step 2:
[1018] The device compresses the captured image and uploads it to the server, which includes the user ID and session information.
[1019] Step 3:
[1020] The server inputs the received image into an OCR engine, which extracts the book title, author name, and visual elements from the image.
[1021] Step 4:
[1022] The server compares the book title and author name extracted by OCR with a book database to obtain accurate bibliographic information. If the bibliographic information is not found as a result of the comparison with the book database, it presents several alternative title candidates and asks the user for confirmation.
[1023] Step 5:
[1024] The user enters their thoughts about the book they have read and keywords into the input form of the terminal app. The input information is converted into JSON format and sent to the server.
[1025] Step 6:
[1026] The server uses an emotion engine to analyze the impressions and keywords entered by the user and recognize emotions. The emotion engine uses natural language processing technology to extract emotions from text. If the user also provides facial images or voice data, the engine can also recognize emotions from that data.
[1027] Step 7:
[1028] The server generates POP text based on the recognized sentiment and bibliographic information. The generated text reflects the essence obtained from the sentiment analysis results.
[1029] Step 8:
[1030] The server runs an automatic design engine to generate POP images based on bibliographic information, sentiment analysis results, and cover images. The automatic design engine creates POP images using design templates.
[1031] Step 9:
[1032] The server then refers to the emotion analysis results and selects appropriate songs from a music database that match the emotional tone.
[1033] Step 10:
[1034] The server combines the POP image with the selected music and generates a short video using a video generation tool such as FFmpeg, which is then sent to the device.
[1035] Step 11:
[1036] The device receives the short video sent from the server and displays a confirmation screen to the user, who can then confirm the video and indicate that it is ready to be posted.
[1037] Step 12:
[1038] The user presses the "Post" button on the device app and selects the social media platform (e.g., Instagram, Twitter) to post the generated short video to.
[1039] Step 13:
[1040] The device will call the upload API of the selected social networking service to upload the short video, and a notification will be displayed to the user once the post is successful.
[1041] This completes the entire process, allowing users to easily generate a POP for the book and share it on social media.
[1042] Example 2
[1043] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1044] While much information is shared on social media these days, there are limited ways for users to easily and effectively share their reading experiences. There is a particular need for sharing book reviews and ratings in a visually appealing format, but existing technologies require manual input and processing of information, which is time-consuming. Furthermore, it is difficult to generate personalized content that reflects user sentiment using sentiment analysis technology.
[1045] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an image analysis means, a bibliographic information acquisition means, an emotion analysis means, an electronic design generation means, a music selection means, a short video generation means, a short video distribution means, an image compression and upload means, an OCR analysis means, a user interaction means, a voice and face image analysis means, a display confirmation means, and an SNS posting means. This enables users to easily and attractively visualize their impressions and reviews of books they have finished reading and share them on SNS.
[1046] "Image analysis means" refers to means having the function of analyzing an image and extracting information.
[1047] The "bibliographic information acquisition means" is a means having a function for acquiring accurate book information related to a book category.
[1048] An "emotion analysis means" is a means that has the function of analyzing and recognizing emotions from impressions and keywords entered by the user.
[1049] "Electronic design generation means" means a means having a function for generating a design electronically.
[1050] The "music selection means" is a means having a function for selecting appropriate music based on the emotion analysis results.
[1051] The "short video generation means" is a means having a function for generating a short video by integrating a POP image with a selected piece of music.
[1052] "Short video distribution means" refers to a means that has the function of distributing the generated short video to social media and other platforms.
[1053] The "image compression and uploading means" is a means having a function for compressing a captured image and uploading it to a server.
[1054] "OCR Analysis Means" means a means capable of extracting book titles, author names, and visual elements from an image using optical character recognition technology.
[1055] "User interaction means" refers to means that has the function of providing an interface for users to input their thoughts and keywords into the system.
[1056] The "voice and facial image analysis means" is a means having a function for analyzing emotions from the voice and facial image data provided by the user.
[1057] The "display confirmation means" is a means having a function of allowing the user to confirm the generated content.
[1058] "SNS posting means" refers to a means that has the function of allowing users to easily post content they have viewed to SNS.
[1059] This invention is a system that takes a photo of the cover of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, recognizes the user's emotions using an emotion engine, and obtains accurate bibliographic information by matching it with a book database. Furthermore, it analyzes the impressions and keywords entered by the user using emotion analysis, automatically generates a POP design based on the bibliographic information and the emotion analysis results, selects an appropriate song, and generates a short video by integrating the generated design with the selected song, which is then distributed to social media.
[1060] The system is programmed as follows:
[1061] A user uses their smartphone camera to take a photo of the cover of a book they have just read. The captured image is saved in the device app. The device compresses the image and uploads it to the server, along with the user ID and session information. The server then inputs the uploaded image into an OCR engine (e.g., Google Cloud Vision API) to extract the book title, author name, and visual elements from the image.
[1062] The server compares the book title and author name extracted by OCR with a book database (e.g., Google Books API, Open Library) to obtain accurate bibliographic information. If bibliographic information is not found, alternative title suggestions are presented and the user is asked for confirmation. The user then enters their thoughts and keywords about the book they read into the input form of the device app. The input is automatically sent to the server. The server uses an emotion engine (e.g., IBM Watson NLU) to analyze the thoughts and keywords entered by the user and recognize emotions. If the user provides facial images or voice, emotions can also be recognized from this data, if necessary.
[1063] The server generates POP text based on the results of sentiment analysis and bibliographic information. The generated text is structured to reflect the emotional emphasis. The server then runs an automated design engine (e.g., Adobe Creative Cloud API) to generate a POP image based on the bibliographic information, sentiment analysis results, and cover image. The POP image also incorporates key design elements from the cover.
[1064] The server then refers to the emotion analysis results and selects an appropriate song from a music database (e.g., Spotify API, Apple Music API). The song is selected to match the emotional tone. The server then combines the POP image with the selected song and generates a short video using a video generation tool such as FFmpeg. The generated video data is then sent to the device.
[1065] The user can view the short video sent from the server on their device. After viewing, the user presses the "Post" button on the device app to post the video to a social networking platform (e.g., Instagram, Twitter). The device then calls the upload API of the selected social networking platform and uploads the short video. Once posting is complete, a notification is displayed to the user.
[1066] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the author's name from the image. It then compares the image with a book database to obtain accurate bibliographic information. If the user writes their impression, "This is a work filled with adventure and emotion," the emotion engine analyzes it and extracts the key essences of "adventure" and "emotion" using a sentiment analysis algorithm. From these results, the server creates a POP text such as "An inspiring story with adventurous elements."
[1067] The automated design engine then generates a POP image based on the text, the acquired bibliographic information, and the cover image. Furthermore, it selects an appropriate song from a music database based on the results of sentiment analysis. Finally, a short video is generated by integrating the POP image and the selected song, which users can easily post to social media.
[1068] The following are examples of prompt sentences:
[1069] "Please explain in natural language the process of a system that takes a photo of a book cover, extracts the book title and author name using OCR technology, analyzes the emotions felt after reading using an emotion engine, compares it with a book database to obtain bibliographic information, generates POP text and images based on the emotion analysis results, selects music, creates a video using a short video creation tool (e.g., FFmpeg), and posts it to social media."
[1070] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1071] Step 1: Capture and upload images
[1072] The user takes a photo of the book cover using the smartphone camera. The captured image is automatically saved in the device app. The device compresses the saved image and uploads it to the server. When uploading, the user ID and session information are also sent. This sends the compressed image data and user information to the server.
[1073] Step 2: OCR
[1074] The server inputs the uploaded image into an OCR engine (e.g., optical character recognition software). The OCR engine extracts the book title, author name, and visual elements from the image. It receives image data as input and obtains text data of the book title, author name, and visual elements as output. Specifically, the OCR engine identifies character regions in the image and reads character data from those regions.
[1075] Step 3: Obtain bibliographic information
[1076] The server uses the book title and author name extracted by OCR to check against a book database (e.g., a book information service). This allows accurate bibliographic information to be obtained. It receives text data of the book title and author name as input, and obtains detailed book information (e.g., publication year, genre, summary) as output. Specifically, it sends a request to the book database via an API and obtains the corresponding book information.
[1077] Step 4: Emotion Recognition
[1078] The user enters their thoughts and keywords about the book they have read into an input form in the device app. The input is automatically sent to the server. The server uses an emotion engine (e.g., natural language processing software) to analyze the thoughts and keywords entered by the user and recognize the emotion. It receives the text data of the thoughts and keywords as input and obtains the type of emotion (e.g., joy, sadness) as output. Specifically, the emotion engine analyzes the text and runs an algorithm to classify the emotion.
[1079] Step 5: Sentiment Analysis
[1080] The server analyzes the emotion results recognized by the emotion engine and the user's impressions, and extracts the main essence using an emotion analysis algorithm. It receives emotion data and impression text as input, and obtains the extracted essence (e.g., adventure, emotion) as output. Specifically, it uses a text analysis algorithm to extract important keywords and themes from the impressions.
[1081] Step 6: POP Text Generation
[1082] The server generates POP text based on the results of sentiment analysis and bibliographic information. It receives the sentiment essence and bibliographic information as input and obtains a catchy slogan and description as output. Specifically, it uses a template engine to compose the text and adjust it to reflect the emotional emphasis.
[1083] Step 7: Design Generation
[1084] The server runs an automatic design engine (e.g., design software API) to generate POP images based on bibliographic information, sentiment analysis results, and cover images. It receives design elements (e.g., book title, author name, emotional essence) as input and obtains POP image data as output. Specifically, it automatically adjusts the layout and coloring to generate visually appealing images.
[1085] Step 8: Music Selection
[1086] The server refers to the emotion analysis results and selects appropriate songs from a music database (e.g., a music streaming service). It receives the emotion essence as input and obtains song information as output. Specifically, it searches for songs that match the emotion and executes an algorithm to select the most suitable one.
[1087] Step 9: Short video generation
[1088] The server combines the POP image with the selected music and generates a short video using a video generation tool (e.g., FFmpeg). It receives POP image data and music information as input and obtains a short video file as output. Specifically, it adjusts the timing of the image and music and encodes them as continuous visual content.
[1089] Step 10: Social Media Distribution
[1090] The device receives the short video sent from the server and displays a confirmation screen to the user. The user checks the video and presses the "Post" button to post the video to the SNS platform (e.g., SNS service API). The device receives the short video file as input and receives a notification that posting to the SNS has been completed as output. Specifically, the device calls the SNS upload API and displays a notification to the user when posting is successful.
[1091] (Application example 2)
[1092] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1093] Traditionally, creating recommended book POPs in bookstores required a lot of time and effort, and the content of the POPs was not emotionally personalized, limiting their appeal to customers. Furthermore, updating promotional materials displayed on in-store displays was done manually, resulting in a lack of immediacy. For these reasons, a method was needed for bookstore staff to quickly and effectively promote books.
[1094] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an image analysis means, a bibliographic information acquisition means, a sentiment analysis means, an electronic design generation means, a music selection means, a short video generation means, a short video distribution means, and an automatic display means. This allows bookstore staff to simply take a photo of a book cover with their smartphone, and automatically generate an individual recommended POP based on related bibliographic information, user reviews, and sentiment analysis, and display it in real time on an electronic display in the store or post it to social media.
[1095] "Image analysis means" refers to means for extracting information from captured images.
[1096] "Bibliographic information acquisition means" refers to a means for acquiring information related to a book (e.g., book title, author name) from a database.
[1097] "Sentiment analysis means" means means for analyzing emotions from user input or other data.
[1098] "Electronic design generation means" refers to a means for automatically generating designs based on book information and sentiment analysis results.
[1099] The "music selection means" is a means for selecting appropriate music based on the emotion analysis results.
[1100] The "short video generation means" is a means for integrating a POP image with selected music to generate a short video.
[1101] "Short video distribution means" refers to a means for distributing the generated short video to an SNS platform.
[1102] "Automatic display means" refers to a means for displaying the generated POP images and short videos on electronic displays in the store in real time.
[1103] This system allows bookstore staff to take a photo of a book cover with their smartphone, automatically generating recommended POPs based on related bibliographic information, user reviews, and sentiment analysis, and posting them on in-store electronic displays and social media. The details of this system are described below.
[1104] The system starts by having the user take a photo of the book cover with their smartphone. The image is compressed and uploaded to a server, where it is analyzed using image analysis tools (e.g., pytesseract) to extract the book title, author, and visual elements.
[1105] Next, the server uses the bibliographic information acquisition means to acquire accurate bibliographic information from the book database, including the book title, author name, publication year, genre, etc.
[1106] The server then uses sentiment analysis tools to analyze the reviews and keywords entered by the user and recognize emotions. This process includes an emotion engine using natural language processing techniques. Facial images and voice data may also be used for emotion recognition.
[1107] Based on the results of the sentiment analysis and the bibliographic information, the server automatically generates POP text using an electronic design generation tool, and then creates a POP image using a design engine. This POP image reflects the bibliographic information and the user's sentiment.
[1108] Furthermore, the server uses a music selection means to select appropriate music from a music database based on the result of the emotion analysis, and the music is selected to match the emotional tone.
[1109] The server combines these POP images with the selected music and generates a short video using a short video generation tool (e.g., FFmpeg). This video is automatically generated and delivered to the user's smartphone.
[1110] Finally, the generated short video is posted to a social networking site using a short video distribution method. Users can select a social networking site (e.g., Instagram or Twitter) on their smartphone to complete the posting.
[1111] To accommodate in-store promotions, the system is equipped with an automatic display means, which displays the generated POP images on electronic displays in the store in real time.
[1112] Hardware and Software Use
[1113] The main hardware used in this invention is a smartphone and an electronic display in a bookstore, and the main software is an image analysis engine (pytesseract), a sentiment analysis engine, a design generation API, a video generation tool (FFmpeg), and a SNS upload API.
[1114] Specific examples
[1115] For example, if a user takes a photo of the cover of a "fantasy novel" and writes in their review that it is "a work filled with adventure and emotion," the system will analyze this information and generate a pop-up image that emphasizes "adventure" and "emotion." This image is then combined with a selected song to create a short video. The video can then be displayed on in-store displays and posted to social media.
[1116] Prompt Sentence Examples
[1117] "Users take a photo of the cover of a book they have read and enter their thoughts on the book. The system analyzes the information and generates a short video that combines recommended pop music and posts it to social media."
[1118] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1119] Step 1:
[1120] Image capture and upload
[1121] The user takes a photo of the book cover with their smartphone. This image is saved in the smartphone app. The device compresses the image file and uploads it to the server. The user ID and session information are also sent at the time of uploading.
[1122] Step 2:
[1123] OCR processing
[1124] The server inputs the received image into an OCR engine (pytesseract), which extracts text from the image and obtains information such as the book title and author name. The extracted text data is sent to the next processing step.
[1125] Step 3:
[1126] Bibliographic information acquisition
[1127] The server uses the book title and author name extracted by OCR to retrieve detailed bibliographic information from a book database. Using the book title and author name as input data, the server compares the book title, author name, publication year, genre, and other bibliographic information to output.
[1128] Step 4:
[1129] emotion recognition
[1130] The user enters their thoughts and keywords about the book they have read into an input form on a smartphone app. This text data is sent to a server. The server uses an emotion analysis tool (emotion engine) to analyze the thoughts and keywords entered by the user and recognize the emotion. The emotion analysis engine takes this data as input and outputs the type of emotion (e.g., joy, sadness).
[1131] Step 5:
[1132] Emotion analysis
[1133] The server analyzes the sentiment data, including the emotion recognition results, and uses a sentiment analysis algorithm to extract the key essence and generate the data needed to generate POPs. The input is the emotion recognition results and user sentiment data, and the output is the essence extraction results.
[1134] Step 6:
[1135] POP Text Generation
[1136] The server automatically generates POP text based on the results of sentiment analysis and bibliographic information. The generated text reflects the emotional emphasis. The input is the sentiment analysis results and bibliographic information, and the output is POP text.
[1137] Step 7:
[1138] Design Generation
[1139] The server runs an electronic design generator to automatically generate a POP image based on the sentiment analysis results, bibliographic information, and cover image. This POP image also incorporates the main design elements of the cover. The input is the POP text, bibliographic information, and cover image, and the output is the POP image.
[1140] Step 8:
[1141] Music Selection
[1142] The server refers to the results of the emotion analysis and selects an appropriate song from a music database. The selected song matches the emotional tone. The input is the emotion analysis result, and the output is the selected song.
[1143] Step 9:
[1144] Short video generation
[1145] The server combines the POP image and the selected music and generates a short video using a short video generation tool (FFmpeg). The input is the POP image and music URL, and the output is the short video data.
[1146] Step 10:
[1147] SNS distribution
[1148] The device receives the short video sent from the server and displays a confirmation screen to the user. The user reviews the video and is ready to post. When the user selects an SNS platform (e.g., Instagram, Twitter) and presses the "Post" button, the device calls the upload API of the selected SNS and uploads the short video. The input is the short video data and the SNS information selected by the user, and the output is a notification of successful posting to the SNS.
[1149] Step 11:
[1150] Automatic display
[1151] The server uses an automatic display means to display the generated POP image on an electronic display in the store in real time. The input is the POP image, and the output is the display on the in-store display.
[1152] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1153] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1154] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1155] [Fourth embodiment]
[1156] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1157] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1158] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1159] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1160] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1161] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1162] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1163] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1164] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1165] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1166] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1167] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1168] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1169] This invention is a system that takes a photo of the cover of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, compares this with a bibliographic database to obtain accurate bibliographic information, and then uses sentiment analysis to analyze the impressions and keywords entered by the user.It then automatically generates a POP design based on the bibliographic information and the results of the sentiment analysis, selects an appropriate song, and combines the generated design with the selected song to generate a short video, which is then distributed to social media.
[1170] Program processing
[1171] 1. Taking and uploading images
[1172] User: Take a photo of the cover of a book you have just read using your smartphone or other device.
[1173] Device: Upload the captured image to the server via the device app.
[1174] 2. OCR processing
[1175] Server: Receives the uploaded image and inputs it into the OCR engine, which extracts the book title, author name, and visual elements from the image.
[1176] 3. Obtaining bibliographic information
[1177] Server: Based on the extracted book title and author name, it checks the book database to get the correct bibliographic information. If no bibliographic information is found, it generates feedback to the user requesting manual input.
[1178] 4. Enter your thoughts
[1179] User: Enter their thoughts and keywords about the book they read into the input form in the device app.
[1180] 5. Emotion analysis
[1181] Server: Analyzes the user-entered comments and extracts the main essence using a sentiment analysis algorithm.
[1182] 6. POP Text Generation
[1183] Server: Generates POP text based on the results of sentiment analysis and bibliographic information.
[1184] 7. Design Generation
[1185] Server: Runs the automatic design engine and generates POP images based on bibliographic information, sentiment analysis results, and cover images.
[1186] 8. Music Selection
[1187] Server: Refers to the results of the sentiment analysis and selects appropriate songs from music databases such as LINE MUSIC.
[1188] 9. Short video generation
[1189] Server: Integrates POP images with selected music to generate short videos.
[1190] Device: Receives short video data sent from the server and converts it into a format that can be posted to social media.
[1191] 10. Social Media Distribution
[1192] User: Press the "Post" button on the device app and select the generated short video.
[1193] Device: Calls an API to upload the video to the social networking site of the user's choice, and notifies the user when the post is complete.
[1194] Specific examples
[1195] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the author's name from the image. It then compares the image with a bibliographic database to obtain accurate bibliographic information. If the user inputs a sentiment such as "This is a work filled with adventure and emotion," the sentiment analysis algorithm analyzes this and generates POP text. Using the keyword "adventure," a song that evokes the image of adventure is selected. Finally, the cover image, POP text, and selected song are integrated to generate a short video, which users can post on social media to share the appeal of the book with many people.
[1196] In this way, the system of the present invention allows users to easily create POPs like those found in bookstores at home and share them on social media.
[1197] The processing flow will be explained below.
[1198] Step 1:
[1199] Users take a photo of the cover of a book they have finished reading with their smartphone, and the image is saved in the device app.
[1200] Step 2:
[1201] The device compresses the captured image and uploads it to the server, along with the user ID and session information.
[1202] Step 3:
[1203] The server inputs the received image into an OCR engine, which extracts the book title, author name, and visual elements from the image.
[1204] Step 4:
[1205] The server compares the OCR-extracted book title and author name with a book database to obtain accurate bibliographic information. If no bibliographic information is found, it presents alternative title suggestions and asks the user for confirmation.
[1206] Step 5:
[1207] Users enter their impressions of the book they have read and keywords into the input form of the terminal app, and the input is automatically sent to the server.
[1208] Step 6:
[1209] The server analyzes the user's input and uses a sentiment analysis algorithm to extract the main essence, thereby determining which sentiment prevails.
[1210] Step 7:
[1211] The server generates POP text based on the results of sentiment analysis and bibliographic information, and the generated text is structured to reflect the emotional emphasis.
[1212] Step 8:
[1213] The server runs an automated design engine that generates a POP image based on bibliographic information, sentiment analysis results, and the cover image, incorporating key design elements from the cover.
[1214] Step 9:
[1215] The server then refers to the results of the emotion analysis and selects appropriate songs from a music database, which are chosen to match the emotional tone.
[1216] Step 10:
[1217] The server combines the POP image with the selected music and generates a short video using a video generation tool such as FFmpeg, which is then sent to the device.
[1218] Step 11:
[1219] The device receives the short video sent from the server and displays a confirmation screen to the user, who then confirms the video and it is ready to post.
[1220] Step 12:
[1221] The user presses the "Post" button on the device app and selects the social media platform (e.g., Instagram, Twitter) to post the generated short video to.
[1222] Step 13:
[1223] The device will call the upload API of the selected social networking site to upload the short video, and the user will receive a notification when the post is complete.
[1224] This completes the entire process, allowing users to easily generate a POP for the book and share it on social media.
[1225] Example 1
[1226] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1227] Conventional systems lack a way for users to efficiently and accurately record information about books they have read and share that content in an engaging way on social media. This requires manual input and editing, which takes a lot of time and effort. Furthermore, analyzing impressions and keywords and selecting design and music based on those must also be done manually, resulting in an inconsistent user experience. Furthermore, distributing the generated content to social media can be difficult.
[1228] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1229] In this invention, the server includes image analysis means, bibliographic information acquisition means, emotion analysis means, electronic design generation means, music selection means, short video generation means, short video distribution means, a terminal for uploading images to the server via a user interface, a terminal for accepting user feedback, means for comparing extracted book titles and author names with a database to acquire bibliographic information, text generation means for generating POP text based on the emotion analysis results, and means for the short video generation means to integrate the images, generated text, and selected music. This allows users to easily and efficiently record information about books they have read, create attractive short videos that reflect the emotion analysis results, and share them on social media.
[1230] An "image analysis tool" is a device or software that uses machine learning algorithms or optical character recognition techniques to extract textual information and visual elements from an image.
[1231] The "bibliographic information acquisition means" refers to a device or software that compares the extracted book title and author name with a database to acquire accurate bibliographic information.
[1232] "Emotion analysis means" refers to a device or software that analyzes text data entered by a user and extracts emotions and key essences.
[1233] An "electronic design generator" is a device or software that automatically creates an attractive design based on POP text and visual elements.
[1234] A "music selection means" is a device or software that selects appropriate music based on the results of emotion analysis.
[1235] The "short video generation means" refers to a device or software that integrates the generated design and selected music to create a short video.
[1236] "Short video distribution means" refers to devices or software used to upload and distribute completed short videos to platforms such as social media.
[1237] A "terminal that uploads images to a server via a user interface" is a device that has hardware or software for sending images taken by a user to a server.
[1238] A "terminal that accepts user feedback input" is a device that provides an input form in which users can input their feedback or keywords, and that is equipped with hardware or software that accepts that data.
[1239] "Means for obtaining bibliographic information by comparing extracted book titles and author names with a database" refers to devices or software that executes the process of using book titles and author names extracted using OCR technology to compare them with an external database and obtain accurate bibliographic information.
[1240] "Text generation means for generating POP text based on the results of sentiment analysis" refers to a device or software that uses the results of sentiment analysis to automatically generate text that succinctly conveys the features and appeal of a book.
[1241] This system photographs the cover image of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, compares this with a bibliographic database to obtain accurate bibliographic information, and then analyzes the impressions and keywords entered by the user using sentiment analysis.The system then automatically generates a POP design based on the bibliographic information and the results of the sentiment analysis, selects an appropriate song, and combines the generated design with the selected song to create a short video, which is then distributed to social media.
[1242] Hardware and software used
[1243] Image analysis methods
[1244] The server stores the cover image uploaded by the user in Google Cloud Storage, then uses the Google Cloud Vision OCR API to extract the book title, author name, and cover image. The OCR engine takes image data as input and outputs text information.
[1245] Bibliographic information acquisition method
[1246] The server sends the extracted book title and author name to the Google Books API to get the exact bibliographic information, which is then returned from the database and stored in an internal database.
[1247] Emotion analysis means
[1248] The server inputs the user's feedback text into the IBM Watson Natural Language Understanding API to extract sentiment and key points, and the analysis results are also stored in a database.
[1249] Electronic Design Generator
[1250] The server uses the Canva API to generate POP images based on bibliographic information, sentiment analysis results, and cover images. The resulting designs are visually appealing and have a compelling layout.
[1251] Music selection method
[1252] Based on the results of the emotion analysis, the server accesses the LINE MUSIC API to select the appropriate song, and records the selected song data in a database.
[1253] Short video creation method
[1254] The server then combines the POP image and the selected music using video editing software such as Adobe Premiere API to create a short video, which is then stored in Google Cloud Storage.
[1255] Short video distribution methods
[1256] The device converts the short video data received from the server into a format that can be posted to SNS. When the user taps the "Post" button in the app, the device calls the SNS's API and uploads the video.
[1257] Examples of concrete examples and prompts
[1258] For example, a user takes a photo of the cover of a "fantasy novel" with their smartphone. The server sends this image to Google Cloud Vision OCR, which extracts "fantasy novel" and the author's name from the image. Next, it uses the Google Books API to obtain accurate bibliographic information. If the user enters their impression that the work is "full of adventure and emotion," the server performs sentiment analysis using the IBM Watson Natural Language Understanding API and extracts the key essence, "adventure." The server selects music related to "adventure" from the LINE MUSIC API and generates a pop-up image using the Canva API. Finally, the pop-up image and music are integrated using the Adobe Premiere API to generate a short video, which the user can post on social media.
[1259] An example of a prompt is as follows:
[1260] "Please build a system that takes a photo of the cover of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, and compares it with a bibliographic database. Analyzes the user's impressions using sentiment analysis, automatically generates POPs, and distributes short videos with appropriate music on social media."
[1261] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1262] Step 1:
[1263] User: Uses smartphone camera to take a picture of the cover of the book they have just read. Checks that the image is clear. The input is the image of the book cover, and the output is the image data.
[1264] Step 2:
[1265] Terminal: The captured image data is received within a dedicated app, compressed, and uploaded to the server. The input is the image data, and the output is the compressed image data sent to the server.
[1266] Step 3:
[1267] Server: Receives uploaded image data and stores it in Google Cloud Storage. Sends the stored image data to the Google Cloud Vision OCR API to extract the book title, author name, and cover image. The input is compressed image data, and the output is extracted text information.
[1268] Step 4:
[1269] Server: Sends the book title and author name obtained through OCR processing to the Google Books API to obtain bibliographic information. Stores the bibliographic information in an internal database and verifies the accuracy of the book title and author name. The input is the extracted text information, and the output is the bibliographic information stored in the database.
[1270] Step 5:
[1271] User: Follow the prompts from the system and enter your thoughts and keywords about the book you read into the input form within the app. Once you've finished entering your thoughts and keywords, tap the "Submit" button. The input is text data of your thoughts and keywords, and the output is text data sent to the server.
[1272] Step 6:
[1273] Server: Sends the sentiment text data to the IBM Watson Natural Language Understanding API, extracts sentiment and key essence, and stores the resulting analysis data in a database. The input is sentiment text data, and the output is the sentiment analysis results.
[1274] Step 7:
[1275] Server: Automatically generates POP text based on the sentiment analysis results and bibliographic information. The generated text succinctly summarizes the "features of this book" and includes visually appealing phrases. The input is the sentiment analysis results and bibliographic information, and the output is the generated POP text.
[1276] Step 8:
[1277] Server: Automatically generates digital designs using the Canva API. Generates POP images based on bibliographic information, sentiment analysis results, and cover images, creating visually appealing layouts. The inputs are bibliographic information, sentiment analysis results, and cover images, and the output is a POP image.
[1278] Step 9:
[1279] Server: Based on the results of the sentiment analysis, the server uses the LINE MUSIC API to select appropriate songs. The selected song data is recorded in a database. The input is the sentiment analysis results, and the output is the selected song data.
[1280] Step 10:
[1281] Server: The POP images and selected music are integrated using video editing software such as Adobe Premiere API to generate a short video. The generated video data is stored in Google Cloud Storage. The input is the POP images and music data, and the output is the generated short video.
[1282] Step 11:
[1283] User: Press the "Post" button in the app and select the generated short video. Select the social media platform to post to (e.g. Instagram, Twitter) and press the "Send" button. The input is the short video to be uploaded, and the output is the video posted to the social media platform.
[1284] Step 12:
[1285] Terminal: Calls the API of the selected SNS and uploads the video. When the upload is complete, a completion notification is displayed to the user. The input is the short video data, and the output is the video uploaded to the SNS.
[1286] (Application example 1)
[1287] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1288] To effectively share their impressions and reviews of a book after reading it, users need a simple and efficient process for creating a video with appropriate visual content and music, and then sharing it on social networking services. However, current technology requires these steps to be performed individually, which is tedious and can lack accuracy and consistency. An integrated system is needed to solve this problem.
[1289] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1290] In this invention, the server includes an image analysis means, a bibliographic information acquisition means, a sentiment analysis means, an electronic design generation means, a music selection means, a short video generation means, a short video distribution means, a means for taking an image from a user terminal and inputting the data, and a means for sharing and distributing the generated short video to a social networking service. This makes it possible to automatically generate a short video containing visual content that combines bibliographic information and sentiment analysis results and appropriate music based on the user's impressions of a book they have finished reading, and to easily share the generated short video on a social networking service.
[1291] "Image analysis means" refers to technology for extracting information from images taken by a user.
[1292] "Bibliographic information acquisition means" refers to technology for acquiring accurate bibliographic information about a book based on the extracted information.
[1293] "Emotion analysis means" refers to technology that analyzes the impressions and keywords entered by users and understands their emotions.
[1294] "Electronic design generation means" refers to technology for automatically generating POP text and graphics based on sentiment analysis results and bibliographic information.
[1295] "Music selection means" refers to technology for selecting appropriate music based on the results of emotion analysis.
[1296] "Short video generation means" refers to the technology for generating short videos by combining POP images with selected music.
[1297] "Short video distribution means" refers to technology for distributing the generated short videos to social networking services, etc.
[1298] "Means for taking pictures and inputting data from a user device" refers to technology that allows users to take pictures of book covers using devices such as smartphones and input that data into the system.
[1299] "Means for sharing and distributing the generated short video on a social networking service" refers to technology that enables the generated short video to be easily posted and shared on a social networking service.
[1300] A system for implementing this invention utilizes book information and impressions based on a user's reading experience to automatically generate short videos containing visual content and music, and share them on social networking services.
[1301] First, the user takes a photo of the cover of the book they have just finished reading using their smartphone. The user device takes the image and uploads the data to a server via an application. This is where the user device, including the smartphone, comes into play.
[1302] The server then receives the uploaded image and performs optical character recognition (OCR) using image analysis techniques to extract the book title, author name, and visual elements from the image. The OCR engine used is Tesseract or similar.
[1303] The server then uses the extracted title and author name to search the book database to find the correct bibliographic information, or if the bibliographic information is not found, prompt the user to enter it manually.
[1304] The user then enters their thoughts about the book and keywords into the input form of the terminal application. The server receives the thoughts and performs sentiment analysis using a sentiment analysis engine such as the Google Cloud Natural Language API.
[1305] The server generates POP text based on the results of sentiment analysis and bibliographic information. Next, the electronic design generation means operates and automatically generates POP images based on the bibliographic information, sentiment analysis results, and cover image. An automatic design engine is used.
[1306] Furthermore, the results of the sentiment analysis are used to select appropriate songs from a music database. For example, songs that evoke adventures are selected based on impressions that evoke adventures. A suitable music database can be found on a general music distribution service.
[1307] Finally, the server combines the POP image with the selected music to generate a short video. This method of generating short videos may utilize a video editing library. The generated short video is sent from the server to the user's device, and the user can easily share it on social networking services by pressing the "post" button on their device. By utilizing the API of the SNS platform, users can seamlessly share content.
[1308] Specific examples
[1309] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the author's name from the image. It then compares the image with a bibliographic database to obtain accurate bibliographic information. If the user inputs their impression that the book is "full of adventure and emotion," the sentiment analysis algorithm analyzes this and generates a pop-up message with an "adventure" theme. A song that evokes the image of adventure is selected based on the keyword "adventure." Finally, a short video is generated by integrating the cover image, pop-up message, and selected song, which users can post on social media to share the appeal of the book with many people.
[1310] Example prompt sentence:
[1311] "You can take a photo of the cover of a fantasy novel and input your impression that it's 'a work filled with adventure and emotion.' The system will then use OCR to extract information about the book, perform a sentiment analysis, select an adventure-themed POP design and music, and create a short video to post on social media."
[1312] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1313] Step 1:
[1314] Users take a photo of the cover of a book they have finished reading using a device such as a smartphone, and the image is uploaded to a server via a device application.
[1315] Input: An image file taken on the user's device
[1316] Output: Image data uploaded to the server
[1317] Step 2:
[1318] The server receives the uploaded image and performs optical character recognition (OCR) using image analysis techniques. The OCR engine extracts the book title, author name, and visual elements from the image.
[1319] Input: Uploaded image data
[1320] Output: Text data of book title, author name, and visual elements
[1321] Specific operation: The OCR engine running on the server extracts text from the image and generates analysis results.
[1322] Step 3:
[1323] The server then searches the bibliographic database based on the extracted title and author name to retrieve the correct bibliographic information, and if no bibliographic information is found, generates feedback to the user requesting manual input.
[1324] Input: Text data of book title and author name
[1325] Output: Bibliographic information (e.g., publisher name, publication date, ISBN, etc.)
[1326] Specific operation: The server communicates with the bibliographic database to search and retrieve the corresponding bibliographic information.
[1327] Step 4:
[1328] Users enter their thoughts and keywords about the book they have just read into the input form in the terminal application.
[1329] Input: User-entered comments and keywords
[1330] Output: Text data of impressions and keywords
[1331] Specific operation: The user enters their thoughts and keywords in text format into the input form on the device.
[1332] Step 5:
[1333] The server receives the inputted impressions and performs sentiment analysis using the sentiment analysis engine, which extracts the main essence and classifies the emotions into numerical values and categories.
[1334] Input: Text data of impressions and keywords
[1335] Output: Sentiment analysis results (e.g., positive, negative, main essence)
[1336] What it does: The sentiment analysis engine analyzes the sentiment text and extracts sentiment and key topics.
[1337] Step 6:
[1338] The server generates POP text based on the results of sentiment analysis and bibliographic information. The generated POP text will highlight the book's appeal.
[1339] Input: Sentiment analysis results, bibliographic information
[1340] Output: POP text
[1341] Specific operation: The server uses templates and generation algorithms to automatically generate text that combines sentiment analysis results and bibliographic information.
[1342] Step 7:
[1343] Using an electronic design generation tool, a POP image is generated based on bibliographic information, sentiment analysis results, and cover image.
[1344] Input: Bibliographic information, sentiment analysis results, cover image
[1345] Output: POP image
[1346] How it works: The online design tool automatically generates visually appealing POP images based on bibliographic information and sentiment analysis results.
[1347] Step 8:
[1348] The results of the sentiment analysis are used to select appropriate songs from a music database.
[1349] Input: Sentiment analysis results
[1350] Output: Selected music files
[1351] Specific operation: The server searches the music database and selects music files that match the emotion analysis results.
[1352] Step 9:
[1353] The server combines the POP image with the selected music to generate a short video.
[1354] Input: POP image, selected music file
[1355] Output: Short video file
[1356] What it does: The video editing library merges the POP image with the selected music to generate a short video in a specific format.
[1357] Step 10:
[1358] The user's device receives the short video data sent from the server and converts it into a format that can be posted to SNS. When the user presses the "Post" button, the generated short video is shared and distributed to the selected social networking service.
[1359] Input: Short video file
[1360] Output: Content posted to social media
[1361] Specific operation: The terminal application converts the short video and calls the SNS API to post it.
[1362] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1363] This invention is a system that photographs the cover image of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, recognizes the user's emotions using an emotion engine, and based on this, compares it with a bibliographic database to obtain accurate bibliographic information.It then uses emotion analysis to analyze the impressions and keywords entered by the user, automatically generates a designed POP based on the bibliographic information and the results of the emotion analysis, selects an appropriate song, and combines the generated design with the selected song to generate a short video, which is then distributed to social media.
[1364] Program processing
[1365] 1. Taking and uploading images
[1366] Users take a photo of the cover of a book they have finished reading with their smartphone, and the image is saved in the device app.
[1367] The device compresses the captured image and uploads it to the server, along with the user ID and session information.
[1368] 2. OCR processing
[1369] The server inputs the received image into an OCR engine, which extracts the book title, author name, and visual elements from the image.
[1370] 3. Obtaining bibliographic information
[1371] The server compares the OCR-extracted book title and author name with a book database to obtain accurate bibliographic information. If no bibliographic information is found, it presents alternative title suggestions and asks the user for confirmation.
[1372] 4. Emotion recognition
[1373] Users enter their impressions of the book they have read and keywords into the input form of the terminal app, and the input is automatically sent to the server.
[1374] The server uses an emotion engine to analyze the impressions and keywords entered by the user and recognize emotions. The emotion engine extracts emotions using natural language processing technology. Furthermore, if the user's facial image or voice is used, emotions can also be recognized from this data.
[1375] 5. Emotion analysis
[1376] The server analyzes the user's impressions, including the emotion recognition results, and extracts the main essence using an emotion analysis algorithm.
[1377] 6. POP Text Generation
[1378] The server generates POP text based on the results of sentiment analysis and bibliographic information, and the generated text is structured to reflect the emotional emphasis.
[1379] 7. Design Generation
[1380] The server runs an automated design engine that generates a POP image based on bibliographic information, sentiment analysis results, and the cover image, incorporating key design elements from the cover.
[1381] 8. Music Selection
[1382] The server then refers to the results of the emotion analysis and selects appropriate songs from a music database, which are chosen to match the emotional tone.
[1383] 9. Short video generation
[1384] The server combines the POP image with the selected music and generates a short video using a video generation tool such as FFmpeg, which is then sent to the device.
[1385] 10. Social Media Distribution
[1386] The device receives the short video sent from the server and displays a confirmation screen to the user, who then confirms the video and it is ready to post.
[1387] The user presses the "Post" button on the device app and selects the social media platform (e.g., Instagram, Twitter) to post the generated short video to.
[1388] The device will call the upload API of the selected social networking site to upload the short video, and the user will receive a notification when the post is complete.
[1389] Specific examples
[1390] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the "author's name" from the image. It then compares the image with a bibliographic database to obtain accurate bibliographic information. If a user writes their impression of a book as "full of adventure and emotion," the emotion engine analyzes it and uses a sentiment analysis algorithm to extract the key essences of "adventure" and "emotion." From this result, a POP text is generated, such as "an inspiring story with adventurous elements."
[1391] The automated design engine then generates a POP image based on the text, the acquired bibliographic information, and the cover image. Furthermore, it selects an appropriate song from a music database based on the results of sentiment analysis. Finally, a short video is generated by integrating the POP image and the selected song, which users can easily post to social media.
[1392] In this way, the system of the present invention allows users to easily create POPs like those found in bookstores at home and share them on social media.By incorporating an emotion engine, it is possible to provide more personalized POPs that reflect the user's emotions.
[1393] The processing flow will be explained below.
[1394] Step 1:
[1395] Users take a photo of the cover of a book they have finished reading with their smartphone, and the image is saved in the device's app storage.
[1396] Step 2:
[1397] The device compresses the captured image and uploads it to the server, which includes the user ID and session information.
[1398] Step 3:
[1399] The server inputs the received image into an OCR engine, which extracts the book title, author name, and visual elements from the image.
[1400] Step 4:
[1401] The server compares the book title and author name extracted by OCR with a book database to obtain accurate bibliographic information. If the bibliographic information is not found as a result of the comparison with the book database, it presents several alternative title candidates and asks the user for confirmation.
[1402] Step 5:
[1403] The user enters their thoughts about the book they have read and keywords into the input form of the terminal app. The input information is converted into JSON format and sent to the server.
[1404] Step 6:
[1405] The server uses an emotion engine to analyze the impressions and keywords entered by the user and recognize emotions. The emotion engine uses natural language processing technology to extract emotions from text. If the user also provides facial images or voice data, the engine can also recognize emotions from that data.
[1406] Step 7:
[1407] The server generates POP text based on the recognized sentiment and bibliographic information. The generated text reflects the essence obtained from the sentiment analysis results.
[1408] Step 8:
[1409] The server runs an automatic design engine to generate POP images based on bibliographic information, sentiment analysis results, and cover images. The automatic design engine creates POP images using design templates.
[1410] Step 9:
[1411] The server then refers to the emotion analysis results and selects appropriate songs from a music database that match the emotional tone.
[1412] Step 10:
[1413] The server combines the POP image with the selected music and generates a short video using a video generation tool such as FFmpeg, which is then sent to the device.
[1414] Step 11:
[1415] The device receives the short video sent from the server and displays a confirmation screen to the user, who can then confirm the video and indicate that it is ready to be posted.
[1416] Step 12:
[1417] The user presses the "Post" button on the device app and selects the social media platform (e.g., Instagram, Twitter) to post the generated short video to.
[1418] Step 13:
[1419] The device will call the upload API of the selected social networking service to upload the short video, and a notification will be displayed to the user once the post is successful.
[1420] This completes the entire process, allowing users to easily generate a POP for the book and share it on social media.
[1421] Example 2
[1422] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1423] While much information is shared on social media these days, there are limited ways for users to easily and effectively share their reading experiences. There is a particular need for sharing book reviews and ratings in a visually appealing format, but existing technologies require manual input and processing of information, which is time-consuming. Furthermore, it is difficult to generate personalized content that reflects user sentiment using sentiment analysis technology.
[1424] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an image analysis means, a bibliographic information acquisition means, an emotion analysis means, an electronic design generation means, a music selection means, a short video generation means, a short video distribution means, an image compression and upload means, an OCR analysis means, a user interaction means, a voice and face image analysis means, a display confirmation means, and an SNS posting means. This enables users to easily and attractively visualize their impressions and reviews of books they have finished reading and share them on SNS.
[1425] "Image analysis means" refers to means having the function of analyzing an image and extracting information.
[1426] The "bibliographic information acquisition means" is a means having a function for acquiring accurate book information related to a book category.
[1427] An "emotion analysis means" is a means that has the function of analyzing and recognizing emotions from impressions and keywords entered by the user.
[1428] "Electronic design generation means" means a means having a function for generating a design electronically.
[1429] The "music selection means" is a means having a function for selecting appropriate music based on the emotion analysis results.
[1430] The "short video generation means" is a means having a function for generating a short video by integrating a POP image with a selected piece of music.
[1431] "Short video distribution means" refers to a means that has the function of distributing the generated short video to social media and other platforms.
[1432] The "image compression and uploading means" is a means having a function for compressing a captured image and uploading it to a server.
[1433] "OCR Analysis Means" means a means capable of extracting book titles, author names, and visual elements from an image using optical character recognition technology.
[1434] "User interaction means" refers to means that has the function of providing an interface for users to input their thoughts and keywords into the system.
[1435] The "voice and facial image analysis means" is a means having a function for analyzing emotions from the voice and facial image data provided by the user.
[1436] The "display confirmation means" is a means having a function of allowing the user to confirm the generated content.
[1437] "SNS posting means" refers to a means that has the function of allowing users to easily post content they have viewed to SNS.
[1438] This invention is a system that takes a photo of the cover of a book that a user has finished reading, extracts the book title, author name, and cover image using OCR technology, recognizes the user's emotions using an emotion engine, and obtains accurate bibliographic information by matching it with a book database. Furthermore, it analyzes the impressions and keywords entered by the user using emotion analysis, automatically generates a POP design based on the bibliographic information and the emotion analysis results, selects an appropriate song, and generates a short video by integrating the generated design with the selected song, which is then distributed to social media.
[1439] The system is programmed as follows:
[1440] A user uses their smartphone camera to take a photo of the cover of a book they have just read. The captured image is saved in the device app. The device compresses the image and uploads it to the server, along with the user ID and session information. The server then inputs the uploaded image into an OCR engine (e.g., Google Cloud Vision API) to extract the book title, author name, and visual elements from the image.
[1441] The server compares the book title and author name extracted by OCR with a book database (e.g., Google Books API, Open Library) to obtain accurate bibliographic information. If bibliographic information is not found, alternative title suggestions are presented and the user is asked for confirmation. The user then enters their thoughts and keywords about the book they read into the input form of the device app. The input is automatically sent to the server. The server uses an emotion engine (e.g., IBM Watson NLU) to analyze the thoughts and keywords entered by the user and recognize emotions. If the user provides facial images or voice, emotions can also be recognized from this data, if necessary.
[1442] The server generates POP text based on the results of sentiment analysis and bibliographic information. The generated text is structured to reflect the emotional emphasis. The server then runs an automated design engine (e.g., Adobe Creative Cloud API) to generate a POP image based on the bibliographic information, sentiment analysis results, and cover image. The POP image also incorporates key design elements from the cover.
[1443] The server then refers to the emotion analysis results and selects an appropriate song from a music database (e.g., Spotify API, Apple Music API). The song is selected to match the emotional tone. The server then combines the POP image with the selected song and generates a short video using a video generation tool such as FFmpeg. The generated video data is then sent to the device.
[1444] The user can view the short video sent from the server on their device. After viewing, the user presses the "Post" button on the device app to post the video to a social networking platform (e.g., Instagram, Twitter). The device then calls the upload API of the selected social networking platform and uploads the short video. Once posting is complete, a notification is displayed to the user.
[1445] For example, if a user photographs the cover of a "fantasy novel," the OCR engine extracts "fantasy novel" and the author's name from the image. It then compares the image with a book database to obtain accurate bibliographic information. If the user writes their impression, "This is a work filled with adventure and emotion," the emotion engine analyzes it and extracts the key essences of "adventure" and "emotion" using a sentiment analysis algorithm. From these results, the server creates a POP text such as "An inspiring story with adventurous elements."
[1446] The automated design engine then generates a POP image based on the text, the acquired bibliographic information, and the cover image. Furthermore, it selects an appropriate song from a music database based on the results of sentiment analysis. Finally, a short video is generated by integrating the POP image and the selected song, which users can easily post to social media.
[1447] The following are examples of prompt sentences:
[1448] "Please explain in natural language the process of a system that takes a photo of a book cover, extracts the book title and author name using OCR technology, analyzes the emotions felt after reading using an emotion engine, compares it with a book database to obtain bibliographic information, generates POP text and images based on the emotion analysis results, selects music, creates a video using a short video creation tool (e.g., FFmpeg), and posts it to social media."
[1449] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1450] Step 1: Capture and upload images
[1451] The user takes a photo of the book cover using the smartphone camera. The captured image is automatically saved in the device app. The device compresses the saved image and uploads it to the server. When uploading, the user ID and session information are also sent. This sends the compressed image data and user information to the server.
[1452] Step 2: OCR
[1453] The server inputs the uploaded image into an OCR engine (e.g., optical character recognition software). The OCR engine extracts the book title, author name, and visual elements from the image. It receives image data as input and obtains text data of the book title, author name, and visual elements as output. Specifically, the OCR engine identifies character regions in the image and reads character data from those regions.
[1454] Step 3: Obtain bibliographic information
[1455] The server uses the book title and author name extracted by OCR to check against a book database (e.g., a book information service). This allows accurate bibliographic information to be obtained. It receives text data of the book title and author name as input, and obtains detailed book information (e.g., publication year, genre, summary) as output. Specifically, it sends a request to the book database via an API and obtains the corresponding book information.
[1456] Step 4: Emotion Recognition
[1457] The user enters their thoughts and keywords about the book they have read into an input form in the device app. The input is automatically sent to the server. The server uses an emotion engine (e.g., natural language processing software) to analyze the thoughts and keywords entered by the user and recognize the emotion. It receives the text data of the thoughts and keywords as input and obtains the type of emotion (e.g., joy, sadness) as output. Specifically, the emotion engine analyzes the text and runs an algorithm to classify the emotion.
[1458] Step 5: Sentiment Analysis
[1459] The server analyzes the emotion results recognized by the emotion engine and the user's impressions, and extracts the main essence using an emotion analysis algorithm. It receives emotion data and impression text as input, and obtains the extracted essence (e.g., adventure, emotion) as output. Specifically, it uses a text analysis algorithm to extract important keywords and themes from the impressions.
[1460] Step 6: POP Text Generation
[1461] The server generates POP text based on the results of sentiment analysis and bibliographic information. It receives the sentiment essence and bibliographic information as input and obtains a catchy slogan and description as output. Specifically, it uses a template engine to compose the text and adjust it to reflect the emotional emphasis.
[1462] Step 7: Design Generation
[1463] The server runs an automatic design engine (e.g., design software API) to generate POP images based on bibliographic information, sentiment analysis results, and cover images. It receives design elements (e.g., book title, author name, emotional essence) as input and obtains POP image data as output. Specifically, it automatically adjusts the layout and coloring to generate visually appealing images.
[1464] Step 8: Music Selection
[1465] The server refers to the emotion analysis results and selects appropriate songs from a music database (e.g., a music streaming service). It receives the emotion essence as input and obtains song information as output. Specifically, it searches for songs that match the emotion and executes an algorithm to select the most suitable one.
[1466] Step 9: Short video generation
[1467] The server combines the POP image with the selected music and generates a short video using a video generation tool (e.g., FFmpeg). It receives POP image data and music information as input and obtains a short video file as output. Specifically, it adjusts the timing of the image and music and encodes them as continuous visual content.
[1468] Step 10: Social Media Distribution
[1469] The device receives the short video sent from the server and displays a confirmation screen to the user. The user checks the video and presses the "Post" button to post the video to the SNS platform (e.g., SNS service API). The device receives the short video file as input and receives a notification that posting to the SNS has been completed as output. Specifically, the device calls the SNS upload API and displays a notification to the user when posting is successful.
[1470] (Application example 2)
[1471] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1472] Traditionally, creating recommended book POPs in bookstores required a lot of time and effort, and the content of the POPs was not emotionally personalized, limiting their appeal to customers. Furthermore, updating promotional materials displayed on in-store displays was done manually, resulting in a lack of immediacy. For these reasons, a method was needed for bookstore staff to quickly and effectively promote books.
[1473] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an image analysis means, a bibliographic information acquisition means, a sentiment analysis means, an electronic design generation means, a music selection means, a short video generation means, a short video distribution means, and an automatic display means. This allows bookstore staff to simply take a photo of a book cover with their smartphone, and automatically generate an individual recommended POP based on related bibliographic information, user reviews, and sentiment analysis, and display it in real time on an electronic display in the store or post it to social media.
[1474] "Image analysis means" refers to means for extracting information from captured images.
[1475] "Bibliographic information acquisition means" refers to a means for acquiring information related to a book (e.g., book title, author name) from a database.
[1476] "Sentiment analysis means" means means for analyzing emotions from user input or other data.
[1477] "Electronic design generation means" refers to a means for automatically generating designs based on book information and sentiment analysis results.
[1478] The "music selection means" is a means for selecting appropriate music based on the emotion analysis results.
[1479] The "short video generation means" is a means for integrating a POP image with selected music to generate a short video.
[1480] "Short video distribution means" refers to a means for distributing the generated short video to an SNS platform.
[1481] "Automatic display means" refers to a means for displaying the generated POP images and short videos on electronic displays in the store in real time.
[1482] This system allows bookstore staff to take a photo of a book cover with their smartphone, automatically generating recommended POPs based on related bibliographic information, user reviews, and sentiment analysis, and posting them on in-store electronic displays and social media. The details of this system are described below.
[1483] The system starts by having the user take a photo of the book cover with their smartphone. The image is compressed and uploaded to a server, where it is analyzed using image analysis tools (e.g., pytesseract) to extract the book title, author, and visual elements.
[1484] Next, the server uses the bibliographic information acquisition means to acquire accurate bibliographic information from the book database, including the book title, author name, publication year, genre, etc.
[1485] The server then uses sentiment analysis tools to analyze the reviews and keywords entered by the user and recognize emotions. This process includes an emotion engine using natural language processing techniques. Facial images and voice data may also be used for emotion recognition.
[1486] Based on the results of the sentiment analysis and the bibliographic information, the server automatically generates POP text using an electronic design generation tool, and then creates a POP image using a design engine. This POP image reflects the bibliographic information and the user's sentiment.
[1487] Furthermore, the server uses a music selection means to select appropriate music from a music database based on the result of the emotion analysis, and the music is selected to match the emotional tone.
[1488] The server combines these POP images with the selected music and generates a short video using a short video generation tool (e.g., FFmpeg). This video is automatically generated and delivered to the user's smartphone.
[1489] Finally, the generated short video is posted to a social networking site using a short video distribution method. Users can select a social networking site (e.g., Instagram or Twitter) on their smartphone to complete the posting.
[1490] To accommodate in-store promotions, the system is equipped with an automatic display means, which displays the generated POP images on electronic displays in the store in real time.
[1491] Hardware and Software Use
[1492] The main hardware used in this invention is a smartphone and an electronic display in a bookstore, and the main software is an image analysis engine (pytesseract), a sentiment analysis engine, a design generation API, a video generation tool (FFmpeg), and a SNS upload API.
[1493] Specific examples
[1494] For example, if a user takes a photo of the cover of a "fantasy novel" and writes in their review that it is "a work filled with adventure and emotion," the system will analyze this information and generate a pop-up image that emphasizes "adventure" and "emotion." This image is then combined with a selected song to create a short video. The video can then be displayed on in-store displays and posted to social media.
[1495] Prompt Sentence Examples
[1496] "Users take a photo of the cover of a book they have read and enter their thoughts on the book. The system analyzes the information and generates a short video that combines recommended pop music and posts it to social media."
[1497] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1498] Step 1:
[1499] Image capture and upload
[1500] The user takes a photo of the book cover with their smartphone. This image is saved in the smartphone app. The device compresses the image file and uploads it to the server. The user ID and session information are also sent at the time of uploading.
[1501] Step 2:
[1502] OCR processing
[1503] The server inputs the received image into an OCR engine (pytesseract), which extracts text from the image and obtains information such as the book title and author name. The extracted text data is sent to the next processing step.
[1504] Step 3:
[1505] Bibliographic information acquisition
[1506] The server uses the book title and author name extracted by OCR to retrieve detailed bibliographic information from a book database. Using the book title and author name as input data, the server compares the book title, author name, publication year, genre, and other bibliographic information to output.
[1507] Step 4:
[1508] emotion recognition
[1509] The user enters their thoughts and keywords about the book they have read into an input form on a smartphone app. This text data is sent to a server. The server uses an emotion analysis tool (emotion engine) to analyze the thoughts and keywords entered by the user and recognize the emotion. The emotion analysis engine takes this data as input and outputs the type of emotion (e.g., joy, sadness).
[1510] Step 5:
[1511] Emotion analysis
[1512] The server analyzes the sentiment data, including the emotion recognition results, and uses a sentiment analysis algorithm to extract the key essence and generate the data needed to generate POPs. The input is the emotion recognition results and user sentiment data, and the output is the essence extraction results.
[1513] Step 6:
[1514] POP Text Generation
[1515] The server automatically generates POP text based on the results of sentiment analysis and bibliographic information. The generated text reflects the emotional emphasis. The input is the sentiment analysis results and bibliographic information, and the output is POP text.
[1516] Step 7:
[1517] Design Generation
[1518] The server runs an electronic design generator to automatically generate a POP image based on the sentiment analysis results, bibliographic information, and cover image. This POP image also incorporates the main design elements of the cover. The input is the POP text, bibliographic information, and cover image, and the output is the POP image.
[1519] Step 8:
[1520] Music Selection
[1521] The server refers to the results of the emotion analysis and selects an appropriate song from a music database. The selected song matches the emotional tone. The input is the emotion analysis result, and the output is the selected song.
[1522] Step 9:
[1523] Short video generation
[1524] The server combines the POP image and the selected music and generates a short video using a short video generation tool (FFmpeg). The input is the POP image and music URL, and the output is the short video data.
[1525] Step 10:
[1526] SNS distribution
[1527] The device receives the short video sent from the server and displays a confirmation screen to the user. The user reviews the video and is ready to post. When the user selects an SNS platform (e.g., Instagram, Twitter) and presses the "Post" button, the device calls the upload API of the selected SNS and uploads the short video. The input is the short video data and the SNS information selected by the user, and the output is a notification of successful posting to the SNS.
[1528] Step 11:
[1529] Automatic display
[1530] The server uses an automatic display means to display the generated POP image on an electronic display in the store in real time. The input is the POP image, and the output is the display on the in-store display.
[1531] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1532] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1533] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1534] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1535] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1536] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1537] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1538] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1539] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1540] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1541] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1542] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1543] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1544] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1545] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1546] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1547] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1548] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1549] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1550] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1551] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1552] The following is further disclosed regarding the above embodiment.
[1553] (Claim 1)
[1554] Image analysis means;
[1555] A bibliographic information acquisition means;
[1556] A sentiment analysis means;
[1557] an electronic design generation means;
[1558] A music selection means;
[1559] A short video generation means;
[1560] Short video distribution means,
[1561] A system including:
[1562] (Claim 2)
[1563] 10. The system of claim 1, wherein the image analysis means extracts book titles, author names, and visual elements from the images using optical character recognition techniques.
[1564] (Claim 3)
[1565] 2. The system according to claim 1, wherein the electronic design generating means automatically generates graphic data based on bibliographic information and sentiment analysis results.
[1566] "Example 1"
[1567] (Claim 1)
[1568] Image analysis means;
[1569] A bibliographic information acquisition means;
[1570] A sentiment analysis means;
[1571] an electronic design generation means;
[1572] A music selection means;
[1573] A short video generation means;
[1574] Short video distribution means,
[1575] a terminal for uploading images to a server via a user interface;
[1576] A terminal that accepts feedback input from users;
[1577] A means for collating the extracted book title and author name with a database to obtain bibliographic information;
[1578] a text generation means for generating a POP text based on the sentiment analysis result;
[1579] a means for integrating the image, the generated text, and the selected music piece in the short video generating means;
[1580] A system including:
[1581] (Claim 2)
[1582] 10. The system of claim 1, wherein the image analysis means extracts book titles, author names, and visual elements from the images using optical character recognition techniques.
[1583] (Claim 3)
[1584] 2. The system according to claim 1, wherein the electronic design generating means automatically generates graphic data based on bibliographic information and sentiment analysis results.
[1585] "Application Example 1"
[1586] (Claim 1)
[1587] Image analysis means;
[1588] A bibliographic information acquisition means;
[1589] A sentiment analysis means;
[1590] an electronic design generation means;
[1591] A music selection means;
[1592] A short video generation means;
[1593] Short video distribution means,
[1594] A means for taking pictures and inputting data from a user terminal;
[1595] A means for sharing and distributing the generated short video on a social networking service;
[1596] A system including:
[1597] (Claim 2)
[1598] 10. The system of claim 1, wherein the image analysis means extracts book titles, author names, and visual elements from the images using optical character recognition techniques.
[1599] (Claim 3)
[1600] 2. The system according to claim 1, wherein the electronic design generating means automatically generates graphic data based on bibliographic information and sentiment analysis results.
[1601] "Example 2: Combining Emotion Engines"
[1602] (Claim 1)
[1603] Image analysis means;
[1604] A bibliographic information acquisition means;
[1605] A sentiment analysis means;
[1606] an electronic design generation means;
[1607] A music selection means;
[1608] A short video generation means;
[1609] Short video distribution means,
[1610] Image compression and uploading means;
[1611] OCR analysis means;
[1612] User interaction means;
[1613] voice and face image analysis means;
[1614] A display confirmation means;
[1615] SNS posting methods and
[1616] A system including:
[1617] (Claim 2)
[1618] 10. The system of claim 1, wherein the image analysis means extracts book titles, author names, and visual elements from the images using optical character recognition techniques.
[1619] (Claim 3)
[1620] 2. The system according to claim 1, wherein the electronic design generating means automatically generates graphic data based on bibliographic information and sentiment analysis results.
[1621] "Application example 2 when combining emotion engines"
[1622] (Claim 1)
[1623] Image analysis means;
[1624] A bibliographic information acquisition means;
[1625] A sentiment analysis means;
[1626] an electronic design generation means;
[1627] A music selection means;
[1628] A short video generation means;
[1629] Short video distribution means,
[1630] automatic display means;
[1631] A system including:
[1632] (Claim 2)
[1633] 10. The system of claim 1, wherein the image analysis means extracts book titles, author names, and visual elements from the images using optical character recognition techniques.
[1634] (Claim 3)
[1635] 2. The system according to claim 1, wherein the electronic design generating means automatically generates graphic data based on bibliographic information and sentiment analysis results.
[1636] (Claim 4)
[1637] 2. The system according to claim 1, wherein the automatic display means displays the generated POP image on an electronic display in the bookstore in real time. [Explanation of symbols]
[1638] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. Image analysis means; A bibliographic information acquisition means; A sentiment analysis means; an electronic design generation means; A music selection means; A short video generation means; Short video distribution means, A system including:
2. 2. The system of claim 1, wherein the image analysis means extracts book titles, author names, and visual elements from the images using optical character recognition techniques.
3. 2. The system according to claim 1, wherein the electronic design generating means automatically generates graphic data based on bibliographic information and sentiment analysis results.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A