System

The system addresses the challenge of transforming images into text by allowing users to input images, analyze key elements, specify text style and keywords, and generate personalized text, enhancing writing skills and creative output.

JP2026019850APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121598
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Aspiring writers face difficulties in transforming vague images into concrete sentences, especially when describing people or scenery, and there is a lack of systems that allow for easy editing and personalization of generated text based on user preferences.

Method used

A system that includes image input, analysis to extract key elements, user-specified text style and keywords, and a generative AI to create personalized text, allowing editing and sharing on various platforms.

Benefits of technology

Enables users to easily generate novel-like text from images, improve writing skills, and personalize the text to their preferences, facilitating creative activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019850000001_ABST
    Figure 2026019850000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for inputting an image; means for analyzing the input image and extracting a main element in the image and a feature thereof; means for inputting a taste and a keyword of a sentence designated by a user; means for generating a sentence using the extracted element and feature of the image based on the taste and the keyword; and means for presenting the generated sentence to the user so that the sentence can be edited.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] For those aspiring to become writers, turning vague images into concrete sentences is a difficult task. This hurdle becomes even higher when the ability to effectively describe people or scenery is lacking. It is also difficult for beginners to generate sentences by specifying a taste or style. There is a need to solve this problem and lower the hurdle for aspiring writers. [Means for solving the problem]

[0005] The present invention solves the aforementioned problems by providing a system including a means for inputting an image, a means for analyzing the input image and extracting key elements and their characteristics, and a means for inputting a user-specified text taste and keywords, and a means for generating text using the extracted image elements and characteristics based on the taste and keywords. Furthermore, by including a means for presenting the generated text to the user and making it editable, the user can easily create text tailored to their preferences. Furthermore, by including a means for personalizing the generated text and tuning it according to the user's preferences, the accuracy of the text generation can be improved with repeated use. In this way, users can easily generate text based on images and improve their writing skills. Furthermore, the image analysis means includes a means for identifying key elements in the image, such as people, animals, and scenery, using image recognition technology, resulting in high analysis accuracy.

[0006] "Means for inputting images" refers to an interface or mechanism that allows a user to use a terminal to upload image files to a server.

[0007] "Means for analyzing images" refers to algorithms or modules that use image recognition technology to extract key elements and their features from input images.

[0008] "Key elements" refer to important parts of an image that are the subject of analysis, such as people, animals, and scenery.

[0009] "Means of extracting features" refers to the process of analyzing detailed information such as the color, shape, and position of each element in an image and extracting it as data.

[0010] "User-specified text style" refers to options that allow the user to select the style or atmosphere of the generated text. Examples include light novel style, horror novel style, fantasy novel style, etc.

[0011] "Means for entering keywords" refers to an interface that allows a user to enter specific words or themes that they want to include in their writing.

[0012] "Means for generating text" refers to algorithms or generation AI that automatically generate text based on the results of image analysis and the tastes and keywords specified by the user.

[0013] "Means for presenting and editing text" refers to an interface that allows the user to view the generated text and edit it as needed.

[0014] "Means for personalization and tuning" refers to a mechanism for adjusting the style and content of generated text to suit the user's preferences, thereby improving the accuracy of subsequent text generation.

[0015] "Image recognition technology" refers to technology for automatically recognizing and classifying specific objects or features from input images, including methods that use machine learning and deep learning. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention relates to a technology for generating novel-like text from images. This system analyzes images provided by the user and generates text based on the analysis results, according to the taste and keywords specified by the user. Such a system promotes the improvement of writing skills and lowers the barrier to becoming a writer. Furthermore, the generated text can be edited by the user, and a personalization function allows adjustments to suit the user's preferences.

[0038] A natural language description of the system's programmatic processing

[0039] 1. Upload an image

[0040] User:

[0041] Users can select any image file from their device and upload it to the server through a specified interface. For example, users can upload photos of their pets at home or scenery from their travels.

[0042] Device:

[0043] The terminal transmits the selected image file to the server as an HTTP request.

[0044] server:

[0045] The server temporarily stores the received image file for further processing.

[0046] 2. Select a style

[0047] User:

[0048] The interface allows users to select the style of their writing, such as "light novel style," "horror novel style," or "fantasy novel style," and also allows users to enter keywords and themes they want to include in their writing.

[0049] Device:

[0050] The terminal transmits the selected taste and keywords to the server as form data.

[0051] server:

[0052] The server stores the received tastes and keywords and uses them for the next process.

[0053] 3. Image Analysis

[0054] server:

[0055] The server analyzes the uploaded images using an image analysis module, which uses image recognition technology to identify key elements in the image (e.g., people, animals, landscapes) and extract their features.

[0056] 4. Sentence Generation

[0057] server:

[0058] The server provides the analyzed image elements and features, as well as the taste and keywords selected by the user, to the multimodal generative AI. The generative AI generates text that matches the specified taste based on the provided information.

[0059] For example, if you set the theme to fantasy and adventure, the following text will be generated:

[0060] "A young adventurer named Aria wakes up in a forest bathed in dazzling light. She deepens her friendships with the companions she meets along the way and renews her resolve to venture into the unknown."

[0061] 5. Check and edit the text

[0062] User:

[0063] The user can view the generated text in the interface and, if necessary, edit parts of the text to suit their preferences.

[0064] Device:

[0065] The terminal transmits the text edited by the user to the server.

[0066] server:

[0067] The server stores the user's edits to serve as a reference for future personalization and as feedback data to improve the quality of the generated text.

[0068] 6. Save and share your writing

[0069] User:

[0070] Users can use the save and share buttons on the interface to save their completed writing and share it across various platforms.

[0071] Device:

[0072] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[0073] server:

[0074] The server stores the final text in a database as needed, allowing the user to use it comfortably the next time they access the site.

[0075] This embodiment allows users to easily create novel-like texts based on images and edit them to their liking. The generated texts are extremely useful for users who aspire to be writers, helping them improve their writing skills and assist in creative activities.

[0076] The processing flow will be explained below.

[0077] Step 1:

[0078] User:

[0079] The user selects an image file from their device, clicks the upload button in the web interface or app, selects the image, and then presses the send button to upload it to the server.

[0080] Step 2:

[0081] Device:

[0082] The terminal transmits the image file selected by the user to the server as an HTTP request.

[0083] Step 3:

[0084] server:

[0085] The server temporarily stores the received image file and prepares it for processing by the image analysis module.

[0086] Step 4:

[0087] User:

[0088] The user selects the style of the generated text on the interface, such as "light novel style," "horror novel style," or "fantasy novel style," and also inputs keywords or themes (e.g., "adventure" or "friendship") they want to include in the generated text.

[0089] Step 5:

[0090] Device:

[0091] The terminal transmits the selected taste and keywords to the server as form data.

[0092] Step 6:

[0093] server:

[0094] The server stores the received tastes and keywords and launches an image analysis module, which analyzes the received image, identifies key elements in the image (e.g., people, animals, landscapes), and extracts their features.

[0095] Step 7:

[0096] server:

[0097] Next, the server uses the image analysis results and the tastes and keywords specified by the user to provide data to a multimodal generation AI, which then generates text that matches the specified taste based on the provided data.

[0098] Step 8:

[0099] server:

[0100] The generated text is temporarily saved and prepared for presentation to the user.

[0101] Step 9:

[0102] User:

[0103] The user can check the generated text on the interface and, if necessary, edit parts of the text on the interface to suit their preferences.

[0104] Step 10:

[0105] Device:

[0106] The terminal transmits the text edited by the user to the server.

[0107] Step 11:

[0108] server:

[0109] The server stores the edited text by the user and uses it for future personalization and as feedback data to improve the quality of the generated text.

[0110] Step 12:

[0111] User:

[0112] Users can click the save button on the interface to save the completed text, and can also click the share button to share the text on social media.

[0113] Step 13:

[0114] Device:

[0115] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[0116] Step 14:

[0117] server:

[0118] The server stores the final text in a database for future reference and reuse, and when the user returns, it provides personalized results based on the previously saved data.

[0119] Through these steps, users can easily create novel-like texts based on images, and then edit, save, and share them according to their preferences.

[0120] Example 1

[0121] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0122] Currently, technology for generating text from images is rapidly developing, but most systems are limited to simply generating text that describes the content of an image, and their ability to generate novel-like creative writing is limited. In particular, generating personalized text based on user-specified tastes and keywords is difficult and time-consuming. Furthermore, there is a lack of functionality that allows users to easily edit, save, and share the generated text. These challenges have resulted in a significant amount of effort required by users engaged in creative activities.

[0123] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0124] In this invention, the server includes means for inputting an image, means for analyzing the input image and extracting key elements and their features within the image, means for inputting text tastes and keywords specified by the user, means for using a multimodal generative model to generate text using the extracted image elements and features based on the tastes and keywords, means for presenting the generated text to the user and making it editable, and means for saving the generated text and sharing it on various platforms. This allows users to easily generate novel-like text based on their specified tastes and keywords, and to edit, save, and share it.

[0125] The "means for inputting images" is a device that includes an interface and a protocol for users to upload any image file from their own terminal to the server.

[0126] "Means for analyzing an input image and extracting key elements and their features" refers to a device that includes software modules and algorithms for using image recognition technology to identify key elements, such as people, animals, and scenery, in an image and extract their features.

[0127] "Means for inputting the style and keywords of the text specified by the user" refers to a device that allows the user to input the style of the text they desire (e.g., light novel style, horror novel style) and the keywords they want to include through an interface.

[0128] A "means for using a multimodal generative model" is a device that uses a generative AI model (e.g., a generative AI model) to process image and text data as input and generate personalized text based on specified tastes and keywords.

[0129] The "means for presenting the generated text to the user and enabling the user to edit it" is a device that includes an interface and software that displays the generated text to the user and allows the user to edit it.

[0130] "Means for saving generated text and sharing it on various platforms" refers to a device that includes an API and functionality for saving generated text and allowing users to download it or post it on social media or other sharing platforms.

[0131] This invention relates to a system that generates novel-like text from images. The system analyzes images provided by the user and generates text based on the analysis results according to the taste and keywords specified by the user. This promotes the improvement of writing skills and lowers the barrier to becoming a writer. The generated text can also be edited by the user, and a personalization function allows for adjustments to suit the user's preferences.

[0132] First, the user selects an image using their device and uploads it to the server through a specified interface. For example, the user might upload a photo of a pet at home or a landscape from a travel destination. This image file is sent to the server by the device as an HTTP request. The server temporarily stores the received image file for further processing. In this case, the server uses a storage solution such as AWS S3 or Google Cloud Storage.

[0133] Next, the user enters the style of the text (e.g., light novel style, horror novel style, fantasy novel style) and keywords (e.g., "adventure," "friendship," "hero") on the interface. The terminal sends the selected style and keywords as form data to the server, and the server stores the received data. For example, a database system such as MySQL or Firebase is used for this purpose.

[0134] The server analyzes the uploaded images using an image analysis module. Specifically, it uses image recognition technologies such as TensorFlow and OpenCV to identify key elements in the image (e.g., people, animals, landscapes) and extract their features. For example, it uses a TensorFlow image classification model to scan the image and identify key elements (e.g., dogs, mountains, rivers, etc.).

[0135] Next, the server provides the analyzed image elements and features, as well as the taste and keywords selected by the user, to a multimodal generative AI model (e.g., OpenAI's generative AI model).The generative AI model generates text that matches the specified taste based on the provided information.

[0136] For example, if you set the theme to fantasy and adventure, the following text will be generated:

[0137] "A young adventurer named Aria wakes up in a forest bathed in dazzling light. She deepens her friendships with the companions she meets along the way and renews her resolve to venture into the unknown."

[0138] The user can view the generated text on the interface and, if necessary, edit parts of the text to suit their preferences. The edited text is then sent from the device to the server, where it is stored. The stored data is used as a reference for future personalization and as feedback to improve the quality of the generative AI model.

[0139] Finally, the user can save the completed sentence and share it on various platforms. When the save button is pressed, the final sentence is downloaded to the user's device. When the SNS share button is pressed, the SNS sharing API is called and the sentence is posted. The server stores the final sentence in a database as needed, allowing the user to use it comfortably the next time they access the site.

[0140] An example of a specific prompt would be something like this:

[0141] Image: Upload your own image (e.g., a photo of your pet, a travel photo, etc.)

[0142] Style: Fantasy novel-style, adventure theme

[0143] Keywords: "Hero", "Friendship"

[0144] Generated text:

[0145] Waffles, a brave dog, embarks on a quest to save the magical kingdom with his best friend Claire. Together, they overcome many obstacles and rediscover the power of true friendship.

[0146] This system allows users to easily generate novel-like texts based on images and edit them to their liking. The generated texts are extremely useful for aspiring writers, helping them improve their writing skills and their creative activities.

[0147] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0148] Step 1:

[0149] User: The user selects an image file from their device and uploads it to the server through a specified interface. For example, the user may select a photo of a pet at home or a landscape from a travel destination.

[0150] Input: An image file selected by the user.

[0151] How it works: The user selects an image using the file selection dialog and clicks the upload button.

[0152] Output: The selected image file is sent to the next step.

[0153] Step 2:

[0154] Terminal: The terminal sends the selected image file to the server as an HTTP request.

[0155] Input: An image file selected by the user.

[0156] How it works: The device generates an HTML form, converts the selected image file to multipart / form-data format, and creates an HTTP POST request to send to the server.

[0157] Output: The image file is uploaded to the server.

[0158] Step 3:

[0159] Server: The server temporarily stores the received image files for further processing.

[0160] Input: Image file sent from the device.

[0161] How it works: The server parses the incoming HTTP request, extracts the image file, and saves it to storage (e.g., AWS S3 or Google Cloud Storage). After saving, it generates a data path for the next analysis step.

[0162] Output: The path of the image file saved in storage will be sent to the next step.

[0163] Step 4:

[0164] User: The user inputs the style of the text (e.g., light novel style, horror novel style, fantasy novel style) and keywords on the interface.

[0165] Input: Tastes and keywords entered by the user.

[0166] How it works: The user enters the desired taste and keywords using the drop-down menus and text fields.

[0167] Output: The input tastes and keywords are sent to the next step.

[0168] Step 5:

[0169] Terminal: The terminal sends the selected tastes and keywords to the server as form data.

[0170] Input: Tastes and keywords entered by the user.

[0171] Operation: The device converts the input data of tastes and keywords into JSON format and sends it to the server as an HTTP POST request.

[0172] Output: The taste and keyword information is sent to the server.

[0173] Step 6:

[0174] Server: The server stores the received tastes and keywords and uses them for the next process.

[0175] Input: Tastes and keywords entered by the user.

[0176] How it works: The server parses the received JSON data and stores it in a database (e.g., MySQL or Firebase). After saving, it proceeds to the next image analysis step.

[0177] Output: The saved taste and keyword information will be used in the next step.

[0178] Step 7:

[0179] Server: The server uses an image analysis module to analyze the uploaded images.

[0180] Input: The path of the image file retrieved from storage.

[0181] How it works: The server uses image recognition technology from TensorFlow and OpenCV to scan an image, identify key elements (e.g. people, animals, landscapes) and extract their features.

[0182] Output: Key elements and feature data in the image are sent to the next step.

[0183] Step 8:

[0184] Server: The server provides the analyzed image elements and features, as well as user-selected tastes and keywords, to the multimodal generative AI model as input.

[0185] Input: Extracted image elements and feature data, user-selected tastes and keywords.

[0186] How it works: The server inputs this data in JSON format into a generative AI model and sends a request to generate the appropriate sentence.

[0187] Output: The generated sentence is obtained as a response from the generative AI model.

[0188] Step 9:

[0189] User: The user checks the generated text on the interface and edits it as necessary.

[0190] Input: The generated sentence.

[0191] How it works: In the interface that displays the text, the user edits parts of the text as needed.

[0192] Output: The edited text data is sent to the next step.

[0193] Step 10:

[0194] Terminal: The terminal transmits the text edited by the user to the server.

[0195] Input: User edited text.

[0196] Operation: The device converts the edited text data into JSON format and sends it to the server as an HTTP POST request.

[0197] Output: The edited text data is sent to the server.

[0198] Step 11:

[0199] Server: The server stores the user's edits and uses them as a reference for future personalization.

[0200] Input: Edited text data.

[0201] How it works: The server parses the received JSON data and stores it in a database. It also uses it as feedback data to improve the quality of the generative AI model.

[0202] Output: Saved edited text data.

[0203] Step 12:

[0204] User: Users save their completed writing and share it on various platforms.

[0205] Input: Final sentence data.

[0206] How it works: Click the save or share button on the interface to share your text on social media or other platforms.

[0207] Output: The saved text is downloaded to the user's device or posted to a sharing platform.

[0208] Step 13:

[0209] Server: The server stores the final text in a database if necessary, making it easier for the user to use the text the next time they access the site.

[0210] Input: Final sentence data.

[0211] How it works: The server stores the final text data in a database and associates it with the user's profile.

[0212] Output: The saved text data is stored for future use.

[0213] (Application example 1)

[0214] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0215] While there have been technologies that generate text from images, there has been a lack of a way to save the text and easily share it across various platforms. It has also been difficult to personalize the generated text and tune it to the user's preferences. As a result, it has been difficult to provide a personalized creative experience for users.

[0216] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0217] In this invention, the server includes means for inputting an image, means for analyzing the input image and extracting key elements and their characteristics from the image, means for inputting text tastes and keywords specified by the user, means for generating text using the extracted image elements and characteristics based on the tastes and keywords, means for presenting the generated text to the user and making it editable, and means for saving the generated text and sharing it on various platforms. This allows users to easily save and share the generated text and provide a personalized creation experience tailored to their preferences.

[0218] "Means for inputting images" refers to a function that allows a user to select any image from the terminal and upload it to the server through a specified interface.

[0219] "Means for analyzing an image and extracting the main elements and their features within the image" refers to a function that uses image analysis technology to analyze an input image, identify the main elements within it, such as people, animals, and scenery, and extract their features.

[0220] "Means for inputting text style and keywords" is a function that allows the user to input the style and theme of the text to be generated on the interface, as well as the keywords that the user wants to include.

[0221] "Means for generating text using extracted image elements and features" refers to a function in which a multimodal generative AI generates text based on the elements and features obtained through image analysis and the tastes and keywords specified by the user.

[0222] The "means for presenting the generated text to the user and enabling it to be edited" is a function that provides the generated text on an interface so that the user can check it and edit it as necessary.

[0223] "Means for saving the generated text and sharing it on various platforms" refers to a function that allows users to save the final text and share it on platforms such as social media and email.

[0224] Overall system flow

[0225] Users operate the system using a smartphone. Specifically, they upload photographed or existing images to the application and input the text style and keywords. The system then analyzes the images and generates text using a generative AI model. The generated text can be reviewed and edited by the user, and can also be saved and shared.

[0226] Hardware and Software

[0227] Hardware

[0228] Smartphone: The device where the user selects and uploads images to the application.

[0229] Server: The central processing unit that analyzes images and generates text.

[0230] software

[0231] Image Analysis Module: Software for analyzing elements in images using TensorFlow and PyTorch.

[0232] Generative AI model: Using OpenAI's GPT-4 and other models, sentences are generated based on the analysis results and conditions specified by the user.

[0233] Front-end interface: The interface through which users upload images, select settings, and review and edit the generated text.

[0234] SNS sharing API: An API for sharing generated text on the SNS platform of the user's choice.

[0235] Data processing and calculation

[0236] When a user uploads an image from their smartphone, the server first temporarily stores the image. Next, an image analysis module is used to identify the main elements of the uploaded image (e.g., people, animals, landscapes) and extract their characteristics. Next, a generative AI model combines the analysis results with the text to generate a sentence based on the user-specified text style (e.g., "fantasy," "mystery," etc.).

[0237] Examples of concrete examples and prompts

[0238] For example, if a user uploads a photo of a beach at sunset, selects "fantasy" as the text style, and enters "adventure" and "magic" as keywords, the system will input the following prompt sentences into the generative AI model:

[0239] Generates the prompt statement:

[0240] The image shows a beach at sunset. Create a fantasy story using this image. Theme: Adventure, Magic.

[0241] Based on this prompt, the generative AI model might generate a sentence like this:

[0242] Examples of generated sentences:

[0243] "A golden sunset sunk into the beach. The wizard set off towards the boat waiting on the shore, eager for new adventures. The faint light from the beach shone on his magic wand."

[0244] The generated text can be viewed and edited on the interface, and can then be shared via social media or email. Through this series of operations, users can easily manage the generated text and enjoy creative activities.

[0245] In this way, the system of the present invention provides users with a comprehensive solution for generating, editing, and sharing novel-like text from images.

[0246] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0247] Step 1:

[0248] User: Selects an image from the smartphone's photo gallery and uploads it to the application.

[0249] Input: Image files on your smartphone

[0250] Output: Image file upload request to the server

[0251] Terminal: The selected image file is sent to the server as an HTTP request and temporarily stored.

[0252] Input: An image file selected by the user

[0253] Output: Image data transferred to the server

[0254] Server: Temporarily stores received image files and prepares them for the next processing step.

[0255] Input: Image data transferred from the device

[0256] Output: Temporarily saved image file

[0257] Step 2:

[0258] User: Select and input text style and keywords in the application interface.

[0259] Input: Text style (e.g., "fantasy"), keywords (e.g., "adventure", "magic")

[0260] Output: Input data of tastes and keywords

[0261] Device: Sends the selected tastes and keywords to the server.

[0262] Input: User input data of tastes and keywords

[0263] Output: Taste and keyword data sent to the server

[0264] Server: Save the received tastes and keywords and use them for the next process.

[0265] Input: Taste and keyword data sent from the device

[0266] Output: Saved tastes and keyword data

[0267] Step 3:

[0268] Server: Analyzes the uploaded image using an image analysis module. This analysis module uses image recognition techniques such as TensorFlow and PyTorch to identify key elements in the image and extract their features.

[0269] Input: Temporarily saved image file

[0270] Output: Identified key elements and their characteristic data

[0271] Step 4:

[0272] Server: Provides the analyzed image elements and features, as well as user-specified tastes and keywords, to the generative AI model. The generative AI model (e.g., OpenAI's GPT-4) generates prompts based on the provided information and creates full sentences.

[0273] Input: Analysis data (image elements and features), taste, keywords

[0274] Output: Generated sentence

[0275] Step 5:

[0276] Server: Sends the generated text to an interface that presents it to the user so that it can be edited by the user.

[0277] Input: Generated sentence

[0278] Output: Text displayed on the user's terminal

[0279] User: Check the generated text on the interface and edit it as necessary.

[0280] Input: Generated text sent from the server

[0281] Output: Edited text

[0282] Terminal: The user sends the edited results to the server.

[0283] Input: Text data edited by the user

[0284] Output: Edited text data sent to the server

[0285] Step 6:

[0286] Server: Saves edited text and connects to the sharing APIs of various platforms to enable sharing on social media and other platforms.

[0287] Input: Edited text data

[0288] Output: Final text data saved and text shared via the sharing API

[0289] User: Use the save and share buttons to share the generated text on various platforms.

[0290] Input: Final sentence

[0291] Output: Content shared across platforms

[0292] This series of steps creates a system that allows users to easily create, review, edit, save, and share text generated from images.

[0293] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0294] This invention combines technology for generating novel-like text from images with an emotion engine, enabling text generation that responds to the user's emotional state. This system analyzes images provided by the user and, based on the analysis results, generates and adjusts text according to the user's emotional state, incorporating tastes and keywords specified by the user. This system promotes the improvement of writing skills and lowers the barrier to becoming a writer. Furthermore, the generated text can be edited by the user, and a personalization function allows adjustments to the user's preferences.

[0295] A natural language description of the system's programmatic processing

[0296] 1. Upload an image

[0297] User:

[0298] Users can select any image file from their device and upload it to the server through a specified interface. For example, users can upload photos of their pets at home or scenery from their travels.

[0299] Device:

[0300] The terminal transmits the selected image file to the server as an HTTP request.

[0301] server:

[0302] The server temporarily stores the received image file for further processing.

[0303] 2. Select a style

[0304] User:

[0305] The interface allows users to select the style of their writing, such as "light novel style," "horror novel style," or "fantasy novel style," and also allows users to enter keywords and themes they want to include in their writing.

[0306] Device:

[0307] The terminal transmits the selected taste and keywords to the server as form data.

[0308] server:

[0309] The server stores the received tastes and keywords and launches an image analysis module, which analyzes the received image, identifies key elements in the image (e.g., people, animals, landscapes), and extracts their features.

[0310] 3. Emotional Recognition

[0311] User:

[0312] Users provide their emotional state to the system through input devices such as a webcam and microphone, which is collected through analysis of facial expressions, voice, and keyboard typing speed.

[0313] server:

[0314] The server analyzes the user's emotional state using an emotion engine, which identifies emotions such as joy, sadness, and anger based on biosignal data provided by the user.

[0315] 4. Sentence Generation

[0316] server:

[0317] The server provides data to a multimodal generation AI using the image analysis results, user-specified tastes and keywords, and the analyzed emotional state. Based on the information provided, the generation AI generates text that matches the user's emotional state in addition to the specified tastes.

[0318] For example, if the user selects "fantasy novel-style" and "adventure" as keywords, and the emotional state is recognized as "joy," the following sentence will be generated:

[0319] "In a meadow bathed in dazzling light, the little adventurer Leah stood up with a smile on her face. A fun-filled adventure with new friends awaited her."

[0320] 5. Check and edit the text

[0321] User:

[0322] The user can review the generated text on the interface and, if necessary, edit parts of the text to suit their preferences.

[0323] Device:

[0324] The terminal transmits the text edited by the user to the server.

[0325] server:

[0326] The server stores the text edited by the user and uses it as feedback data for future personalization and to improve the quality of the generated text.

[0327] 6. Save and share your writing

[0328] User:

[0329] Users can use the save and share buttons on the interface to save their completed writing and share it across various platforms.

[0330] Device:

[0331] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[0332] server:

[0333] The server stores the final text in a database for future reference and reuse, and when the user returns, it provides personalized results based on the previously saved data.

[0334] This embodiment allows users to easily create novel-like texts based on images, and then edit, save, and share them in a more personalized way that reflects their emotional state, further enriching the user's creative experience and helping to improve their writing skills and creative activities.

[0335] The processing flow will be explained below.

[0336] Step 1:

[0337] User:

[0338] Users simply select an image file from their device and click the upload button on the web interface or app. For example, a user can select a photo of a landscape from a travel destination or a pet, and then press the send button to upload it to the server.

[0339] Step 2:

[0340] Device:

[0341] The selected image file is sent to the server as an HTTP request.

[0342] Step 3:

[0343] server:

[0344] The received image file is temporarily saved and the image analysis module is prepared.

[0345] Step 4:

[0346] User:

[0347] The interface allows you to select the style of your writing, such as "light novel style," "horror novel style," or "fantasy novel style," and also enter keywords or themes you want to include in your writing (e.g., "adventure" or "friendship").

[0348] Step 5:

[0349] Device:

[0350] The selected tastes and keywords are sent to the server as form data.

[0351] Step 6:

[0352] server:

[0353] The received tastes and keywords are saved and the image analysis module is launched. The image analysis module analyzes the received image, identifies the main elements in the image (e.g., people, animals, landscapes), and extracts their features (color, shape, position, etc.).

[0354] Step 7:

[0355] User:

[0356] Through the interface, users provide their emotional state via input devices such as a webcam and microphone. Users input emotional data into the system through facial expressions, voice, and keyboard input speed.

[0357] Step 8:

[0358] server:

[0359] An emotion engine is activated to analyze the user's emotional state, and the emotion engine identifies emotions such as joy, sadness, anger, etc. based on the biosignal data provided by the user.

[0360] Step 9:

[0361] server:

[0362] The image analysis results, the user's emotional data, and the user's specified tastes and keywords are provided to a multimodal generation AI, which then generates text based on the given information, in addition to the specified tastes, that corresponds to the user's emotional state.

[0363] Step 10:

[0364] server:

[0365] The generated text is temporarily saved and prepared for presentation to the user.

[0366] Step 11:

[0367] User:

[0368] Review the generated text in the interface and, if necessary, edit parts of the text to suit your preferences.

[0369] Step 12:

[0370] Device:

[0371] The text edited by the user is sent to the server.

[0372] Step 13:

[0373] server:

[0374] User edits are saved and used for future personalization, and also as feedback data to improve the quality of generated text.

[0375] Step 14:

[0376] User:

[0377] To save your completed essay, click the Save button on the interface. You can also share your essay on social media by clicking the Share button.

[0378] Step 15:

[0379] Device:

[0380] When the save button is pressed, the final text is downloaded to the user's device. When the SNS share button is pressed, the SNS sharing API is called and the text is posted.

[0381] Step 16:

[0382] server:

[0383] The final text is stored in a database for future reference and reuse, and when the user returns, the previously stored data is used to provide personalized results.

[0384] Through these steps, users can easily generate novel-like texts based on images, and then edit, save, and share them with personalized content that reflects the user's emotional state. This enriches the user's creative experience, improves their writing skills, and helps with creative activities.

[0385] Example 2

[0386] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0387] While conventional text generation systems can generate text that takes into account user-specified tastes and keywords, they face the problem of difficulty in generating personalized text that reflects the user's emotional state. Furthermore, they lack the functionality to edit and adjust generated text to suit the user's preferences, limiting the user's creative experience. Furthermore, the complicated process of sharing and saving generated text reduces user convenience.

[0388] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0389] In this invention, the server includes: means for inputting an image; means for analyzing the input image and extracting key elements and their features within the image; means for inputting text tastes and keywords specified by the user; means for recognizing the user's emotional state; means for generating and adjusting text based on the recognized user's emotional state; means for presenting the generated text to the user and making it editable; means for analyzing biosignal data such as the user's facial expression, voice, and keyboard input speed; means for generating text based on the user's emotional state based on given information using a multimodal generative AI model; means for providing means for sharing the generated text on various platforms and for invoking a sharing API based on the user's instructions. This allows users to generate personalized text based on information extracted from the image and their own emotional state, enabling a richer creative experience through editing and sharing.

[0390] "Means for inputting images" is a function that allows a user to upload an image file selected from their own terminal to the system.

[0391] "Means for extracting the main elements and their features in an image" refers to a function that identifies the main elements, such as people, animals, and scenery, from the analyzed image and extracts these features as data.

[0392] The "means for inputting text style and keywords" is an input interface that allows the user to specify to the system the style of text to be generated and the content to be included.

[0393] The "means for recognizing the user's emotional state" is a function for analyzing biosignal data such as the user's facial expression, voice, keyboard input speed, etc., and identifying the user's current emotional state.

[0394] "Means for generating and adjusting sentences" refers to a function that uses specific algorithms and AI technology based on collected data to generate sentences that conform to specified tastes and keywords, and further adjusts them according to the user's emotional state.

[0395] The "means for presenting the generated text and making it editable" is an interface that displays the generated text to the user and allows the user to modify and edit parts of the text as necessary.

[0396] A "multimodal generative AI model" is an advanced artificial intelligence technology that integrates and processes information from multiple data sources (e.g., images, text, and emotional data) to generate natural language sentences.

[0397] The "means for analyzing biosignal data" is a function for analyzing biosignals such as facial expressions and voice obtained from the user and determining the emotional state from that data.

[0398] "Sharing API" means an application program interface used to share generated text across various platforms.

[0399] "Means for calling a sharing API based on user instructions" refers to a function that automatically calls the appropriate sharing API and posts text to a specified platform when a user performs an operation such as pressing a share button on the interface.

[0400] MODE FOR CARRYING OUT THE INVENTION

[0401] This system extracts key elements from an image provided by a user and generates text based on the user's preferences and keywords. It also has the ability to personalize, edit, save, and share text based on the user's emotional state. This system is implemented using the following hardware and software:

[0402] 1. Upload an image

[0403] User:

[0404] Users simply select an image file from their device and press the upload button through the system interface. For example, users can upload a photo of their pet at home or a landscape from a travel destination.

[0405] Device:

[0406] The device sends the selected image file as an HTTP request to the server, where the image data is temporarily stored in memory and sent to the server in the appropriate encoding format.

[0407] server:

[0408] The server temporarily stores the received image file in a secure storage location, and after storing it, assigns a unique identifier to the image file and uses that identifier for further processing.

[0409] 2. Select a style

[0410] User:

[0411] Users select the style of their writing (e.g., "fantasy novel-style") from the options provided on the interface, and then enter the keywords and themes they want to include in the writing.

[0412] Device:

[0413] The device sends the tastes and keywords selected by the user to the server as form data, where the input data is encoded in JSON format or similar.

[0414] server:

[0415] The server stores the received tastes and keywords in a database and then launches an image analysis module, which reads pre-saved image files, identifies key elements (e.g., people, animals, landscapes), and extracts their features.

[0416] 3. Emotional Recognition

[0417] User:

[0418] Users provide their emotional state to the system using a webcam or microphone, which is collected through facial expression recognition, voice tone analysis, keyboard input speed, etc.

[0419] server:

[0420] The server then analyzes the received biometric data using an emotion engine to identify specific emotions such as joy, sadness, anger, etc. The emotion analysis results are then used in the subsequent sentence generation process.

[0421] 4. Sentence Generation

[0422] server:

[0423] The server provides data to a multimodal generative AI model based on the results of image analysis, user-specified tastes and keywords, and the analyzed emotional state.

[0424] The generative AI model generates sentences based on the given information, in addition to the specified taste, according to the user's emotional state.

[0425] Examples:

[0426] If the user selects the keywords "fantasy novel-style" and "adventure" and the emotional state is recognized as "joy," the following sentence will be generated:

[0427] "In a meadow bathed in dazzling light, the little adventurer Leah stood up with a smile on her face. A fun-filled adventure with new friends awaited her."

[0428] 5. Check and edit the text

[0429] User:

[0430] The user can review the generated text in the interface and, if necessary, edit parts of the text to suit their preferences.

[0431] Device:

[0432] The terminal sends the text edited by the user to the server, again in JSON format or similar.

[0433] server:

[0434] The server stores the text edited by the user and uses it as feedback data for future personalization and system improvements.

[0435] 6. Save and share your writing

[0436] User:

[0437] Users can use the save and share buttons on the interface to save their completed writing and share it on various platforms.

[0438] Device:

[0439] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[0440] server:

[0441] The server stores the final text in a database for future reference and reuse, and when the user returns, it provides personalized results based on the previously saved data.

[0442] Prompt Sentence Examples

[0443] Below are some example prompts to input to a generative AI model:

[0444] "Right now, your emotional state is 'joy.' Create a fantasy-style adventure story based on a landscape photo. The keywords are 'light,' 'adventure,' and 'friends.'"

[0445] The system described above allows users to easily create novel-like texts based on images, and then edit, save, and share them with personalized content that reflects their emotional state. This will further enrich users' creative experiences, improve their writing skills, and help them with creative activities.

[0446] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0447] Step 1: Upload an image

[0448] input:

[0449] Users simply select any image file from their device and press the upload button through the system interface.

[0450] Specific behavior:

[0451] The user selects the image file of their choice through the specified interface and performs the upload operation, at which point the image file is temporarily stored on the user's device.

[0452] output:

[0453] The device sends the selected image file to the server as an HTTP request.

[0454] Step 2: Receive and save the image file

[0455] input:

[0456] Image files sent from the device

[0457] Specific behavior:

[0458] The server temporarily stores the received image file in storage and assigns a unique identifier (ID) to the file.

[0459] output:

[0460] The image file is given a unique identifier and stored in storage, which is used in the next processing step.

[0461] Step 3: Selecting Styles and Keywords

[0462] input:

[0463] The user specifies the style and keywords of the text.

[0464] Specific behavior:

[0465] The user selects the style of the text (e.g., "fantasy novel-style") from the options provided on the interface and enters keywords and themes.

[0466] output:

[0467] The device sends the selected taste and entered keywords as form data to the server. The sent data is encoded in JSON format.

[0468] Step 4: Begin image analysis

[0469] input:

[0470] Taste and keyword data, as well as identifiers of saved image files

[0471] Specific behavior:

[0472] The server receives the taste and keyword data and launches an image analysis module, which reads the saved image file, identifies key elements (e.g., people, animals, landscapes, etc.), and extracts their features.

[0473] output:

[0474] Data on the extracted image elements and their features are generated and used in the next processing step.

[0475] Step 5: Recognizing your emotional state

[0476] input:

[0477] Biosignal data collected from users (facial expressions, voice, keyboard typing speed, etc.)

[0478] Specific behavior:

[0479] Users provide their emotional state to the system using a webcam, microphone, etc. Emotional data is collected.

[0480] output:

[0481] The terminal transmits the collected biosignal data to the server.

[0482] Step 6: Analyze the sentiment data

[0483] input:

[0484] Collected biosignal data

[0485] Specific behavior:

[0486] The server analyzes the received biosignal data using an emotion engine to identify the user's emotional state (joy, sadness, anger, etc.).

[0487] output:

[0488] Parsed emotional state data is generated and used in the next processing step.

[0489] Step 7: Sentence generation

[0490] input:

[0491] Image analysis results, tastes and keywords, analyzed emotional states

[0492] Specific behavior:

[0493] The server provides the above data to a multimodal generative AI model, which uses the information to generate text that reflects the user's emotional state and the user's specified taste.

[0494] output:

[0495] The generated text is temporarily stored on the server.

[0496] Examples:

[0497] If the user sets the keywords "fantasy novel style" and "adventure" and the emotional state is recognized as "joy," the following sentence will be generated: "In a meadow bathed in dazzling light, the little adventurer Leah stood up with a smile on her face. A fun adventure with new friends awaited her."

[0498] Step 8: Presenting and editing the generated text

[0499] input:

[0500] Generated sentence data

[0501] Specific behavior:

[0502] The server sends the generated text to the interface, allowing the user to check the text content.

[0503] output:

[0504] The generated text is presented on the interface, and the user can view and edit it.

[0505] Step 9: Save your edits

[0506] input:

[0507] Text data edited by the user

[0508] Specific behavior:

[0509] The user can check the generated text on the interface, make corrections as necessary, and then click the save button when editing is complete.

[0510] output:

[0511] The terminal transmits the text data edited by the user to the server.

[0512] Step 10: Save edits and feedback

[0513] input:

[0514] Edited text data

[0515] Specific behavior:

[0516] The server stores the text data edited by the user and records it as feedback data to be used for future personalization and system improvement.

[0517] output:

[0518] The edited text data is stored in a database.

[0519] Step 11: Save and share your writing

[0520] input:

[0521] Completed sentence data

[0522] Specific behavior:

[0523] The user saves or shares the text using the save or share buttons on the interface.

[0524] output:

[0525] When the save button is pressed, the completed sentence is downloaded to the user's device. When the share button is pressed, the appropriate SNS sharing API is called and the sentence is posted.

[0526] Step 12: Save the final data

[0527] input:

[0528] Shared or stored text data

[0529] Specific behavior:

[0530] The server stores the final text in a database for future reference and reuse.

[0531] output:

[0532] The final text data is securely stored in a database and can be reused or referenced as needed.

[0533] In this way, this system allows users to easily create novel-like texts based on images, and then edit, save, and share them with personalized content that reflects their emotional state. This will further enrich the user's creative experience, improve their writing skills, and help with creative activities.

[0534] (Application example 2)

[0535] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0536] Conventional text generation systems generate text without considering the user's emotional state, which can result in text that does not match the user's emotions. This makes it difficult to generate text that the user can truly empathize with. Furthermore, conventional systems do not fully utilize the information in the image, limiting the quality of the generated text. This limits the user's creative experience.

[0537] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0538] In this invention, the server includes means for inputting an image, means for analyzing the input image and extracting the main elements and features of the image, means for inputting the taste and keywords of a sentence specified by the user, means for recognizing the analyzed emotional state of the user, means for generating a sentence using the extracted elements and features of the image based on the taste and keywords and the analyzed emotional state, and means for presenting the generated sentence to the user and making it editable. This makes it possible to generate sentences that correspond to the emotional state of the user and to generate high-quality sentences that the user can empathize with.

[0539] The "means for inputting images" refers to a method by which a user can provide a digital image to the system, and is a means by which image files can be uploaded using a smartphone or computer.

[0540] "Means for analyzing an input image and extracting the main elements and their features within the image" refers to a method that uses image recognition technology to identify the main elements of people, objects, landscapes, etc. that appear in an image and extract their features as data.

[0541] The "means for inputting the style and keywords of the text specified by the user" is an interface that allows the user to input the style of the text to be generated, the themes they wish to include, and keywords into the system.

[0542] "Means for recognizing an analyzed emotional state of a user" refers to a method that utilizes technology to identify a user's current emotional state through facial expression or voice analysis of the user.

[0543] "Means for generating sentences using extracted image elements and features based on the tastes and keywords and the analyzed emotional state" refers to a technology for generating sentences by incorporating tastes and keywords specified by the user and the recognized emotional state, and utilizing elements and features extracted from the image.

[0544] "Means for presenting the generated text to the user and making it editable" refers to an interface that displays the generated text to the user and allows the user to freely modify and edit the text.

[0545] The present invention relates to a system for generating sentences in accordance with the emotional state of a user, and a specific implementation method thereof will be described below.

[0546] Hardware and Software Configuration

[0547] This system includes an image input means, an image analysis means, a taste and keyword input means, an emotional state recognition means, a sentence generation means, and a means for making the generated sentences editable. The main hardware and software used are as follows:

[0548] Smartphones: iPhone, Android devices, etc.

[0549] Image recognition APIs: Google Cloud Vision, Amazon Rekognition.

[0550] Emotion analysis API: Microsoft Azure Emotion API, Affectiva.

[0551] Generative AI model: OpenAI GPT-4.

[0552] Database: Firebase, SQLite.

[0553] Processing flow

[0554] 1. Upload an image

[0555] Users select any image from their smartphone and use the app's interface to upload it to the server, which temporarily stores the received image for subsequent analysis.

[0556] 2. Selecting the theme and keywords

[0557] The user selects or inputs the text's style and keywords on the interface. For example, options such as "moving story," "pet," and "support" are displayed. This information is sent to the server as form data.

[0558] 3. Recognizing the user's emotional state

[0559] The user's real-time emotional state is analyzed through the smartphone's camera and microphone. An emotion analysis API is used to identify emotions such as joy, sadness, and anger from facial expressions and voice. This emotional state is then sent to the server.

[0560] 4. Sentence Generation

[0561] The server sends prompts to a generative AI model (OpenAI GPT-4) using the image analysis results, user-selected tastes and keywords, and the analyzed emotional state. The generative AI model generates high-quality novel-like text based on the given information.

[0562] For example, if a user uploads a picture of their pet, selects "Inspirational Stories," and the emotional state is recognized as "Joy," the prompt might look like this:

[0563] Image: A pet dog smiling and holding a ball

[0564] Taste: Inspirational story

[0565] Emotional state: Joy

[0566] Prompt: "This dog has supported me through some tough times. When he brought me the ball that day with a smile on his face..."

[0567] 5. Edit and save your text

[0568] The generated sentences are presented to the user on the interface. The user can review the generated sentences and edit them as necessary. The edited sentences are sent to the server for storage and used as feedback for future personalized generation results.

[0569] This allows users to generate, edit and save personalized text based on images and emotional states, enhancing their creativity.

[0570] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0571] Step 1:

[0572] Uploading an image

[0573] A user selects an image file from their smartphone and uploads it to the server through a specified interface. The device sends the selected image file to the server as an HTTP request, and the server temporarily stores the received image file. The input here is the image file selected by the user, and the output is the image file temporarily stored on the server.

[0574] Step 2:

[0575] Selecting styles and keywords

[0576] The user inputs the taste of the text (e.g., "moving story," "suspense," etc.) and keywords (e.g., "pet," "adventure," etc.) on the interface. The terminal sends the selected taste and keywords to the server as form data, which the server receives and stores. The input here is the taste and keywords selected by the user, and the output is the taste and keyword data stored on the server.

[0577] Step 3:

[0578] Recognizing the user's emotional state

[0579] The user's real-time emotional state is analyzed through the smartphone's camera and microphone. The device sends the acquired biosignal data to an emotion analysis API, which identifies the emotional state (e.g., joy, sadness, anger). The server receives and stores this emotional state data. The input here is the user's biosignal data (images, voice, etc.), and the output is analyzed emotional state data.

[0580] Step 4:

[0581] Image analysis

[0582] The server uses an image recognition API (e.g., Google Cloud Vision, Amazon Rekognition) to analyze the image uploaded by the user. The main elements in the image (e.g., people, animals, landscapes) and their features are extracted and saved as text data. The input here is the uploaded image file, and the output is the text data of the extracted main elements and features.

[0583] Step 5:

[0584] Sentence generation

[0585] The server sends prompts to a generative AI model (e.g., OpenAI GPT-4) based on the image analysis results, the tastes and keywords selected by the user, and the analyzed emotional state, to generate a sentence. The generated sentence is stored on the server. The inputs here are the image analysis results, tastes and keywords, and the emotional state, and the output is the text data of the generated sentence.

[0586] Step 6:

[0587] Editing and saving the generated text

[0588] The user can check the generated text on the interface and edit it as needed. The terminal sends the edited text to the server, which stores it. The input is the generated text and the user's edits, and the output is the final edited text that has been saved.

[0589] Step 7:

[0590] Sharing and saving text

[0591] Users use the save and share buttons on the interface to save their completed writing and share it on various platforms. When the save button is pressed, the device downloads the final writing to the user's device, and when the SNS share button is pressed, it calls the SNS's sharing API to post the writing. The server saves the final writing in a database for future reference and reuse. The input here is the user's save and share instructions, and the output is the saved writing and shared content.

[0592] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0593] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0594] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0595] [Second embodiment]

[0596] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0597] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0598] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0599] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0600] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0601] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0602] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0603] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0604] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0605] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0606] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0607] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0608] This invention relates to a technology for generating novel-like text from images. This system analyzes images provided by the user and generates text based on the analysis results, according to the taste and keywords specified by the user. Such a system promotes the improvement of writing skills and lowers the barrier to becoming a writer. Furthermore, the generated text can be edited by the user, and a personalization function allows adjustments to suit the user's preferences.

[0609] A natural language description of the system's programmatic processing

[0610] 1. Upload an image

[0611] User:

[0612] Users can select any image file from their device and upload it to the server through a specified interface. For example, users can upload photos of their pets at home or scenery from their travels.

[0613] Device:

[0614] The terminal transmits the selected image file to the server as an HTTP request.

[0615] server:

[0616] The server temporarily stores the received image file for further processing.

[0617] 2. Select a style

[0618] User:

[0619] The interface allows users to select the style of their writing, such as "light novel style," "horror novel style," or "fantasy novel style," and also allows users to enter keywords and themes they want to include in their writing.

[0620] Device:

[0621] The terminal transmits the selected taste and keywords to the server as form data.

[0622] server:

[0623] The server stores the received tastes and keywords and uses them for the next process.

[0624] 3. Image Analysis

[0625] server:

[0626] The server analyzes the uploaded images using an image analysis module, which uses image recognition technology to identify key elements in the image (e.g., people, animals, landscapes) and extract their features.

[0627] 4. Sentence Generation

[0628] server:

[0629] The server provides the analyzed image elements and features, as well as the taste and keywords selected by the user, to the multimodal generative AI. The generative AI generates text that matches the specified taste based on the provided information.

[0630] For example, if you set the theme to fantasy and adventure, the following text will be generated:

[0631] "A young adventurer named Aria wakes up in a forest bathed in dazzling light. She deepens her friendships with the companions she meets along the way and renews her resolve to venture into the unknown."

[0632] 5. Check and edit the text

[0633] User:

[0634] The user can view the generated text in the interface and, if necessary, edit parts of the text to suit their preferences.

[0635] Device:

[0636] The terminal transmits the text edited by the user to the server.

[0637] server:

[0638] The server stores the user's edits to serve as a reference for future personalization and as feedback data to improve the quality of the generated text.

[0639] 6. Save and share your writing

[0640] User:

[0641] Users can use the save and share buttons on the interface to save their completed writing and share it across various platforms.

[0642] Device:

[0643] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[0644] server:

[0645] The server stores the final text in a database as needed, allowing the user to use it comfortably the next time they access the site.

[0646] This embodiment allows users to easily create novel-like texts based on images and edit them to their liking. The generated texts are extremely useful for users who aspire to be writers, helping them improve their writing skills and assist in creative activities.

[0647] The processing flow will be explained below.

[0648] Step 1:

[0649] User:

[0650] The user selects an image file from their device, clicks the upload button in the web interface or app, selects the image, and then presses the send button to upload it to the server.

[0651] Step 2:

[0652] Device:

[0653] The terminal transmits the image file selected by the user to the server as an HTTP request.

[0654] Step 3:

[0655] server:

[0656] The server temporarily stores the received image file and prepares it for processing by the image analysis module.

[0657] Step 4:

[0658] User:

[0659] The user selects the style of the generated text on the interface, such as "light novel style," "horror novel style," or "fantasy novel style," and also inputs keywords or themes (e.g., "adventure" or "friendship") they want to include in the generated text.

[0660] Step 5:

[0661] Device:

[0662] The terminal transmits the selected taste and keywords to the server as form data.

[0663] Step 6:

[0664] server:

[0665] The server stores the received tastes and keywords and launches an image analysis module, which analyzes the received image, identifies key elements in the image (e.g., people, animals, landscapes), and extracts their features.

[0666] Step 7:

[0667] server:

[0668] Next, the server uses the image analysis results and the tastes and keywords specified by the user to provide data to a multimodal generation AI, which then generates text that matches the specified taste based on the provided data.

[0669] Step 8:

[0670] server:

[0671] The generated text is temporarily saved and prepared for presentation to the user.

[0672] Step 9:

[0673] User:

[0674] The user can check the generated text on the interface and, if necessary, edit parts of the text on the interface to suit their preferences.

[0675] Step 10:

[0676] Device:

[0677] The terminal transmits the text edited by the user to the server.

[0678] Step 11:

[0679] server:

[0680] The server stores the edited text by the user and uses it for future personalization and as feedback data to improve the quality of the generated text.

[0681] Step 12:

[0682] User:

[0683] Users can click the save button on the interface to save the completed text, and can also click the share button to share the text on social media.

[0684] Step 13:

[0685] Device:

[0686] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[0687] Step 14:

[0688] server:

[0689] The server stores the final text in a database for future reference and reuse, and when the user returns, it provides personalized results based on the previously saved data.

[0690] Through these steps, users can easily create novel-like texts based on images, and then edit, save, and share them according to their preferences.

[0691] Example 1

[0692] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0693] Currently, technology for generating text from images is rapidly developing, but most systems are limited to simply generating text that describes the content of an image, and their ability to generate novel-like creative writing is limited. In particular, generating personalized text based on user-specified tastes and keywords is difficult and time-consuming. Furthermore, there is a lack of functionality that allows users to easily edit, save, and share the generated text. These challenges have resulted in a significant amount of effort required by users engaged in creative activities.

[0694] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0695] In this invention, the server includes means for inputting an image, means for analyzing the input image and extracting key elements and their features within the image, means for inputting text tastes and keywords specified by the user, means for using a multimodal generative model to generate text using the extracted image elements and features based on the tastes and keywords, means for presenting the generated text to the user and making it editable, and means for saving the generated text and sharing it on various platforms. This allows users to easily generate novel-like text based on their specified tastes and keywords, and to edit, save, and share it.

[0696] The "means for inputting images" is a device that includes an interface and a protocol for users to upload any image file from their own terminal to the server.

[0697] "Means for analyzing an input image and extracting key elements and their features" refers to a device that includes software modules and algorithms for using image recognition technology to identify key elements, such as people, animals, and scenery, in an image and extract their features.

[0698] "Means for inputting the style and keywords of the text specified by the user" refers to a device that allows the user to input the style of the text they desire (e.g., light novel style, horror novel style) and the keywords they want to include through an interface.

[0699] A "means for using a multimodal generative model" is a device that uses a generative AI model (e.g., a generative AI model) to process image and text data as input and generate personalized text based on specified tastes and keywords.

[0700] The "means for presenting the generated text to the user and enabling the user to edit it" is a device that includes an interface and software that displays the generated text to the user and allows the user to edit it.

[0701] "Means for saving generated text and sharing it on various platforms" refers to a device that includes an API and functionality for saving generated text and allowing users to download it or post it on social media or other sharing platforms.

[0702] This invention relates to a system that generates novel-like text from images. The system analyzes images provided by the user and generates text based on the analysis results according to the taste and keywords specified by the user. This promotes the improvement of writing skills and lowers the barrier to becoming a writer. The generated text can also be edited by the user, and a personalization function allows for adjustments to suit the user's preferences.

[0703] First, the user selects an image using their device and uploads it to the server through a specified interface. For example, the user might upload a photo of a pet at home or a landscape from a travel destination. This image file is sent to the server by the device as an HTTP request. The server temporarily stores the received image file for further processing. In this case, the server uses a storage solution such as AWS S3 or Google Cloud Storage.

[0704] Next, the user enters the style of the text (e.g., light novel style, horror novel style, fantasy novel style) and keywords (e.g., "adventure," "friendship," "hero") on the interface. The terminal sends the selected style and keywords as form data to the server, and the server stores the received data. For example, a database system such as MySQL or Firebase is used for this purpose.

[0705] The server analyzes the uploaded images using an image analysis module. Specifically, it uses image recognition technologies such as TensorFlow and OpenCV to identify key elements in the image (e.g., people, animals, landscapes) and extract their features. For example, it uses a TensorFlow image classification model to scan the image and identify key elements (e.g., dogs, mountains, rivers, etc.).

[0706] Next, the server provides the analyzed image elements and features, as well as the taste and keywords selected by the user, to a multimodal generative AI model (e.g., OpenAI's generative AI model).The generative AI model generates text that matches the specified taste based on the provided information.

[0707] For example, if you set the theme to fantasy and adventure, the following text will be generated:

[0708] "A young adventurer named Aria wakes up in a forest bathed in dazzling light. She deepens her friendships with the companions she meets along the way and renews her resolve to venture into the unknown."

[0709] The user can view the generated text on the interface and, if necessary, edit parts of the text to suit their preferences. The edited text is then sent from the device to the server, where it is stored. The stored data is used as a reference for future personalization and as feedback to improve the quality of the generative AI model.

[0710] Finally, the user can save the completed sentence and share it on various platforms. When the save button is pressed, the final sentence is downloaded to the user's device. When the SNS share button is pressed, the SNS sharing API is called and the sentence is posted. The server stores the final sentence in a database as needed, allowing the user to use it comfortably the next time they access the site.

[0711] An example of a specific prompt would be something like this:

[0712] Image: Upload your own image (e.g., a photo of your pet, a travel photo, etc.)

[0713] Style: Fantasy novel-style, adventure theme

[0714] Keywords: "Hero", "Friendship"

[0715] Generated text:

[0716] Waffles, a brave dog, embarks on a quest to save the magical kingdom with his best friend Claire. Together, they overcome many obstacles and rediscover the power of true friendship.

[0717] This system allows users to easily generate novel-like texts based on images and edit them to their liking. The generated texts are extremely useful for aspiring writers, helping them improve their writing skills and their creative activities.

[0718] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0719] Step 1:

[0720] User: The user selects an image file from their device and uploads it to the server through a specified interface. For example, the user may select a photo of a pet at home or a landscape from a travel destination.

[0721] Input: An image file selected by the user.

[0722] How it works: The user selects an image using the file selection dialog and clicks the upload button.

[0723] Output: The selected image file is sent to the next step.

[0724] Step 2:

[0725] Terminal: The terminal sends the selected image file to the server as an HTTP request.

[0726] Input: An image file selected by the user.

[0727] How it works: The device generates an HTML form, converts the selected image file to multipart / form-data format, and creates an HTTP POST request to send to the server.

[0728] Output: The image file is uploaded to the server.

[0729] Step 3:

[0730] Server: The server temporarily stores the received image files for further processing.

[0731] Input: Image file sent from the device.

[0732] How it works: The server parses the incoming HTTP request, extracts the image file, and saves it to storage (e.g., AWS S3 or Google Cloud Storage). After saving, it generates a data path for the next analysis step.

[0733] Output: The path of the image file saved in storage will be sent to the next step.

[0734] Step 4:

[0735] User: The user inputs the style of the text (e.g., light novel style, horror novel style, fantasy novel style) and keywords on the interface.

[0736] Input: Tastes and keywords entered by the user.

[0737] How it works: The user enters the desired taste and keywords using the drop-down menus and text fields.

[0738] Output: The input tastes and keywords are sent to the next step.

[0739] Step 5:

[0740] Terminal: The terminal sends the selected tastes and keywords to the server as form data.

[0741] Input: Tastes and keywords entered by the user.

[0742] Operation: The device converts the input data of tastes and keywords into JSON format and sends it to the server as an HTTP POST request.

[0743] Output: The taste and keyword information is sent to the server.

[0744] Step 6:

[0745] Server: The server stores the received tastes and keywords and uses them for the next process.

[0746] Input: Tastes and keywords entered by the user.

[0747] How it works: The server parses the received JSON data and stores it in a database (e.g., MySQL or Firebase). After saving, it proceeds to the next image analysis step.

[0748] Output: The saved taste and keyword information will be used in the next step.

[0749] Step 7:

[0750] Server: The server uses an image analysis module to analyze the uploaded images.

[0751] Input: The path of the image file retrieved from storage.

[0752] How it works: The server uses image recognition technology from TensorFlow and OpenCV to scan an image, identify key elements (e.g. people, animals, landscapes) and extract their features.

[0753] Output: Key elements and feature data in the image are sent to the next step.

[0754] Step 8:

[0755] Server: The server provides the analyzed image elements and features, as well as user-selected tastes and keywords, to the multimodal generative AI model as input.

[0756] Input: Extracted image elements and feature data, user-selected tastes and keywords.

[0757] How it works: The server inputs this data in JSON format into a generative AI model and sends a request to generate the appropriate sentence.

[0758] Output: The generated sentence is obtained as a response from the generative AI model.

[0759] Step 9:

[0760] User: The user checks the generated text on the interface and edits it as necessary.

[0761] Input: The generated sentence.

[0762] How it works: In the interface that displays the text, the user edits parts of the text as needed.

[0763] Output: The edited text data is sent to the next step.

[0764] Step 10:

[0765] Terminal: The terminal transmits the text edited by the user to the server.

[0766] Input: User edited text.

[0767] Operation: The device converts the edited text data into JSON format and sends it to the server as an HTTP POST request.

[0768] Output: The edited text data is sent to the server.

[0769] Step 11:

[0770] Server: The server stores the user's edits and uses them as a reference for future personalization.

[0771] Input: Edited text data.

[0772] How it works: The server parses the received JSON data and stores it in a database. It also uses it as feedback data to improve the quality of the generative AI model.

[0773] Output: Saved edited text data.

[0774] Step 12:

[0775] User: Users save their completed writing and share it on various platforms.

[0776] Input: Final sentence data.

[0777] How it works: Click the save or share button on the interface to share your text on social media or other platforms.

[0778] Output: The saved text is downloaded to the user's device or posted to a sharing platform.

[0779] Step 13:

[0780] Server: The server stores the final text in a database if necessary, making it easier for the user to use the text the next time they access the site.

[0781] Input: Final sentence data.

[0782] How it works: The server stores the final text data in a database and associates it with the user's profile.

[0783] Output: The saved text data is stored for future use.

[0784] (Application example 1)

[0785] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0786] While there have been technologies that generate text from images, there has been a lack of a way to save the text and easily share it across various platforms. It has also been difficult to personalize the generated text and tune it to the user's preferences. As a result, it has been difficult to provide a personalized creative experience for users.

[0787] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0788] In this invention, the server includes means for inputting an image, means for analyzing the input image and extracting key elements and their characteristics from the image, means for inputting text tastes and keywords specified by the user, means for generating text using the extracted image elements and characteristics based on the tastes and keywords, means for presenting the generated text to the user and making it editable, and means for saving the generated text and sharing it on various platforms. This allows users to easily save and share the generated text and provide a personalized creation experience tailored to their preferences.

[0789] "Means for inputting images" refers to a function that allows a user to select any image from the terminal and upload it to the server through a specified interface.

[0790] "Means for analyzing an image and extracting the main elements and their features within the image" refers to a function that uses image analysis technology to analyze an input image, identify the main elements within it, such as people, animals, and scenery, and extract their features.

[0791] "Means for inputting text style and keywords" is a function that allows the user to input the style and theme of the text to be generated on the interface, as well as the keywords that the user wants to include.

[0792] "Means for generating text using extracted image elements and features" refers to a function in which a multimodal generative AI generates text based on the elements and features obtained through image analysis and the tastes and keywords specified by the user.

[0793] The "means for presenting the generated text to the user and enabling it to be edited" is a function that provides the generated text on an interface so that the user can check it and edit it as necessary.

[0794] "Means for saving the generated text and sharing it on various platforms" refers to a function that allows users to save the final text and share it on platforms such as social media and email.

[0795] Overall system flow

[0796] Users operate the system using a smartphone. Specifically, they upload photographed or existing images to the application and input the text style and keywords. The system then analyzes the images and generates text using a generative AI model. The generated text can be reviewed and edited by the user, and can also be saved and shared.

[0797] Hardware and Software

[0798] Hardware

[0799] Smartphone: The device where the user selects and uploads images to the application.

[0800] Server: The central processing unit that analyzes images and generates text.

[0801] software

[0802] Image Analysis Module: Software for analyzing elements in images using TensorFlow and PyTorch.

[0803] Generative AI model: Using OpenAI's GPT-4 and other models, sentences are generated based on the analysis results and conditions specified by the user.

[0804] Front-end interface: The interface through which users upload images, select settings, and review and edit the generated text.

[0805] SNS sharing API: An API for sharing generated text on the SNS platform of the user's choice.

[0806] Data processing and calculation

[0807] When a user uploads an image from their smartphone, the server first temporarily stores the image. Next, an image analysis module is used to identify the main elements of the uploaded image (e.g., people, animals, landscapes) and extract their characteristics. Next, a generative AI model combines the analysis results with the text to generate a sentence based on the user-specified text style (e.g., "fantasy," "mystery," etc.).

[0808] Examples of concrete examples and prompts

[0809] For example, if a user uploads a photo of a beach at sunset, selects "fantasy" as the text style, and enters "adventure" and "magic" as keywords, the system will input the following prompt sentences into the generative AI model:

[0810] Generates the prompt statement:

[0811] The image shows a beach at sunset. Create a fantasy story using this image. Theme: Adventure, Magic.

[0812] Based on this prompt, the generative AI model might generate a sentence like this:

[0813] Examples of generated sentences:

[0814] "A golden sunset sunk into the beach. The wizard set off towards the boat waiting on the shore, eager for new adventures. The faint light from the beach shone on his magic wand."

[0815] The generated text can be viewed and edited on the interface, and can then be shared via social media or email. Through this series of operations, users can easily manage the generated text and enjoy creative activities.

[0816] In this way, the system of the present invention provides users with a comprehensive solution for generating, editing, and sharing novel-like text from images.

[0817] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0818] Step 1:

[0819] User: Selects an image from the smartphone's photo gallery and uploads it to the application.

[0820] Input: Image files on your smartphone

[0821] Output: Image file upload request to the server

[0822] Terminal: The selected image file is sent to the server as an HTTP request and temporarily stored.

[0823] Input: An image file selected by the user

[0824] Output: Image data transferred to the server

[0825] Server: Temporarily stores received image files and prepares them for the next processing step.

[0826] Input: Image data transferred from the device

[0827] Output: Temporarily saved image file

[0828] Step 2:

[0829] User: Select and input text style and keywords in the application interface.

[0830] Input: Text style (e.g., "fantasy"), keywords (e.g., "adventure", "magic")

[0831] Output: Input data of tastes and keywords

[0832] Device: Sends the selected tastes and keywords to the server.

[0833] Input: User input data of tastes and keywords

[0834] Output: Taste and keyword data sent to the server

[0835] Server: Save the received tastes and keywords and use them for the next process.

[0836] Input: Taste and keyword data sent from the device

[0837] Output: Saved tastes and keyword data

[0838] Step 3:

[0839] Server: Analyzes the uploaded image using an image analysis module. This analysis module uses image recognition techniques such as TensorFlow and PyTorch to identify key elements in the image and extract their features.

[0840] Input: Temporarily saved image file

[0841] Output: Identified key elements and their characteristic data

[0842] Step 4:

[0843] Server: Provides the analyzed image elements and features, as well as user-specified tastes and keywords, to the generative AI model. The generative AI model (e.g., OpenAI's GPT-4) generates prompts based on the provided information and creates full sentences.

[0844] Input: Analysis data (image elements and features), taste, keywords

[0845] Output: Generated sentence

[0846] Step 5:

[0847] Server: Sends the generated text to an interface that presents it to the user so that it can be edited by the user.

[0848] Input: Generated sentence

[0849] Output: Text displayed on the user's terminal

[0850] User: Check the generated text on the interface and edit it as necessary.

[0851] Input: Generated text sent from the server

[0852] Output: Edited text

[0853] Terminal: The user sends the edited results to the server.

[0854] Input: Text data edited by the user

[0855] Output: Edited text data sent to the server

[0856] Step 6:

[0857] Server: Saves edited text and connects to the sharing APIs of various platforms to enable sharing on social media and other platforms.

[0858] Input: Edited text data

[0859] Output: Final text data saved and text shared via the sharing API

[0860] User: Use the save and share buttons to share the generated text on various platforms.

[0861] Input: Final sentence

[0862] Output: Content shared across platforms

[0863] This series of steps creates a system that allows users to easily create, review, edit, save, and share text generated from images.

[0864] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0865] This invention combines technology for generating novel-like text from images with an emotion engine, enabling text generation that responds to the user's emotional state. This system analyzes images provided by the user and, based on the analysis results, generates and adjusts text according to the user's emotional state, incorporating tastes and keywords specified by the user. This system promotes the improvement of writing skills and lowers the barrier to becoming a writer. Furthermore, the generated text can be edited by the user, and a personalization function allows adjustments to the user's preferences.

[0866] A natural language description of the system's programmatic processing

[0867] 1. Upload an image

[0868] User:

[0869] Users can select any image file from their device and upload it to the server through a specified interface. For example, users can upload photos of their pets at home or scenery from their travels.

[0870] Device:

[0871] The terminal transmits the selected image file to the server as an HTTP request.

[0872] server:

[0873] The server temporarily stores the received image file for further processing.

[0874] 2. Select a style

[0875] User:

[0876] The interface allows users to select the style of their writing, such as "light novel style," "horror novel style," or "fantasy novel style," and also allows users to enter keywords and themes they want to include in their writing.

[0877] Device:

[0878] The terminal transmits the selected taste and keywords to the server as form data.

[0879] server:

[0880] The server stores the received tastes and keywords and launches an image analysis module, which analyzes the received image, identifies key elements in the image (e.g., people, animals, landscapes), and extracts their features.

[0881] 3. Emotional Recognition

[0882] User:

[0883] Users provide their emotional state to the system through input devices such as a webcam and microphone, which is collected through analysis of facial expressions, voice, and keyboard typing speed.

[0884] server:

[0885] The server analyzes the user's emotional state using an emotion engine, which identifies emotions such as joy, sadness, and anger based on biosignal data provided by the user.

[0886] 4. Sentence Generation

[0887] server:

[0888] The server provides data to a multimodal generation AI using the image analysis results, user-specified tastes and keywords, and the analyzed emotional state. Based on the information provided, the generation AI generates text that matches the user's emotional state in addition to the specified tastes.

[0889] For example, if the user selects "fantasy novel-style" and "adventure" as keywords, and the emotional state is recognized as "joy," the following sentence will be generated:

[0890] "In a meadow bathed in dazzling light, the little adventurer Leah stood up with a smile on her face. A fun-filled adventure with new friends awaited her."

[0891] 5. Check and edit the text

[0892] User:

[0893] The user can review the generated text on the interface and, if necessary, edit parts of the text to suit their preferences.

[0894] Device:

[0895] The terminal transmits the text edited by the user to the server.

[0896] server:

[0897] The server stores the text edited by the user and uses it as feedback data for future personalization and to improve the quality of the generated text.

[0898] 6. Save and share your writing

[0899] User:

[0900] Users can use the save and share buttons on the interface to save their completed writing and share it across various platforms.

[0901] Device:

[0902] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[0903] server:

[0904] The server stores the final text in a database for future reference and reuse, and when the user returns, it provides personalized results based on the previously saved data.

[0905] This embodiment allows users to easily create novel-like texts based on images, and then edit, save, and share them in a more personalized way that reflects their emotional state, further enriching the user's creative experience and helping to improve their writing skills and creative activities.

[0906] The processing flow will be explained below.

[0907] Step 1:

[0908] User:

[0909] Users simply select an image file from their device and click the upload button on the web interface or app. For example, a user can select a photo of a landscape from a travel destination or a pet, and then press the send button to upload it to the server.

[0910] Step 2:

[0911] Device:

[0912] The selected image file is sent to the server as an HTTP request.

[0913] Step 3:

[0914] server:

[0915] The received image file is temporarily saved and the image analysis module is prepared.

[0916] Step 4:

[0917] User:

[0918] The interface allows you to select the style of your writing, such as "light novel style," "horror novel style," or "fantasy novel style," and also enter keywords or themes you want to include in your writing (e.g., "adventure" or "friendship").

[0919] Step 5:

[0920] Device:

[0921] The selected tastes and keywords are sent to the server as form data.

[0922] Step 6:

[0923] server:

[0924] The received tastes and keywords are saved and the image analysis module is launched. The image analysis module analyzes the received image, identifies the main elements in the image (e.g., people, animals, landscapes), and extracts their features (color, shape, position, etc.).

[0925] Step 7:

[0926] User:

[0927] Through the interface, users provide their emotional state via input devices such as a webcam and microphone. Users input emotional data into the system through facial expressions, voice, and keyboard input speed.

[0928] Step 8:

[0929] server:

[0930] An emotion engine is activated to analyze the user's emotional state, and the emotion engine identifies emotions such as joy, sadness, anger, etc. based on the biosignal data provided by the user.

[0931] Step 9:

[0932] server:

[0933] The image analysis results, the user's emotional data, and the user's specified tastes and keywords are provided to a multimodal generation AI, which then generates text based on the given information, in addition to the specified tastes, that corresponds to the user's emotional state.

[0934] Step 10:

[0935] server:

[0936] The generated text is temporarily saved and prepared for presentation to the user.

[0937] Step 11:

[0938] User:

[0939] Review the generated text in the interface and, if necessary, edit parts of the text to suit your preferences.

[0940] Step 12:

[0941] Device:

[0942] The text edited by the user is sent to the server.

[0943] Step 13:

[0944] server:

[0945] User edits are saved and used for future personalization, and also as feedback data to improve the quality of generated text.

[0946] Step 14:

[0947] User:

[0948] To save your completed essay, click the Save button on the interface. You can also share your essay on social media by clicking the Share button.

[0949] Step 15:

[0950] Device:

[0951] When the save button is pressed, the final text is downloaded to the user's device. When the SNS share button is pressed, the SNS sharing API is called and the text is posted.

[0952] Step 16:

[0953] server:

[0954] The final text is stored in a database for future reference and reuse, and when the user returns, the previously stored data is used to provide personalized results.

[0955] Through these steps, users can easily generate novel-like texts based on images, and then edit, save, and share them with personalized content that reflects the user's emotional state. This enriches the user's creative experience, improves their writing skills, and helps with creative activities.

[0956] Example 2

[0957] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0958] While conventional text generation systems can generate text that takes into account user-specified tastes and keywords, they face the problem of difficulty in generating personalized text that reflects the user's emotional state. Furthermore, they lack the functionality to edit and adjust generated text to suit the user's preferences, limiting the user's creative experience. Furthermore, the complicated process of sharing and saving generated text reduces user convenience.

[0959] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0960] In this invention, the server includes: means for inputting an image; means for analyzing the input image and extracting key elements and their features within the image; means for inputting text tastes and keywords specified by the user; means for recognizing the user's emotional state; means for generating and adjusting text based on the recognized user's emotional state; means for presenting the generated text to the user and making it editable; means for analyzing biosignal data such as the user's facial expression, voice, and keyboard input speed; means for generating text based on the user's emotional state based on given information using a multimodal generative AI model; means for providing means for sharing the generated text on various platforms and for invoking a sharing API based on the user's instructions. This allows users to generate personalized text based on information extracted from the image and their own emotional state, enabling a richer creative experience through editing and sharing.

[0961] "Means for inputting images" is a function that allows a user to upload an image file selected from their own terminal to the system.

[0962] "Means for extracting the main elements and their features in an image" refers to a function that identifies the main elements, such as people, animals, and scenery, from the analyzed image and extracts these features as data.

[0963] The "means for inputting text style and keywords" is an input interface that allows the user to specify to the system the style of text to be generated and the content to be included.

[0964] The "means for recognizing the user's emotional state" is a function for analyzing biosignal data such as the user's facial expression, voice, keyboard input speed, etc., and identifying the user's current emotional state.

[0965] "Means for generating and adjusting sentences" refers to a function that uses specific algorithms and AI technology based on collected data to generate sentences that conform to specified tastes and keywords, and further adjusts them according to the user's emotional state.

[0966] The "means for presenting the generated text and making it editable" is an interface that displays the generated text to the user and allows the user to modify and edit parts of the text as necessary.

[0967] A "multimodal generative AI model" is an advanced artificial intelligence technology that integrates and processes information from multiple data sources (e.g., images, text, and emotional data) to generate natural language sentences.

[0968] The "means for analyzing biosignal data" is a function for analyzing biosignals such as facial expressions and voice obtained from the user and determining the emotional state from that data.

[0969] "Sharing API" means an application program interface used to share generated text across various platforms.

[0970] "Means for calling a sharing API based on user instructions" refers to a function that automatically calls the appropriate sharing API and posts text to a specified platform when a user performs an operation such as pressing a share button on the interface.

[0971] MODE FOR CARRYING OUT THE INVENTION

[0972] This system extracts key elements from an image provided by a user and generates text based on the user's preferences and keywords. It also has the ability to personalize, edit, save, and share text based on the user's emotional state. This system is implemented using the following hardware and software:

[0973] 1. Upload an image

[0974] User:

[0975] Users simply select an image file from their device and press the upload button through the system interface. For example, users can upload a photo of their pet at home or a landscape from a travel destination.

[0976] Device:

[0977] The device sends the selected image file as an HTTP request to the server, where the image data is temporarily stored in memory and sent to the server in the appropriate encoding format.

[0978] server:

[0979] The server temporarily stores the received image file in a secure storage location, and after storing it, assigns a unique identifier to the image file and uses that identifier for further processing.

[0980] 2. Select a style

[0981] User:

[0982] Users select the style of their writing (e.g., "fantasy novel-style") from the options provided on the interface, and then enter the keywords and themes they want to include in the writing.

[0983] Device:

[0984] The device sends the tastes and keywords selected by the user to the server as form data, where the input data is encoded in JSON format or similar.

[0985] server:

[0986] The server stores the received tastes and keywords in a database and then launches an image analysis module, which reads pre-saved image files, identifies key elements (e.g., people, animals, landscapes), and extracts their features.

[0987] 3. Emotional Recognition

[0988] User:

[0989] Users provide their emotional state to the system using a webcam or microphone, which is collected through facial expression recognition, voice tone analysis, keyboard input speed, etc.

[0990] server:

[0991] The server then analyzes the received biometric data using an emotion engine to identify specific emotions such as joy, sadness, anger, etc. The emotion analysis results are then used in the subsequent sentence generation process.

[0992] 4. Sentence Generation

[0993] server:

[0994] The server provides data to a multimodal generative AI model based on the results of image analysis, user-specified tastes and keywords, and the analyzed emotional state.

[0995] The generative AI model generates sentences based on the given information, in addition to the specified taste, according to the user's emotional state.

[0996] Examples:

[0997] If the user selects the keywords "fantasy novel-style" and "adventure" and the emotional state is recognized as "joy," the following sentence will be generated:

[0998] "In a meadow bathed in dazzling light, the little adventurer Leah stood up with a smile on her face. A fun-filled adventure with new friends awaited her."

[0999] 5. Check and edit the text

[1000] User:

[1001] The user can review the generated text in the interface and, if necessary, edit parts of the text to suit their preferences.

[1002] Device:

[1003] The terminal sends the text edited by the user to the server, again in JSON format or similar.

[1004] server:

[1005] The server stores the text edited by the user and uses it as feedback data for future personalization and system improvements.

[1006] 6. Save and share your writing

[1007] User:

[1008] Users can use the save and share buttons on the interface to save their completed writing and share it on various platforms.

[1009] Device:

[1010] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[1011] server:

[1012] The server stores the final text in a database for future reference and reuse, and when the user returns, it provides personalized results based on the previously saved data.

[1013] Prompt Sentence Examples

[1014] Below are some example prompts to input to a generative AI model:

[1015] "Right now, your emotional state is 'joy.' Create a fantasy-style adventure story based on a landscape photo. The keywords are 'light,' 'adventure,' and 'friends.'"

[1016] The system described above allows users to easily create novel-like texts based on images, and then edit, save, and share them with personalized content that reflects their emotional state. This will further enrich users' creative experiences, improve their writing skills, and help them with creative activities.

[1017] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1018] Step 1: Upload an image

[1019] input:

[1020] Users simply select any image file from their device and press the upload button through the system interface.

[1021] Specific behavior:

[1022] The user selects the image file of their choice through the specified interface and performs the upload operation, at which point the image file is temporarily stored on the user's device.

[1023] output:

[1024] The device sends the selected image file to the server as an HTTP request.

[1025] Step 2: Receive and save the image file

[1026] input:

[1027] Image files sent from the device

[1028] Specific behavior:

[1029] The server temporarily stores the received image file in storage and assigns a unique identifier (ID) to the file.

[1030] output:

[1031] The image file is given a unique identifier and stored in storage, which is used in the next processing step.

[1032] Step 3: Selecting Styles and Keywords

[1033] input:

[1034] The user specifies the style and keywords of the text.

[1035] Specific behavior:

[1036] The user selects the style of the text (e.g., "fantasy novel-style") from the options provided on the interface and enters keywords and themes.

[1037] output:

[1038] The device sends the selected taste and entered keywords as form data to the server. The sent data is encoded in JSON format.

[1039] Step 4: Begin image analysis

[1040] input:

[1041] Taste and keyword data, as well as identifiers of saved image files

[1042] Specific behavior:

[1043] The server receives the taste and keyword data and launches an image analysis module, which reads the saved image file, identifies key elements (e.g., people, animals, landscapes, etc.), and extracts their features.

[1044] output:

[1045] Data on the extracted image elements and their features are generated and used in the next processing step.

[1046] Step 5: Recognizing your emotional state

[1047] input:

[1048] Biosignal data collected from users (facial expressions, voice, keyboard typing speed, etc.)

[1049] Specific behavior:

[1050] Users provide their emotional state to the system using a webcam, microphone, etc. Emotional data is collected.

[1051] output:

[1052] The terminal transmits the collected biosignal data to the server.

[1053] Step 6: Analyze the sentiment data

[1054] input:

[1055] Collected biosignal data

[1056] Specific behavior:

[1057] The server analyzes the received biosignal data using an emotion engine to identify the user's emotional state (joy, sadness, anger, etc.).

[1058] output:

[1059] Parsed emotional state data is generated and used in the next processing step.

[1060] Step 7: Sentence generation

[1061] input:

[1062] Image analysis results, tastes and keywords, analyzed emotional states

[1063] Specific behavior:

[1064] The server provides the above data to a multimodal generative AI model, which uses the information to generate text that reflects the user's emotional state and the user's specified taste.

[1065] output:

[1066] The generated text is temporarily stored on the server.

[1067] Examples:

[1068] If the user sets the keywords "fantasy novel style" and "adventure" and the emotional state is recognized as "joy," the following sentence will be generated: "In a meadow bathed in dazzling light, the little adventurer Leah stood up with a smile on her face. A fun adventure with new friends awaited her."

[1069] Step 8: Presenting and editing the generated text

[1070] input:

[1071] Generated sentence data

[1072] Specific behavior:

[1073] The server sends the generated text to the interface, allowing the user to check the text content.

[1074] output:

[1075] The generated text is presented on the interface, and the user can view and edit it.

[1076] Step 9: Save your edits

[1077] input:

[1078] Text data edited by the user

[1079] Specific behavior:

[1080] The user can check the generated text on the interface, make corrections as necessary, and then click the save button when editing is complete.

[1081] output:

[1082] The terminal transmits the text data edited by the user to the server.

[1083] Step 10: Save edits and feedback

[1084] input:

[1085] Edited text data

[1086] Specific behavior:

[1087] The server stores the text data edited by the user and records it as feedback data to be used for future personalization and system improvement.

[1088] output:

[1089] The edited text data is stored in a database.

[1090] Step 11: Save and share your writing

[1091] input:

[1092] Completed sentence data

[1093] Specific behavior:

[1094] The user saves or shares the text using the save or share buttons on the interface.

[1095] output:

[1096] When the save button is pressed, the completed sentence is downloaded to the user's device. When the share button is pressed, the appropriate SNS sharing API is called and the sentence is posted.

[1097] Step 12: Save the final data

[1098] input:

[1099] Shared or stored text data

[1100] Specific behavior:

[1101] The server stores the final text in a database for future reference and reuse.

[1102] output:

[1103] The final text data is securely stored in a database and can be reused or referenced as needed.

[1104] In this way, this system allows users to easily create novel-like texts based on images, and then edit, save, and share them with personalized content that reflects their emotional state. This will further enrich the user's creative experience, improve their writing skills, and help with creative activities.

[1105] (Application example 2)

[1106] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1107] Conventional text generation systems generate text without considering the user's emotional state, which can result in text that does not match the user's emotions. This makes it difficult to generate text that the user can truly empathize with. Furthermore, conventional systems do not fully utilize the information in the image, limiting the quality of the generated text. This limits the user's creative experience.

[1108] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1109] In this invention, the server includes means for inputting an image, means for analyzing the input image and extracting the main elements and features of the image, means for inputting the taste and keywords of a sentence specified by the user, means for recognizing the analyzed emotional state of the user, means for generating a sentence using the extracted elements and features of the image based on the taste and keywords and the analyzed emotional state, and means for presenting the generated sentence to the user and making it editable. This makes it possible to generate sentences that correspond to the emotional state of the user and to generate high-quality sentences that the user can empathize with.

[1110] The "means for inputting images" refers to a method by which a user can provide a digital image to the system, and is a means by which image files can be uploaded using a smartphone or computer.

[1111] "Means for analyzing an input image and extracting the main elements and their features within the image" refers to a method that uses image recognition technology to identify the main elements of people, objects, landscapes, etc. that appear in an image and extract their features as data.

[1112] The "means for inputting the style and keywords of the text specified by the user" is an interface that allows the user to input the style of the text to be generated, the themes they wish to include, and keywords into the system.

[1113] "Means for recognizing an analyzed emotional state of a user" refers to a method that utilizes technology to identify a user's current emotional state through facial expression or voice analysis of the user.

[1114] "Means for generating sentences using extracted image elements and features based on the tastes and keywords and the analyzed emotional state" refers to a technology for generating sentences by incorporating tastes and keywords specified by the user and the recognized emotional state, and utilizing elements and features extracted from the image.

[1115] "Means for presenting the generated text to the user and making it editable" refers to an interface that displays the generated text to the user and allows the user to freely modify and edit the text.

[1116] The present invention relates to a system for generating sentences in accordance with the emotional state of a user, and a specific implementation method thereof will be described below.

[1117] Hardware and Software Configuration

[1118] This system includes an image input means, an image analysis means, a taste and keyword input means, an emotional state recognition means, a sentence generation means, and a means for making the generated sentences editable. The main hardware and software used are as follows:

[1119] Smartphones: iPhone, Android devices, etc.

[1120] Image recognition APIs: Google Cloud Vision, Amazon Rekognition.

[1121] Emotion analysis API: Microsoft Azure Emotion API, Affectiva.

[1122] Generative AI model: OpenAI GPT-4.

[1123] Database: Firebase, SQLite.

[1124] Processing flow

[1125] 1. Upload an image

[1126] Users select any image from their smartphone and use the app's interface to upload it to the server, which temporarily stores the received image for subsequent analysis.

[1127] 2. Selecting the theme and keywords

[1128] The user selects or inputs the text's style and keywords on the interface. For example, options such as "moving story," "pet," and "support" are displayed. This information is sent to the server as form data.

[1129] 3. Recognizing the user's emotional state

[1130] The user's real-time emotional state is analyzed through the smartphone's camera and microphone. An emotion analysis API is used to identify emotions such as joy, sadness, and anger from facial expressions and voice. This emotional state is then sent to the server.

[1131] 4. Sentence Generation

[1132] The server sends prompts to a generative AI model (OpenAI GPT-4) using the image analysis results, user-selected tastes and keywords, and the analyzed emotional state. The generative AI model generates high-quality novel-like text based on the given information.

[1133] For example, if a user uploads a picture of their pet, selects "Inspirational Stories," and the emotional state is recognized as "Joy," the prompt might look like this:

[1134] Image: A pet dog smiling and holding a ball

[1135] Taste: Inspirational story

[1136] Emotional state: Joy

[1137] Prompt: "This dog has supported me through some tough times. When he brought me the ball that day with a smile on his face..."

[1138] 5. Edit and save your text

[1139] The generated sentences are presented to the user on the interface. The user can review the generated sentences and edit them as necessary. The edited sentences are sent to the server for storage and used as feedback for future personalized generation results.

[1140] This allows users to generate, edit and save personalized text based on images and emotional states, enhancing their creativity.

[1141] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1142] Step 1:

[1143] Uploading an image

[1144] A user selects an image file from their smartphone and uploads it to the server through a specified interface. The device sends the selected image file to the server as an HTTP request, and the server temporarily stores the received image file. The input here is the image file selected by the user, and the output is the image file temporarily stored on the server.

[1145] Step 2:

[1146] Selecting styles and keywords

[1147] The user inputs the taste of the text (e.g., "moving story," "suspense," etc.) and keywords (e.g., "pet," "adventure," etc.) on the interface. The terminal sends the selected taste and keywords to the server as form data, which the server receives and stores. The input here is the taste and keywords selected by the user, and the output is the taste and keyword data stored on the server.

[1148] Step 3:

[1149] Recognizing the user's emotional state

[1150] The user's real-time emotional state is analyzed through the smartphone's camera and microphone. The device sends the acquired biosignal data to an emotion analysis API, which identifies the emotional state (e.g., joy, sadness, anger). The server receives and stores this emotional state data. The input here is the user's biosignal data (images, voice, etc.), and the output is analyzed emotional state data.

[1151] Step 4:

[1152] Image analysis

[1153] The server uses an image recognition API (e.g., Google Cloud Vision, Amazon Rekognition) to analyze the image uploaded by the user. The main elements in the image (e.g., people, animals, landscapes) and their features are extracted and saved as text data. The input here is the uploaded image file, and the output is the text data of the extracted main elements and features.

[1154] Step 5:

[1155] Sentence generation

[1156] The server sends prompts to a generative AI model (e.g., OpenAI GPT-4) based on the image analysis results, the tastes and keywords selected by the user, and the analyzed emotional state, to generate a sentence. The generated sentence is stored on the server. The inputs here are the image analysis results, tastes and keywords, and the emotional state, and the output is the text data of the generated sentence.

[1157] Step 6:

[1158] Editing and saving the generated text

[1159] The user can check the generated text on the interface and edit it as needed. The terminal sends the edited text to the server, which stores it. The input is the generated text and the user's edits, and the output is the final edited text that has been saved.

[1160] Step 7:

[1161] Sharing and saving text

[1162] Users use the save and share buttons on the interface to save their completed writing and share it on various platforms. When the save button is pressed, the device downloads the final writing to the user's device, and when the SNS share button is pressed, it calls the SNS's sharing API to post the writing. The server saves the final writing in a database for future reference and reuse. The input here is the user's save and share instructions, and the output is the saved writing and shared content.

[1163] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1164] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1165] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1166] [Third embodiment]

[1167] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1168] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1169] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1170] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1171] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1172] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1173] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1174] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1175] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1176] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1177] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1178] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1179] This invention relates to a technology for generating novel-like text from images. This system analyzes images provided by the user and generates text based on the analysis results, according to the taste and keywords specified by the user. Such a system promotes the improvement of writing skills and lowers the barrier to becoming a writer. Furthermore, the generated text can be edited by the user, and a personalization function allows adjustments to suit the user's preferences.

[1180] A natural language description of the system's programmatic processing

[1181] 1. Upload an image

[1182] User:

[1183] Users can select any image file from their device and upload it to the server through a specified interface. For example, users can upload photos of their pets at home or scenery from their travels.

[1184] Device:

[1185] The terminal transmits the selected image file to the server as an HTTP request.

[1186] server:

[1187] The server temporarily stores the received image file for further processing.

[1188] 2. Select a style

[1189] User:

[1190] The interface allows users to select the style of their writing, such as "light novel style," "horror novel style," or "fantasy novel style," and also allows users to enter keywords and themes they want to include in their writing.

[1191] Device:

[1192] The terminal transmits the selected taste and keywords to the server as form data.

[1193] server:

[1194] The server stores the received tastes and keywords and uses them for the next process.

[1195] 3. Image Analysis

[1196] server:

[1197] The server analyzes the uploaded images using an image analysis module, which uses image recognition technology to identify key elements in the image (e.g., people, animals, landscapes) and extract their features.

[1198] 4. Sentence Generation

[1199] server:

[1200] The server provides the analyzed image elements and features, as well as the taste and keywords selected by the user, to the multimodal generative AI. The generative AI generates text that matches the specified taste based on the provided information.

[1201] For example, if you set the theme to fantasy and adventure, the following text will be generated:

[1202] "A young adventurer named Aria wakes up in a forest bathed in dazzling light. She deepens her friendships with the companions she meets along the way and renews her resolve to venture into the unknown."

[1203] 5. Check and edit the text

[1204] User:

[1205] The user can view the generated text in the interface and, if necessary, edit parts of the text to suit their preferences.

[1206] Device:

[1207] The terminal transmits the text edited by the user to the server.

[1208] server:

[1209] The server stores the user's edits to serve as a reference for future personalization and as feedback data to improve the quality of the generated text.

[1210] 6. Save and share your writing

[1211] User:

[1212] Users can use the save and share buttons on the interface to save their completed writing and share it across various platforms.

[1213] Device:

[1214] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[1215] server:

[1216] The server stores the final text in a database as needed, allowing the user to use it comfortably the next time they access the site.

[1217] This embodiment allows users to easily create novel-like texts based on images and edit them to their liking. The generated texts are extremely useful for users who aspire to be writers, helping them improve their writing skills and assist in creative activities.

[1218] The processing flow will be explained below.

[1219] Step 1:

[1220] User:

[1221] The user selects an image file from their device, clicks the upload button in the web interface or app, selects the image, and then presses the send button to upload it to the server.

[1222] Step 2:

[1223] Device:

[1224] The terminal transmits the image file selected by the user to the server as an HTTP request.

[1225] Step 3:

[1226] server:

[1227] The server temporarily stores the received image file and prepares it for processing by the image analysis module.

[1228] Step 4:

[1229] User:

[1230] The user selects the style of the generated text on the interface, such as "light novel style," "horror novel style," or "fantasy novel style," and also inputs keywords or themes (e.g., "adventure" or "friendship") they want to include in the generated text.

[1231] Step 5:

[1232] Device:

[1233] The terminal transmits the selected taste and keywords to the server as form data.

[1234] Step 6:

[1235] server:

[1236] The server stores the received tastes and keywords and launches an image analysis module, which analyzes the received image, identifies key elements in the image (e.g., people, animals, landscapes), and extracts their features.

[1237] Step 7:

[1238] server:

[1239] Next, the server uses the image analysis results and the tastes and keywords specified by the user to provide data to a multimodal generation AI, which then generates text that matches the specified taste based on the provided data.

[1240] Step 8:

[1241] server:

[1242] The generated text is temporarily saved and prepared for presentation to the user.

[1243] Step 9:

[1244] User:

[1245] The user can check the generated text on the interface and, if necessary, edit parts of the text on the interface to suit their preferences.

[1246] Step 10:

[1247] Device:

[1248] The terminal transmits the text edited by the user to the server.

[1249] Step 11:

[1250] server:

[1251] The server stores the edited text by the user and uses it for future personalization and as feedback data to improve the quality of the generated text.

[1252] Step 12:

[1253] User:

[1254] Users can click the save button on the interface to save the completed text, and can also click the share button to share the text on social media.

[1255] Step 13:

[1256] Device:

[1257] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[1258] Step 14:

[1259] server:

[1260] The server stores the final text in a database for future reference and reuse, and when the user returns, it provides personalized results based on the previously saved data.

[1261] Through these steps, users can easily create novel-like texts based on images, and then edit, save, and share them according to their preferences.

[1262] Example 1

[1263] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1264] Currently, technology for generating text from images is rapidly developing, but most systems are limited to simply generating text that describes the content of an image, and their ability to generate novel-like creative writing is limited. In particular, generating personalized text based on user-specified tastes and keywords is difficult and time-consuming. Furthermore, there is a lack of functionality that allows users to easily edit, save, and share the generated text. These challenges have resulted in a significant amount of effort required by users engaged in creative activities.

[1265] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1266] In this invention, the server includes means for inputting an image, means for analyzing the input image and extracting key elements and their features within the image, means for inputting text tastes and keywords specified by the user, means for using a multimodal generative model to generate text using the extracted image elements and features based on the tastes and keywords, means for presenting the generated text to the user and making it editable, and means for saving the generated text and sharing it on various platforms. This allows users to easily generate novel-like text based on their specified tastes and keywords, and to edit, save, and share it.

[1267] The "means for inputting images" is a device that includes an interface and a protocol for users to upload any image file from their own terminal to the server.

[1268] "Means for analyzing an input image and extracting key elements and their features" refers to a device that includes software modules and algorithms for using image recognition technology to identify key elements, such as people, animals, and scenery, in an image and extract their features.

[1269] "Means for inputting the style and keywords of the text specified by the user" refers to a device that allows the user to input the style of the text they desire (e.g., light novel style, horror novel style) and the keywords they want to include through an interface.

[1270] A "means for using a multimodal generative model" is a device that uses a generative AI model (e.g., a generative AI model) to process image and text data as input and generate personalized text based on specified tastes and keywords.

[1271] The "means for presenting the generated text to the user and enabling the user to edit it" is a device that includes an interface and software that displays the generated text to the user and allows the user to edit it.

[1272] "Means for saving generated text and sharing it on various platforms" refers to a device that includes an API and functionality for saving generated text and allowing users to download it or post it on social media or other sharing platforms.

[1273] This invention relates to a system that generates novel-like text from images. The system analyzes images provided by the user and generates text based on the analysis results according to the taste and keywords specified by the user. This promotes the improvement of writing skills and lowers the barrier to becoming a writer. The generated text can also be edited by the user, and a personalization function allows for adjustments to suit the user's preferences.

[1274] First, the user selects an image using their device and uploads it to the server through a specified interface. For example, the user might upload a photo of a pet at home or a landscape from a travel destination. This image file is sent to the server by the device as an HTTP request. The server temporarily stores the received image file for further processing. In this case, the server uses a storage solution such as AWS S3 or Google Cloud Storage.

[1275] Next, the user enters the style of the text (e.g., light novel style, horror novel style, fantasy novel style) and keywords (e.g., "adventure," "friendship," "hero") on the interface. The terminal sends the selected style and keywords as form data to the server, and the server stores the received data. For example, a database system such as MySQL or Firebase is used for this purpose.

[1276] The server analyzes the uploaded images using an image analysis module. Specifically, it uses image recognition technologies such as TensorFlow and OpenCV to identify key elements in the image (e.g., people, animals, landscapes) and extract their features. For example, it uses a TensorFlow image classification model to scan the image and identify key elements (e.g., dogs, mountains, rivers, etc.).

[1277] Next, the server provides the analyzed image elements and features, as well as the taste and keywords selected by the user, to a multimodal generative AI model (e.g., OpenAI's generative AI model).The generative AI model generates text that matches the specified taste based on the provided information.

[1278] For example, if you set the theme to fantasy and adventure, the following text will be generated:

[1279] "A young adventurer named Aria wakes up in a forest bathed in dazzling light. She deepens her friendships with the companions she meets along the way and renews her resolve to venture into the unknown."

[1280] The user can view the generated text on the interface and, if necessary, edit parts of the text to suit their preferences. The edited text is then sent from the device to the server, where it is stored. The stored data is used as a reference for future personalization and as feedback to improve the quality of the generative AI model.

[1281] Finally, the user can save the completed sentence and share it on various platforms. When the save button is pressed, the final sentence is downloaded to the user's device. When the SNS share button is pressed, the SNS sharing API is called and the sentence is posted. The server stores the final sentence in a database as needed, allowing the user to use it comfortably the next time they access the site.

[1282] An example of a specific prompt would be something like this:

[1283] Image: Upload your own image (e.g., a photo of your pet, a travel photo, etc.)

[1284] Style: Fantasy novel-style, adventure theme

[1285] Keywords: "Hero", "Friendship"

[1286] Generated text:

[1287] Waffles, a brave dog, embarks on a quest to save the magical kingdom with his best friend Claire. Together, they overcome many obstacles and rediscover the power of true friendship.

[1288] This system allows users to easily generate novel-like texts based on images and edit them to their liking. The generated texts are extremely useful for aspiring writers, helping them improve their writing skills and their creative activities.

[1289] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1290] Step 1:

[1291] User: The user selects an image file from their device and uploads it to the server through a specified interface. For example, the user may select a photo of a pet at home or a landscape from a travel destination.

[1292] Input: An image file selected by the user.

[1293] How it works: The user selects an image using the file selection dialog and clicks the upload button.

[1294] Output: The selected image file is sent to the next step.

[1295] Step 2:

[1296] Terminal: The terminal sends the selected image file to the server as an HTTP request.

[1297] Input: An image file selected by the user.

[1298] How it works: The device generates an HTML form, converts the selected image file to multipart / form-data format, and creates an HTTP POST request to send to the server.

[1299] Output: The image file is uploaded to the server.

[1300] Step 3:

[1301] Server: The server temporarily stores the received image files for further processing.

[1302] Input: Image file sent from the device.

[1303] How it works: The server parses the incoming HTTP request, extracts the image file, and saves it to storage (e.g., AWS S3 or Google Cloud Storage). After saving, it generates a data path for the next analysis step.

[1304] Output: The path of the image file saved in storage will be sent to the next step.

[1305] Step 4:

[1306] User: The user inputs the style of the text (e.g., light novel style, horror novel style, fantasy novel style) and keywords on the interface.

[1307] Input: Tastes and keywords entered by the user.

[1308] How it works: The user enters the desired taste and keywords using the drop-down menus and text fields.

[1309] Output: The input tastes and keywords are sent to the next step.

[1310] Step 5:

[1311] Terminal: The terminal sends the selected tastes and keywords to the server as form data.

[1312] Input: Tastes and keywords entered by the user.

[1313] Operation: The device converts the input data of tastes and keywords into JSON format and sends it to the server as an HTTP POST request.

[1314] Output: The taste and keyword information is sent to the server.

[1315] Step 6:

[1316] Server: The server stores the received tastes and keywords and uses them for the next process.

[1317] Input: Tastes and keywords entered by the user.

[1318] How it works: The server parses the received JSON data and stores it in a database (e.g., MySQL or Firebase). After saving, it proceeds to the next image analysis step.

[1319] Output: The saved taste and keyword information will be used in the next step.

[1320] Step 7:

[1321] Server: The server uses an image analysis module to analyze the uploaded images.

[1322] Input: The path of the image file retrieved from storage.

[1323] How it works: The server uses image recognition technology from TensorFlow and OpenCV to scan an image, identify key elements (e.g. people, animals, landscapes) and extract their features.

[1324] Output: Key elements and feature data in the image are sent to the next step.

[1325] Step 8:

[1326] Server: The server provides the analyzed image elements and features, as well as user-selected tastes and keywords, to the multimodal generative AI model as input.

[1327] Input: Extracted image elements and feature data, user-selected tastes and keywords.

[1328] How it works: The server inputs this data in JSON format into a generative AI model and sends a request to generate the appropriate sentence.

[1329] Output: The generated sentence is obtained as a response from the generative AI model.

[1330] Step 9:

[1331] User: The user checks the generated text on the interface and edits it as necessary.

[1332] Input: The generated sentence.

[1333] How it works: In the interface that displays the text, the user edits parts of the text as needed.

[1334] Output: The edited text data is sent to the next step.

[1335] Step 10:

[1336] Terminal: The terminal transmits the text edited by the user to the server.

[1337] Input: User edited text.

[1338] Operation: The device converts the edited text data into JSON format and sends it to the server as an HTTP POST request.

[1339] Output: The edited text data is sent to the server.

[1340] Step 11:

[1341] Server: The server stores the user's edits and uses them as a reference for future personalization.

[1342] Input: Edited text data.

[1343] How it works: The server parses the received JSON data and stores it in a database. It also uses it as feedback data to improve the quality of the generative AI model.

[1344] Output: Saved edited text data.

[1345] Step 12:

[1346] User: Users save their completed writing and share it on various platforms.

[1347] Input: Final sentence data.

[1348] How it works: Click the save or share button on the interface to share your text on social media or other platforms.

[1349] Output: The saved text is downloaded to the user's device or posted to a sharing platform.

[1350] Step 13:

[1351] Server: The server stores the final text in a database if necessary, making it easier for the user to use the text the next time they access the site.

[1352] Input: Final sentence data.

[1353] How it works: The server stores the final text data in a database and associates it with the user's profile.

[1354] Output: The saved text data is stored for future use.

[1355] (Application example 1)

[1356] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1357] While there have been technologies that generate text from images, there has been a lack of a way to save the text and easily share it across various platforms. It has also been difficult to personalize the generated text and tune it to the user's preferences. As a result, it has been difficult to provide a personalized creative experience for users.

[1358] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1359] In this invention, the server includes means for inputting an image, means for analyzing the input image and extracting key elements and their characteristics from the image, means for inputting text tastes and keywords specified by the user, means for generating text using the extracted image elements and characteristics based on the tastes and keywords, means for presenting the generated text to the user and making it editable, and means for saving the generated text and sharing it on various platforms. This allows users to easily save and share the generated text and provide a personalized creation experience tailored to their preferences.

[1360] "Means for inputting images" refers to a function that allows a user to select any image from the terminal and upload it to the server through a specified interface.

[1361] "Means for analyzing an image and extracting the main elements and their features within the image" refers to a function that uses image analysis technology to analyze an input image, identify the main elements within it, such as people, animals, and scenery, and extract their features.

[1362] "Means for inputting text style and keywords" is a function that allows the user to input the style and theme of the text to be generated on the interface, as well as the keywords that the user wants to include.

[1363] "Means for generating text using extracted image elements and features" refers to a function in which a multimodal generative AI generates text based on the elements and features obtained through image analysis and the tastes and keywords specified by the user.

[1364] The "means for presenting the generated text to the user and enabling it to be edited" is a function that provides the generated text on an interface so that the user can check it and edit it as necessary.

[1365] "Means for saving the generated text and sharing it on various platforms" refers to a function that allows users to save the final text and share it on platforms such as social media and email.

[1366] Overall system flow

[1367] Users operate the system using a smartphone. Specifically, they upload photographed or existing images to the application and input the text style and keywords. The system then analyzes the images and generates text using a generative AI model. The generated text can be reviewed and edited by the user, and can also be saved and shared.

[1368] Hardware and Software

[1369] Hardware

[1370] Smartphone: The device where the user selects and uploads images to the application.

[1371] Server: The central processing unit that analyzes images and generates text.

[1372] software

[1373] Image Analysis Module: Software for analyzing elements in images using TensorFlow and PyTorch.

[1374] Generative AI model: Using OpenAI's GPT-4 and other models, sentences are generated based on the analysis results and conditions specified by the user.

[1375] Front-end interface: The interface through which users upload images, select settings, and review and edit the generated text.

[1376] SNS sharing API: An API for sharing generated text on the SNS platform of the user's choice.

[1377] Data processing and calculation

[1378] When a user uploads an image from their smartphone, the server first temporarily stores the image. Next, an image analysis module is used to identify the main elements of the uploaded image (e.g., people, animals, landscapes) and extract their characteristics. Next, a generative AI model combines the analysis results with the text to generate a sentence based on the user-specified text style (e.g., "fantasy," "mystery," etc.).

[1379] Examples of concrete examples and prompts

[1380] For example, if a user uploads a photo of a beach at sunset, selects "fantasy" as the text style, and enters "adventure" and "magic" as keywords, the system will input the following prompt sentences into the generative AI model:

[1381] Generates the prompt statement:

[1382] The image shows a beach at sunset. Create a fantasy story using this image. Theme: Adventure, Magic.

[1383] Based on this prompt, the generative AI model might generate a sentence like this:

[1384] Examples of generated sentences:

[1385] "A golden sunset sunk into the beach. The wizard set off towards the boat waiting on the shore, eager for new adventures. The faint light from the beach shone on his magic wand."

[1386] The generated text can be viewed and edited on the interface, and can then be shared via social media or email. Through this series of operations, users can easily manage the generated text and enjoy creative activities.

[1387] In this way, the system of the present invention provides users with a comprehensive solution for generating, editing, and sharing novel-like text from images.

[1388] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1389] Step 1:

[1390] User: Selects an image from the smartphone's photo gallery and uploads it to the application.

[1391] Input: Image files on your smartphone

[1392] Output: Image file upload request to the server

[1393] Terminal: The selected image file is sent to the server as an HTTP request and temporarily stored.

[1394] Input: An image file selected by the user

[1395] Output: Image data transferred to the server

[1396] Server: Temporarily stores received image files and prepares them for the next processing step.

[1397] Input: Image data transferred from the device

[1398] Output: Temporarily saved image file

[1399] Step 2:

[1400] User: Select and input text style and keywords in the application interface.

[1401] Input: Text style (e.g., "fantasy"), keywords (e.g., "adventure", "magic")

[1402] Output: Input data of tastes and keywords

[1403] Device: Sends the selected tastes and keywords to the server.

[1404] Input: User input data of tastes and keywords

[1405] Output: Taste and keyword data sent to the server

[1406] Server: Save the received tastes and keywords and use them for the next process.

[1407] Input: Taste and keyword data sent from the device

[1408] Output: Saved tastes and keyword data

[1409] Step 3:

[1410] Server: Analyzes the uploaded image using an image analysis module. This analysis module uses image recognition techniques such as TensorFlow and PyTorch to identify key elements in the image and extract their features.

[1411] Input: Temporarily saved image file

[1412] Output: Identified key elements and their characteristic data

[1413] Step 4:

[1414] Server: Provides the analyzed image elements and features, as well as user-specified tastes and keywords, to the generative AI model. The generative AI model (e.g., OpenAI's GPT-4) generates prompts based on the provided information and creates full sentences.

[1415] Input: Analysis data (image elements and features), taste, keywords

[1416] Output: Generated sentence

[1417] Step 5:

[1418] Server: Sends the generated text to an interface that presents it to the user so that it can be edited by the user.

[1419] Input: Generated sentence

[1420] Output: Text displayed on the user's terminal

[1421] User: Check the generated text on the interface and edit it as necessary.

[1422] Input: Generated text sent from the server

[1423] Output: Edited text

[1424] Terminal: The user sends the edited results to the server.

[1425] Input: Text data edited by the user

[1426] Output: Edited text data sent to the server

[1427] Step 6:

[1428] Server: Saves edited text and connects to the sharing APIs of various platforms to enable sharing on social media and other platforms.

[1429] Input: Edited text data

[1430] Output: Final text data saved and text shared via the sharing API

[1431] User: Use the save and share buttons to share the generated text on various platforms.

[1432] Input: Final sentence

[1433] Output: Content shared across platforms

[1434] This series of steps creates a system that allows users to easily create, review, edit, save, and share text generated from images.

[1435] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1436] This invention combines technology for generating novel-like text from images with an emotion engine, enabling text generation that responds to the user's emotional state. This system analyzes images provided by the user and, based on the analysis results, generates and adjusts text according to the user's emotional state, incorporating tastes and keywords specified by the user. This system promotes the improvement of writing skills and lowers the barrier to becoming a writer. Furthermore, the generated text can be edited by the user, and a personalization function allows adjustments to the user's preferences.

[1437] A natural language description of the system's programmatic processing

[1438] 1. Upload an image

[1439] User:

[1440] Users can select any image file from their device and upload it to the server through a specified interface. For example, users can upload photos of their pets at home or scenery from their travels.

[1441] Device:

[1442] The terminal transmits the selected image file to the server as an HTTP request.

[1443] server:

[1444] The server temporarily stores the received image file for further processing.

[1445] 2. Select a style

[1446] User:

[1447] The interface allows users to select the style of their writing, such as "light novel style," "horror novel style," or "fantasy novel style," and also allows users to enter keywords and themes they want to include in their writing.

[1448] Device:

[1449] The terminal transmits the selected taste and keywords to the server as form data.

[1450] server:

[1451] The server stores the received tastes and keywords and launches an image analysis module, which analyzes the received image, identifies key elements in the image (e.g., people, animals, landscapes), and extracts their features.

[1452] 3. Emotional Recognition

[1453] User:

[1454] Users provide their emotional state to the system through input devices such as a webcam and microphone, which is collected through analysis of facial expressions, voice, and keyboard typing speed.

[1455] server:

[1456] The server analyzes the user's emotional state using an emotion engine, which identifies emotions such as joy, sadness, and anger based on biosignal data provided by the user.

[1457] 4. Sentence Generation

[1458] server:

[1459] The server provides data to a multimodal generation AI using the image analysis results, user-specified tastes and keywords, and the analyzed emotional state. Based on the information provided, the generation AI generates text that matches the user's emotional state in addition to the specified tastes.

[1460] For example, if the user selects "fantasy novel-style" and "adventure" as keywords, and the emotional state is recognized as "joy," the following sentence will be generated:

[1461] "In a meadow bathed in dazzling light, the little adventurer Leah stood up with a smile on her face. A fun-filled adventure with new friends awaited her."

[1462] 5. Check and edit the text

[1463] User:

[1464] The user can review the generated text on the interface and, if necessary, edit parts of the text to suit their preferences.

[1465] Device:

[1466] The terminal transmits the text edited by the user to the server.

[1467] server:

[1468] The server stores the text edited by the user and uses it as feedback data for future personalization and to improve the quality of the generated text.

[1469] 6. Save and share your writing

[1470] User:

[1471] Users can use the save and share buttons on the interface to save their completed writing and share it across various platforms.

[1472] Device:

[1473] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[1474] server:

[1475] The server stores the final text in a database for future reference and reuse, and when the user returns, it provides personalized results based on the previously saved data.

[1476] This embodiment allows users to easily create novel-like texts based on images, and then edit, save, and share them in a more personalized way that reflects their emotional state, further enriching the user's creative experience and helping to improve their writing skills and creative activities.

[1477] The processing flow will be explained below.

[1478] Step 1:

[1479] User:

[1480] Users simply select an image file from their device and click the upload button on the web interface or app. For example, a user can select a photo of a landscape from a travel destination or a pet, and then press the send button to upload it to the server.

[1481] Step 2:

[1482] Device:

[1483] The selected image file is sent to the server as an HTTP request.

[1484] Step 3:

[1485] server:

[1486] The received image file is temporarily saved and the image analysis module is prepared.

[1487] Step 4:

[1488] User:

[1489] The interface allows you to select the style of your writing, such as "light novel style," "horror novel style," or "fantasy novel style," and also enter keywords or themes you want to include in your writing (e.g., "adventure" or "friendship").

[1490] Step 5:

[1491] Device:

[1492] The selected tastes and keywords are sent to the server as form data.

[1493] Step 6:

[1494] server:

[1495] The received tastes and keywords are saved and the image analysis module is launched. The image analysis module analyzes the received image, identifies the main elements in the image (e.g., people, animals, landscapes), and extracts their features (color, shape, position, etc.).

[1496] Step 7:

[1497] User:

[1498] Through the interface, users provide their emotional state via input devices such as a webcam and microphone. Users input emotional data into the system through facial expressions, voice, and keyboard input speed.

[1499] Step 8:

[1500] server:

[1501] An emotion engine is activated to analyze the user's emotional state, and the emotion engine identifies emotions such as joy, sadness, anger, etc. based on the biosignal data provided by the user.

[1502] Step 9:

[1503] server:

[1504] The image analysis results, the user's emotional data, and the user's specified tastes and keywords are provided to a multimodal generation AI, which then generates text based on the given information, in addition to the specified tastes, that corresponds to the user's emotional state.

[1505] Step 10:

[1506] server:

[1507] The generated text is temporarily saved and prepared for presentation to the user.

[1508] Step 11:

[1509] User:

[1510] Review the generated text in the interface and, if necessary, edit parts of the text to suit your preferences.

[1511] Step 12:

[1512] Device:

[1513] The text edited by the user is sent to the server.

[1514] Step 13:

[1515] server:

[1516] User edits are saved and used for future personalization, and also as feedback data to improve the quality of generated text.

[1517] Step 14:

[1518] User:

[1519] To save your completed essay, click the Save button on the interface. You can also share your essay on social media by clicking the Share button.

[1520] Step 15:

[1521] Device:

[1522] When the save button is pressed, the final text is downloaded to the user's device. When the SNS share button is pressed, the SNS sharing API is called and the text is posted.

[1523] Step 16:

[1524] server:

[1525] The final text is stored in a database for future reference and reuse, and when the user returns, the previously stored data is used to provide personalized results.

[1526] Through these steps, users can easily generate novel-like texts based on images, and then edit, save, and share them with personalized content that reflects the user's emotional state. This enriches the user's creative experience, improves their writing skills, and helps with creative activities.

[1527] Example 2

[1528] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1529] While conventional text generation systems can generate text that takes into account user-specified tastes and keywords, they face the problem of difficulty in generating personalized text that reflects the user's emotional state. Furthermore, they lack the functionality to edit and adjust generated text to suit the user's preferences, limiting the user's creative experience. Furthermore, the complicated process of sharing and saving generated text reduces user convenience.

[1530] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1531] In this invention, the server includes: means for inputting an image; means for analyzing the input image and extracting key elements and their features within the image; means for inputting text tastes and keywords specified by the user; means for recognizing the user's emotional state; means for generating and adjusting text based on the recognized user's emotional state; means for presenting the generated text to the user and making it editable; means for analyzing biosignal data such as the user's facial expression, voice, and keyboard input speed; means for generating text based on the user's emotional state based on given information using a multimodal generative AI model; means for providing means for sharing the generated text on various platforms and for invoking a sharing API based on the user's instructions. This allows users to generate personalized text based on information extracted from the image and their own emotional state, enabling a richer creative experience through editing and sharing.

[1532] "Means for inputting images" is a function that allows a user to upload an image file selected from their own terminal to the system.

[1533] "Means for extracting the main elements and their features in an image" refers to a function that identifies the main elements, such as people, animals, and scenery, from the analyzed image and extracts these features as data.

[1534] The "means for inputting text style and keywords" is an input interface that allows the user to specify to the system the style of text to be generated and the content to be included.

[1535] The "means for recognizing the user's emotional state" is a function for analyzing biosignal data such as the user's facial expression, voice, keyboard input speed, etc., and identifying the user's current emotional state.

[1536] "Means for generating and adjusting sentences" refers to a function that uses specific algorithms and AI technology based on collected data to generate sentences that conform to specified tastes and keywords, and further adjusts them according to the user's emotional state.

[1537] The "means for presenting the generated text and making it editable" is an interface that displays the generated text to the user and allows the user to modify and edit parts of the text as necessary.

[1538] A "multimodal generative AI model" is an advanced artificial intelligence technology that integrates and processes information from multiple data sources (e.g., images, text, and emotional data) to generate natural language sentences.

[1539] The "means for analyzing biosignal data" is a function for analyzing biosignals such as facial expressions and voice obtained from the user and determining the emotional state from that data.

[1540] "Sharing API" means an application program interface used to share generated text across various platforms.

[1541] "Means for calling a sharing API based on user instructions" refers to a function that automatically calls the appropriate sharing API and posts text to a specified platform when a user performs an operation such as pressing a share button on the interface.

[1542] MODE FOR CARRYING OUT THE INVENTION

[1543] This system extracts key elements from an image provided by a user and generates text based on the user's preferences and keywords. It also has the ability to personalize, edit, save, and share text based on the user's emotional state. This system is implemented using the following hardware and software:

[1544] 1. Upload an image

[1545] User:

[1546] Users simply select an image file from their device and press the upload button through the system interface. For example, users can upload a photo of their pet at home or a landscape from a travel destination.

[1547] Device:

[1548] The device sends the selected image file as an HTTP request to the server, where the image data is temporarily stored in memory and sent to the server in the appropriate encoding format.

[1549] server:

[1550] The server temporarily stores the received image file in a secure storage location, and after storing it, assigns a unique identifier to the image file and uses that identifier for further processing.

[1551] 2. Select a style

[1552] User:

[1553] Users select the style of their writing (e.g., "fantasy novel-style") from the options provided on the interface, and then enter the keywords and themes they want to include in the writing.

[1554] Device:

[1555] The device sends the tastes and keywords selected by the user to the server as form data, where the input data is encoded in JSON format or similar.

[1556] server:

[1557] The server stores the received tastes and keywords in a database and then launches an image analysis module, which reads pre-saved image files, identifies key elements (e.g., people, animals, landscapes), and extracts their features.

[1558] 3. Emotional Recognition

[1559] User:

[1560] Users provide their emotional state to the system using a webcam or microphone, which is collected through facial expression recognition, voice tone analysis, keyboard input speed, etc.

[1561] server:

[1562] The server then analyzes the received biometric data using an emotion engine to identify specific emotions such as joy, sadness, anger, etc. The emotion analysis results are then used in the subsequent sentence generation process.

[1563] 4. Sentence Generation

[1564] server:

[1565] The server provides data to a multimodal generative AI model based on the results of image analysis, user-specified tastes and keywords, and the analyzed emotional state.

[1566] The generative AI model generates sentences based on the given information, in addition to the specified taste, according to the user's emotional state.

[1567] Examples:

[1568] If the user selects the keywords "fantasy novel-style" and "adventure" and the emotional state is recognized as "joy," the following sentence will be generated:

[1569] "In a meadow bathed in dazzling light, the little adventurer Leah stood up with a smile on her face. A fun-filled adventure with new friends awaited her."

[1570] 5. Check and edit the text

[1571] User:

[1572] The user can review the generated text in the interface and, if necessary, edit parts of the text to suit their preferences.

[1573] Device:

[1574] The terminal sends the text edited by the user to the server, again in JSON format or similar.

[1575] server:

[1576] The server stores the text edited by the user and uses it as feedback data for future personalization and system improvements.

[1577] 6. Save and share your writing

[1578] User:

[1579] Users can use the save and share buttons on the interface to save their completed writing and share it on various platforms.

[1580] Device:

[1581] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[1582] server:

[1583] The server stores the final text in a database for future reference and reuse, and when the user returns, it provides personalized results based on the previously saved data.

[1584] Prompt Sentence Examples

[1585] Below are some example prompts to input to a generative AI model:

[1586] "Right now, your emotional state is 'joy.' Create a fantasy-style adventure story based on a landscape photo. The keywords are 'light,' 'adventure,' and 'friends.'"

[1587] The system described above allows users to easily create novel-like texts based on images, and then edit, save, and share them with personalized content that reflects their emotional state. This will further enrich users' creative experiences, improve their writing skills, and help them with creative activities.

[1588] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1589] Step 1: Upload an image

[1590] input:

[1591] Users simply select any image file from their device and press the upload button through the system interface.

[1592] Specific behavior:

[1593] The user selects the image file of their choice through the specified interface and performs the upload operation, at which point the image file is temporarily stored on the user's device.

[1594] output:

[1595] The device sends the selected image file to the server as an HTTP request.

[1596] Step 2: Receive and save the image file

[1597] input:

[1598] Image files sent from the device

[1599] Specific behavior:

[1600] The server temporarily stores the received image file in storage and assigns a unique identifier (ID) to the file.

[1601] output:

[1602] The image file is given a unique identifier and stored in storage, which is used in the next processing step.

[1603] Step 3: Selecting Styles and Keywords

[1604] input:

[1605] The user specifies the style and keywords of the text.

[1606] Specific behavior:

[1607] The user selects the style of the text (e.g., "fantasy novel-style") from the options provided on the interface and enters keywords and themes.

[1608] output:

[1609] The device sends the selected taste and entered keywords as form data to the server. The sent data is encoded in JSON format.

[1610] Step 4: Begin image analysis

[1611] input:

[1612] Taste and keyword data, as well as identifiers of saved image files

[1613] Specific behavior:

[1614] The server receives the taste and keyword data and launches an image analysis module, which reads the saved image file, identifies key elements (e.g., people, animals, landscapes, etc.), and extracts their features.

[1615] output:

[1616] Data on the extracted image elements and their features are generated and used in the next processing step.

[1617] Step 5: Recognizing your emotional state

[1618] input:

[1619] Biosignal data collected from users (facial expressions, voice, keyboard typing speed, etc.)

[1620] Specific behavior:

[1621] Users provide their emotional state to the system using a webcam, microphone, etc. Emotional data is collected.

[1622] output:

[1623] The terminal transmits the collected biosignal data to the server.

[1624] Step 6: Analyze the sentiment data

[1625] input:

[1626] Collected biosignal data

[1627] Specific behavior:

[1628] The server analyzes the received biosignal data using an emotion engine to identify the user's emotional state (joy, sadness, anger, etc.).

[1629] output:

[1630] Parsed emotional state data is generated and used in the next processing step.

[1631] Step 7: Sentence generation

[1632] input:

[1633] Image analysis results, tastes and keywords, analyzed emotional states

[1634] Specific behavior:

[1635] The server provides the above data to a multimodal generative AI model, which uses the information to generate text that reflects the user's emotional state and the user's specified taste.

[1636] output:

[1637] The generated text is temporarily stored on the server.

[1638] Examples:

[1639] If the user sets the keywords "fantasy novel style" and "adventure" and the emotional state is recognized as "joy," the following sentence will be generated: "In a meadow bathed in dazzling light, the little adventurer Leah stood up with a smile on her face. A fun adventure with new friends awaited her."

[1640] Step 8: Presenting and editing the generated text

[1641] input:

[1642] Generated sentence data

[1643] Specific behavior:

[1644] The server sends the generated text to the interface, allowing the user to check the text content.

[1645] output:

[1646] The generated text is presented on the interface, and the user can view and edit it.

[1647] Step 9: Save your edits

[1648] input:

[1649] Text data edited by the user

[1650] Specific behavior:

[1651] The user can check the generated text on the interface, make corrections as necessary, and then click the save button when editing is complete.

[1652] output:

[1653] The terminal transmits the text data edited by the user to the server.

[1654] Step 10: Save edits and feedback

[1655] input:

[1656] Edited text data

[1657] Specific behavior:

[1658] The server stores the text data edited by the user and records it as feedback data to be used for future personalization and system improvement.

[1659] output:

[1660] The edited text data is stored in a database.

[1661] Step 11: Save and share your writing

[1662] input:

[1663] Completed sentence data

[1664] Specific behavior:

[1665] The user saves or shares the text using the save or share buttons on the interface.

[1666] output:

[1667] When the save button is pressed, the completed sentence is downloaded to the user's device. When the share button is pressed, the appropriate SNS sharing API is called and the sentence is posted.

[1668] Step 12: Save the final data

[1669] input:

[1670] Shared or stored text data

[1671] Specific behavior:

[1672] The server stores the final text in a database for future reference and reuse.

[1673] output:

[1674] The final text data is securely stored in a database and can be reused or referenced as needed.

[1675] In this way, this system allows users to easily create novel-like texts based on images, and then edit, save, and share them with personalized content that reflects their emotional state. This will further enrich the user's creative experience, improve their writing skills, and help with creative activities.

[1676] (Application example 2)

[1677] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1678] Conventional text generation systems generate text without considering the user's emotional state, which can result in text that does not match the user's emotions. This makes it difficult to generate text that the user can truly empathize with. Furthermore, conventional systems do not fully utilize the information in the image, limiting the quality of the generated text. This limits the user's creative experience.

[1679] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1680] In this invention, the server includes means for inputting an image, means for analyzing the input image and extracting the main elements and features of the image, means for inputting the taste and keywords of a sentence specified by the user, means for recognizing the analyzed emotional state of the user, means for generating a sentence using the extracted elements and features of the image based on the taste and keywords and the analyzed emotional state, and means for presenting the generated sentence to the user and making it editable. This makes it possible to generate sentences that correspond to the emotional state of the user and to generate high-quality sentences that the user can empathize with.

[1681] The "means for inputting images" refers to a method by which a user can provide a digital image to the system, and is a means by which image files can be uploaded using a smartphone or computer.

[1682] "Means for analyzing an input image and extracting the main elements and their features within the image" refers to a method that uses image recognition technology to identify the main elements of people, objects, landscapes, etc. that appear in an image and extract their features as data.

[1683] The "means for inputting the style and keywords of the text specified by the user" is an interface that allows the user to input the style of the text to be generated, the themes they wish to include, and keywords into the system.

[1684] "Means for recognizing an analyzed emotional state of a user" refers to a method that utilizes technology to identify a user's current emotional state through facial expression or voice analysis of the user.

[1685] "Means for generating sentences using extracted image elements and features based on the tastes and keywords and the analyzed emotional state" refers to a technology for generating sentences by incorporating tastes and keywords specified by the user and the recognized emotional state, and utilizing elements and features extracted from the image.

[1686] "Means for presenting the generated text to the user and making it editable" refers to an interface that displays the generated text to the user and allows the user to freely modify and edit the text.

[1687] The present invention relates to a system for generating sentences in accordance with the emotional state of a user, and a specific implementation method thereof will be described below.

[1688] Hardware and Software Configuration

[1689] This system includes an image input means, an image analysis means, a taste and keyword input means, an emotional state recognition means, a sentence generation means, and a means for making the generated sentences editable. The main hardware and software used are as follows:

[1690] Smartphones: iPhone, Android devices, etc.

[1691] Image recognition APIs: Google Cloud Vision, Amazon Rekognition.

[1692] Emotion analysis API: Microsoft Azure Emotion API, Affectiva.

[1693] Generative AI model: OpenAI GPT-4.

[1694] Database: Firebase, SQLite.

[1695] Processing flow

[1696] 1. Upload an image

[1697] Users select any image from their smartphone and use the app's interface to upload it to the server, which temporarily stores the received image for subsequent analysis.

[1698] 2. Selecting the theme and keywords

[1699] The user selects or inputs the text's style and keywords on the interface. For example, options such as "moving story," "pet," and "support" are displayed. This information is sent to the server as form data.

[1700] 3. Recognizing the user's emotional state

[1701] The user's real-time emotional state is analyzed through the smartphone's camera and microphone. An emotion analysis API is used to identify emotions such as joy, sadness, and anger from facial expressions and voice. This emotional state is then sent to the server.

[1702] 4. Sentence Generation

[1703] The server sends prompts to a generative AI model (OpenAI GPT-4) using the image analysis results, user-selected tastes and keywords, and the analyzed emotional state. The generative AI model generates high-quality novel-like text based on the given information.

[1704] For example, if a user uploads a picture of their pet, selects "Inspirational Stories," and the emotional state is recognized as "Joy," the prompt might look like this:

[1705] Image: A pet dog smiling and holding a ball

[1706] Taste: Inspirational story

[1707] Emotional state: Joy

[1708] Prompt: "This dog has supported me through some tough times. When he brought me the ball that day with a smile on his face..."

[1709] 5. Edit and save your text

[1710] The generated sentences are presented to the user on the interface. The user can review the generated sentences and edit them as necessary. The edited sentences are sent to the server for storage and used as feedback for future personalized generation results.

[1711] This allows users to generate, edit and save personalized text based on images and emotional states, enhancing their creativity.

[1712] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1713] Step 1:

[1714] Uploading an image

[1715] A user selects an image file from their smartphone and uploads it to the server through a specified interface. The device sends the selected image file to the server as an HTTP request, and the server temporarily stores the received image file. The input here is the image file selected by the user, and the output is the image file temporarily stored on the server.

[1716] Step 2:

[1717] Selecting styles and keywords

[1718] The user inputs the taste of the text (e.g., "moving story," "suspense," etc.) and keywords (e.g., "pet," "adventure," etc.) on the interface. The terminal sends the selected taste and keywords to the server as form data, which the server receives and stores. The input here is the taste and keywords selected by the user, and the output is the taste and keyword data stored on the server.

[1719] Step 3:

[1720] Recognizing the user's emotional state

[1721] The user's real-time emotional state is analyzed through the smartphone's camera and microphone. The device sends the acquired biosignal data to an emotion analysis API, which identifies the emotional state (e.g., joy, sadness, anger). The server receives and stores this emotional state data. The input here is the user's biosignal data (images, voice, etc.), and the output is analyzed emotional state data.

[1722] Step 4:

[1723] Image analysis

[1724] The server uses an image recognition API (e.g., Google Cloud Vision, Amazon Rekognition) to analyze the image uploaded by the user. The main elements in the image (e.g., people, animals, landscapes) and their features are extracted and saved as text data. The input here is the uploaded image file, and the output is the text data of the extracted main elements and features.

[1725] Step 5:

[1726] Sentence generation

[1727] The server sends prompts to a generative AI model (e.g., OpenAI GPT-4) based on the image analysis results, the tastes and keywords selected by the user, and the analyzed emotional state, to generate a sentence. The generated sentence is stored on the server. The inputs here are the image analysis results, tastes and keywords, and the emotional state, and the output is the text data of the generated sentence.

[1728] Step 6:

[1729] Editing and saving the generated text

[1730] The user can check the generated text on the interface and edit it as needed. The terminal sends the edited text to the server, which stores it. The input is the generated text and the user's edits, and the output is the final edited text that has been saved.

[1731] Step 7:

[1732] Sharing and saving text

[1733] Users use the save and share buttons on the interface to save their completed writing and share it on various platforms. When the save button is pressed, the device downloads the final writing to the user's device, and when the SNS share button is pressed, it calls the SNS's sharing API to post the writing. The server saves the final writing in a database for future reference and reuse. The input here is the user's save and share instructions, and the output is the saved writing and shared content.

[1734] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1735] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1736] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1737] [Fourth embodiment]

[1738] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1739] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1740] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1741] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1742] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1743] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1744] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1745] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1746] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1747] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1748] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1749] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1750] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1751] This invention relates to a technology for generating novel-like text from images. This system analyzes images provided by the user and generates text based on the analysis results, according to the taste and keywords specified by the user. Such a system promotes the improvement of writing skills and lowers the barrier to becoming a writer. Furthermore, the generated text can be edited by the user, and a personalization function allows adjustments to suit the user's preferences.

[1752] A natural language description of the system's programmatic processing

[1753] 1. Upload an image

[1754] User:

[1755] Users can select any image file from their device and upload it to the server through a specified interface. For example, users can upload photos of their pets at home or scenery from their travels.

[1756] Device:

[1757] The terminal transmits the selected image file to the server as an HTTP request.

[1758] server:

[1759] The server temporarily stores the received image file for further processing.

[1760] 2. Select a style

[1761] User:

[1762] The interface allows users to select the style of their writing, such as "light novel style," "horror novel style," or "fantasy novel style," and also allows users to enter keywords and themes they want to include in their writing.

[1763] Device:

[1764] The terminal transmits the selected taste and keywords to the server as form data.

[1765] server:

[1766] The server stores the received tastes and keywords and uses them for the next process.

[1767] 3. Image Analysis

[1768] server:

[1769] The server analyzes the uploaded images using an image analysis module, which uses image recognition technology to identify key elements in the image (e.g., people, animals, landscapes) and extract their features.

[1770] 4. Sentence Generation

[1771] server:

[1772] The server provides the analyzed image elements and features, as well as the taste and keywords selected by the user, to the multimodal generative AI. The generative AI generates text that matches the specified taste based on the provided information.

[1773] For example, if you set the theme to fantasy and adventure, the following text will be generated:

[1774] "A young adventurer named Aria wakes up in a forest bathed in dazzling light. She deepens her friendships with the companions she meets along the way and renews her resolve to venture into the unknown."

[1775] 5. Check and edit the text

[1776] User:

[1777] The user can view the generated text in the interface and, if necessary, edit parts of the text to suit their preferences.

[1778] Device:

[1779] The terminal transmits the text edited by the user to the server.

[1780] server:

[1781] The server stores the user's edits to serve as a reference for future personalization and as feedback data to improve the quality of the generated text.

[1782] 6. Save and share your writing

[1783] User:

[1784] Users can use the save and share buttons on the interface to save their completed writing and share it across various platforms.

[1785] Device:

[1786] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[1787] server:

[1788] The server stores the final text in a database as needed, allowing the user to use it comfortably the next time they access the site.

[1789] This embodiment allows users to easily create novel-like texts based on images and edit them to their liking. The generated texts are extremely useful for users who aspire to be writers, helping them improve their writing skills and assist in creative activities.

[1790] The processing flow will be explained below.

[1791] Step 1:

[1792] User:

[1793] The user selects an image file from their device, clicks the upload button in the web interface or app, selects the image, and then presses the send button to upload it to the server.

[1794] Step 2:

[1795] Device:

[1796] The terminal transmits the image file selected by the user to the server as an HTTP request.

[1797] Step 3:

[1798] server:

[1799] The server temporarily stores the received image file and prepares it for processing by the image analysis module.

[1800] Step 4:

[1801] User:

[1802] The user selects the style of the generated text on the interface, such as "light novel style," "horror novel style," or "fantasy novel style," and also inputs keywords or themes (e.g., "adventure" or "friendship") they want to include in the generated text.

[1803] Step 5:

[1804] Device:

[1805] The terminal transmits the selected taste and keywords to the server as form data.

[1806] Step 6:

[1807] server:

[1808] The server stores the received tastes and keywords and launches an image analysis module, which analyzes the received image, identifies key elements in the image (e.g., people, animals, landscapes), and extracts their features.

[1809] Step 7:

[1810] server:

[1811] Next, the server uses the image analysis results and the tastes and keywords specified by the user to provide data to a multimodal generation AI, which then generates text that matches the specified taste based on the provided data.

[1812] Step 8:

[1813] server:

[1814] The generated text is temporarily saved and prepared for presentation to the user.

[1815] Step 9:

[1816] User:

[1817] The user can check the generated text on the interface and, if necessary, edit parts of the text on the interface to suit their preferences.

[1818] Step 10:

[1819] Device:

[1820] The terminal transmits the text edited by the user to the server.

[1821] Step 11:

[1822] server:

[1823] The server stores the edited text by the user and uses it for future personalization and as feedback data to improve the quality of the generated text.

[1824] Step 12:

[1825] User:

[1826] Users can click the save button on the interface to save the completed text, and can also click the share button to share the text on social media.

[1827] Step 13:

[1828] Device:

[1829] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[1830] Step 14:

[1831] server:

[1832] The server stores the final text in a database for future reference and reuse, and when the user returns, it provides personalized results based on the previously saved data.

[1833] Through these steps, users can easily create novel-like texts based on images, and then edit, save, and share them according to their preferences.

[1834] Example 1

[1835] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1836] Currently, technology for generating text from images is rapidly developing, but most systems are limited to simply generating text that describes the content of an image, and their ability to generate novel-like creative writing is limited. In particular, generating personalized text based on user-specified tastes and keywords is difficult and time-consuming. Furthermore, there is a lack of functionality that allows users to easily edit, save, and share the generated text. These challenges have resulted in a significant amount of effort required by users engaged in creative activities.

[1837] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1838] In this invention, the server includes means for inputting an image, means for analyzing the input image and extracting key elements and their features within the image, means for inputting text tastes and keywords specified by the user, means for using a multimodal generative model to generate text using the extracted image elements and features based on the tastes and keywords, means for presenting the generated text to the user and making it editable, and means for saving the generated text and sharing it on various platforms. This allows users to easily generate novel-like text based on their specified tastes and keywords, and to edit, save, and share it.

[1839] The "means for inputting images" is a device that includes an interface and a protocol for users to upload any image file from their own terminal to the server.

[1840] "Means for analyzing an input image and extracting key elements and their features" refers to a device that includes software modules and algorithms for using image recognition technology to identify key elements, such as people, animals, and scenery, in an image and extract their features.

[1841] "Means for inputting the style and keywords of the text specified by the user" refers to a device that allows the user to input the style of the text they desire (e.g., light novel style, horror novel style) and the keywords they want to include through an interface.

[1842] A "means for using a multimodal generative model" is a device that uses a generative AI model (e.g., a generative AI model) to process image and text data as input and generate personalized text based on specified tastes and keywords.

[1843] The "means for presenting the generated text to the user and enabling the user to edit it" is a device that includes an interface and software that displays the generated text to the user and allows the user to edit it.

[1844] "Means for saving generated text and sharing it on various platforms" refers to a device that includes an API and functionality for saving generated text and allowing users to download it or post it on social media or other sharing platforms.

[1845] This invention relates to a system that generates novel-like text from images. The system analyzes images provided by the user and generates text based on the analysis results according to the taste and keywords specified by the user. This promotes the improvement of writing skills and lowers the barrier to becoming a writer. The generated text can also be edited by the user, and a personalization function allows for adjustments to suit the user's preferences.

[1846] First, the user selects an image using their device and uploads it to the server through a specified interface. For example, the user might upload a photo of a pet at home or a landscape from a travel destination. This image file is sent to the server by the device as an HTTP request. The server temporarily stores the received image file for further processing. In this case, the server uses a storage solution such as AWS S3 or Google Cloud Storage.

[1847] Next, the user enters the style of the text (e.g., light novel style, horror novel style, fantasy novel style) and keywords (e.g., "adventure," "friendship," "hero") on the interface. The terminal sends the selected style and keywords as form data to the server, and the server stores the received data. For example, a database system such as MySQL or Firebase is used for this purpose.

[1848] The server analyzes the uploaded images using an image analysis module. Specifically, it uses image recognition technologies such as TensorFlow and OpenCV to identify key elements in the image (e.g., people, animals, landscapes) and extract their features. For example, it uses a TensorFlow image classification model to scan the image and identify key elements (e.g., dogs, mountains, rivers, etc.).

[1849] Next, the server provides the analyzed image elements and features, as well as the taste and keywords selected by the user, to a multimodal generative AI model (e.g., OpenAI's generative AI model).The generative AI model generates text that matches the specified taste based on the provided information.

[1850] For example, if you set the theme to fantasy and adventure, the following text will be generated:

[1851] "A young adventurer named Aria wakes up in a forest bathed in dazzling light. She deepens her friendships with the companions she meets along the way and renews her resolve to venture into the unknown."

[1852] The user can view the generated text on the interface and, if necessary, edit parts of the text to suit their preferences. The edited text is then sent from the device to the server, where it is stored. The stored data is used as a reference for future personalization and as feedback to improve the quality of the generative AI model.

[1853] Finally, the user can save the completed sentence and share it on various platforms. When the save button is pressed, the final sentence is downloaded to the user's device. When the SNS share button is pressed, the SNS sharing API is called and the sentence is posted. The server stores the final sentence in a database as needed, allowing the user to use it comfortably the next time they access the site.

[1854] An example of a specific prompt would be something like this:

[1855] Image: Upload your own image (e.g., a photo of your pet, a travel photo, etc.)

[1856] Style: Fantasy novel-style, adventure theme

[1857] Keywords: "Hero", "Friendship"

[1858] Generated text:

[1859] Waffles, a brave dog, embarks on a quest to save the magical kingdom with his best friend Claire. Together, they overcome many obstacles and rediscover the power of true friendship.

[1860] This system allows users to easily generate novel-like texts based on images and edit them to their liking. The generated texts are extremely useful for aspiring writers, helping them improve their writing skills and their creative activities.

[1861] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1862] Step 1:

[1863] User: The user selects an image file from their device and uploads it to the server through a specified interface. For example, the user may select a photo of a pet at home or a landscape from a travel destination.

[1864] Input: An image file selected by the user.

[1865] How it works: The user selects an image using the file selection dialog and clicks the upload button.

[1866] Output: The selected image file is sent to the next step.

[1867] Step 2:

[1868] Terminal: The terminal sends the selected image file to the server as an HTTP request.

[1869] Input: An image file selected by the user.

[1870] How it works: The device generates an HTML form, converts the selected image file to multipart / form-data format, and creates an HTTP POST request to send to the server.

[1871] Output: The image file is uploaded to the server.

[1872] Step 3:

[1873] Server: The server temporarily stores the received image files for further processing.

[1874] Input: Image file sent from the device.

[1875] How it works: The server parses the incoming HTTP request, extracts the image file, and saves it to storage (e.g., AWS S3 or Google Cloud Storage). After saving, it generates a data path for the next analysis step.

[1876] Output: The path of the image file saved in storage will be sent to the next step.

[1877] Step 4:

[1878] User: The user inputs the style of the text (e.g., light novel style, horror novel style, fantasy novel style) and keywords on the interface.

[1879] Input: Tastes and keywords entered by the user.

[1880] How it works: The user enters the desired taste and keywords using the drop-down menus and text fields.

[1881] Output: The input tastes and keywords are sent to the next step.

[1882] Step 5:

[1883] Terminal: The terminal sends the selected tastes and keywords to the server as form data.

[1884] Input: Tastes and keywords entered by the user.

[1885] Operation: The device converts the input data of tastes and keywords into JSON format and sends it to the server as an HTTP POST request.

[1886] Output: The taste and keyword information is sent to the server.

[1887] Step 6:

[1888] Server: The server stores the received tastes and keywords and uses them for the next process.

[1889] Input: Tastes and keywords entered by the user.

[1890] How it works: The server parses the received JSON data and stores it in a database (e.g., MySQL or Firebase). After saving, it proceeds to the next image analysis step.

[1891] Output: The saved taste and keyword information will be used in the next step.

[1892] Step 7:

[1893] Server: The server uses an image analysis module to analyze the uploaded images.

[1894] Input: The path of the image file retrieved from storage.

[1895] How it works: The server uses image recognition technology from TensorFlow and OpenCV to scan an image, identify key elements (e.g. people, animals, landscapes) and extract their features.

[1896] Output: Key elements and feature data in the image are sent to the next step.

[1897] Step 8:

[1898] Server: The server provides the analyzed image elements and features, as well as user-selected tastes and keywords, to the multimodal generative AI model as input.

[1899] Input: Extracted image elements and feature data, user-selected tastes and keywords.

[1900] How it works: The server inputs this data in JSON format into a generative AI model and sends a request to generate the appropriate sentence.

[1901] Output: The generated sentence is obtained as a response from the generative AI model.

[1902] Step 9:

[1903] User: The user checks the generated text on the interface and edits it as necessary.

[1904] Input: The generated sentence.

[1905] How it works: In the interface that displays the text, the user edits parts of the text as needed.

[1906] Output: The edited text data is sent to the next step.

[1907] Step 10:

[1908] Terminal: The terminal transmits the text edited by the user to the server.

[1909] Input: User edited text.

[1910] Operation: The device converts the edited text data into JSON format and sends it to the server as an HTTP POST request.

[1911] Output: The edited text data is sent to the server.

[1912] Step 11:

[1913] Server: The server stores the user's edits and uses them as a reference for future personalization.

[1914] Input: Edited text data.

[1915] How it works: The server parses the received JSON data and stores it in a database. It also uses it as feedback data to improve the quality of the generative AI model.

[1916] Output: Saved edited text data.

[1917] Step 12:

[1918] User: Users save their completed writing and share it on various platforms.

[1919] Input: Final sentence data.

[1920] How it works: Click the save or share button on the interface to share your text on social media or other platforms.

[1921] Output: The saved text is downloaded to the user's device or posted to a sharing platform.

[1922] Step 13:

[1923] Server: The server stores the final text in a database if necessary, making it easier for the user to use the text the next time they access the site.

[1924] Input: Final sentence data.

[1925] How it works: The server stores the final text data in a database and associates it with the user's profile.

[1926] Output: The saved text data is stored for future use.

[1927] (Application example 1)

[1928] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1929] While there have been technologies that generate text from images, there has been a lack of a way to save the text and easily share it across various platforms. It has also been difficult to personalize the generated text and tune it to the user's preferences. As a result, it has been difficult to provide a personalized creative experience for users.

[1930] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1931] In this invention, the server includes means for inputting an image, means for analyzing the input image and extracting key elements and their characteristics from the image, means for inputting text tastes and keywords specified by the user, means for generating text using the extracted image elements and characteristics based on the tastes and keywords, means for presenting the generated text to the user and making it editable, and means for saving the generated text and sharing it on various platforms. This allows users to easily save and share the generated text and provide a personalized creation experience tailored to their preferences.

[1932] "Means for inputting images" refers to a function that allows a user to select any image from the terminal and upload it to the server through a specified interface.

[1933] "Means for analyzing an image and extracting the main elements and their features within the image" refers to a function that uses image analysis technology to analyze an input image, identify the main elements within it, such as people, animals, and scenery, and extract their features.

[1934] "Means for inputting text style and keywords" is a function that allows the user to input the style and theme of the text to be generated on the interface, as well as the keywords that the user wants to include.

[1935] "Means for generating text using extracted image elements and features" refers to a function in which a multimodal generative AI generates text based on the elements and features obtained through image analysis and the tastes and keywords specified by the user.

[1936] The "means for presenting the generated text to the user and enabling it to be edited" is a function that provides the generated text on an interface so that the user can check it and edit it as necessary.

[1937] "Means for saving the generated text and sharing it on various platforms" refers to a function that allows users to save the final text and share it on platforms such as social media and email.

[1938] Overall system flow

[1939] Users operate the system using a smartphone. Specifically, they upload photographed or existing images to the application and input the text style and keywords. The system then analyzes the images and generates text using a generative AI model. The generated text can be reviewed and edited by the user, and can also be saved and shared.

[1940] Hardware and Software

[1941] Hardware

[1942] Smartphone: The device where the user selects and uploads images to the application.

[1943] Server: The central processing unit that analyzes images and generates text.

[1944] software

[1945] Image Analysis Module: Software for analyzing elements in images using TensorFlow and PyTorch.

[1946] Generative AI model: Using OpenAI's GPT-4 and other models, sentences are generated based on the analysis results and conditions specified by the user.

[1947] Front-end interface: The interface through which users upload images, select settings, and review and edit the generated text.

[1948] SNS sharing API: An API for sharing generated text on the SNS platform of the user's choice.

[1949] Data processing and calculation

[1950] When a user uploads an image from their smartphone, the server first temporarily stores the image. Next, an image analysis module is used to identify the main elements of the uploaded image (e.g., people, animals, landscapes) and extract their characteristics. Next, a generative AI model combines the analysis results with the text to generate a sentence based on the user-specified text style (e.g., "fantasy," "mystery," etc.).

[1951] Examples of concrete examples and prompts

[1952] For example, if a user uploads a photo of a beach at sunset, selects "fantasy" as the text style, and enters "adventure" and "magic" as keywords, the system will input the following prompt sentences into the generative AI model:

[1953] Generates the prompt statement:

[1954] The image shows a beach at sunset. Create a fantasy story using this image. Theme: Adventure, Magic.

[1955] Based on this prompt, the generative AI model might generate a sentence like this:

[1956] Examples of generated sentences:

[1957] "A golden sunset sunk into the beach. The wizard set off towards the boat waiting on the shore, eager for new adventures. The faint light from the beach shone on his magic wand."

[1958] The generated text can be viewed and edited on the interface, and can then be shared via social media or email. Through this series of operations, users can easily manage the generated text and enjoy creative activities.

[1959] In this way, the system of the present invention provides users with a comprehensive solution for generating, editing, and sharing novel-like text from images.

[1960] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1961] Step 1:

[1962] User: Selects an image from the smartphone's photo gallery and uploads it to the application.

[1963] Input: Image files on your smartphone

[1964] Output: Image file upload request to the server

[1965] Terminal: The selected image file is sent to the server as an HTTP request and temporarily stored.

[1966] Input: An image file selected by the user

[1967] Output: Image data transferred to the server

[1968] Server: Temporarily stores received image files and prepares them for the next processing step.

[1969] Input: Image data transferred from the device

[1970] Output: Temporarily saved image file

[1971] Step 2:

[1972] User: Select and input text style and keywords in the application interface.

[1973] Input: Text style (e.g., "fantasy"), keywords (e.g., "adventure", "magic")

[1974] Output: Input data of tastes and keywords

[1975] Device: Sends the selected tastes and keywords to the server.

[1976] Input: User input data of tastes and keywords

[1977] Output: Taste and keyword data sent to the server

[1978] Server: Save the received tastes and keywords and use them for the next process.

[1979] Input: Taste and keyword data sent from the device

[1980] Output: Saved tastes and keyword data

[1981] Step 3:

[1982] Server: Analyzes the uploaded image using an image analysis module. This analysis module uses image recognition techniques such as TensorFlow and PyTorch to identify key elements in the image and extract their features.

[1983] Input: Temporarily saved image file

[1984] Output: Identified key elements and their characteristic data

[1985] Step 4:

[1986] Server: Provides the analyzed image elements and features, as well as user-specified tastes and keywords, to the generative AI model. The generative AI model (e.g., OpenAI's GPT-4) generates prompts based on the provided information and creates full sentences.

[1987] Input: Analysis data (image elements and features), taste, keywords

[1988] Output: Generated sentence

[1989] Step 5:

[1990] Server: Sends the generated text to an interface that presents it to the user so that it can be edited by the user.

[1991] Input: Generated sentence

[1992] Output: Text displayed on the user's terminal

[1993] User: Check the generated text on the interface and edit it as necessary.

[1994] Input: Generated text sent from the server

[1995] Output: Edited text

[1996] Terminal: The user sends the edited results to the server.

[1997] Input: Text data edited by the user

[1998] Output: Edited text data sent to the server

[1999] Step 6:

[2000] Server: Saves edited text and connects to the sharing APIs of various platforms to enable sharing on social media and other platforms.

[2001] Input: Edited text data

[2002] Output: Final text data saved and text shared via the sharing API

[2003] User: Use the save and share buttons to share the generated text on various platforms.

[2004] Input: Final sentence

[2005] Output: Content shared across platforms

[2006] This series of steps creates a system that allows users to easily create, review, edit, save, and share text generated from images.

[2007] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2008] This invention combines technology for generating novel-like text from images with an emotion engine, enabling text generation that responds to the user's emotional state. This system analyzes images provided by the user and, based on the analysis results, generates and adjusts text according to the user's emotional state, incorporating tastes and keywords specified by the user. This system promotes the improvement of writing skills and lowers the barrier to becoming a writer. Furthermore, the generated text can be edited by the user, and a personalization function allows adjustments to the user's preferences.

[2009] A natural language description of the system's programmatic processing

[2010] 1. Upload an image

[2011] User:

[2012] Users can select any image file from their device and upload it to the server through a specified interface. For example, users can upload photos of their pets at home or scenery from their travels.

[2013] Device:

[2014] The terminal transmits the selected image file to the server as an HTTP request.

[2015] server:

[2016] The server temporarily stores the received image file for further processing.

[2017] 2. Select a style

[2018] User:

[2019] The interface allows users to select the style of their writing, such as "light novel style," "horror novel style," or "fantasy novel style," and also allows users to enter keywords and themes they want to include in their writing.

[2020] Device:

[2021] The terminal transmits the selected taste and keywords to the server as form data.

[2022] server:

[2023] The server stores the received tastes and keywords and launches an image analysis module, which analyzes the received image, identifies key elements in the image (e.g., people, animals, landscapes), and extracts their features.

[2024] 3. Emotional Recognition

[2025] User:

[2026] Users provide their emotional state to the system through input devices such as a webcam and microphone, which is collected through analysis of facial expressions, voice, and keyboard typing speed.

[2027] server:

[2028] The server analyzes the user's emotional state using an emotion engine, which identifies emotions such as joy, sadness, and anger based on biosignal data provided by the user.

[2029] 4. Sentence Generation

[2030] server:

[2031] The server provides data to a multimodal generation AI using the image analysis results, user-specified tastes and keywords, and the analyzed emotional state. Based on the information provided, the generation AI generates text that matches the user's emotional state in addition to the specified tastes.

[2032] For example, if the user selects "fantasy novel-style" and "adventure" as keywords, and the emotional state is recognized as "joy," the following sentence will be generated:

[2033] "In a meadow bathed in dazzling light, the little adventurer Leah stood up with a smile on her face. A fun-filled adventure with new friends awaited her."

[2034] 5. Check and edit the text

[2035] User:

[2036] The user can review the generated text on the interface and, if necessary, edit parts of the text to suit their preferences.

[2037] Device:

[2038] The terminal transmits the text edited by the user to the server.

[2039] server:

[2040] The server stores the text edited by the user and uses it as feedback data for future personalization and to improve the quality of the generated text.

[2041] 6. Save and share your writing

[2042] User:

[2043] Users can use the save and share buttons on the interface to save their completed writing and share it across various platforms.

[2044] Device:

[2045] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[2046] server:

[2047] The server stores the final text in a database for future reference and reuse, and when the user returns, it provides personalized results based on the previously saved data.

[2048] This embodiment allows users to easily create novel-like texts based on images, and then edit, save, and share them in a more personalized way that reflects their emotional state, further enriching the user's creative experience and helping to improve their writing skills and creative activities.

[2049] The processing flow will be explained below.

[2050] Step 1:

[2051] User:

[2052] Users simply select an image file from their device and click the upload button on the web interface or app. For example, a user can select a photo of a landscape from a travel destination or a pet, and then press the send button to upload it to the server.

[2053] Step 2:

[2054] Device:

[2055] The selected image file is sent to the server as an HTTP request.

[2056] Step 3:

[2057] server:

[2058] The received image file is temporarily saved and the image analysis module is prepared.

[2059] Step 4:

[2060] User:

[2061] The interface allows you to select the style of your writing, such as "light novel style," "horror novel style," or "fantasy novel style," and also enter keywords or themes you want to include in your writing (e.g., "adventure" or "friendship").

[2062] Step 5:

[2063] Device:

[2064] The selected tastes and keywords are sent to the server as form data.

[2065] Step 6:

[2066] server:

[2067] The received tastes and keywords are saved and the image analysis module is launched. The image analysis module analyzes the received image, identifies the main elements in the image (e.g., people, animals, landscapes), and extracts their features (color, shape, position, etc.).

[2068] Step 7:

[2069] User:

[2070] Through the interface, users provide their emotional state via input devices such as a webcam and microphone. Users input emotional data into the system through facial expressions, voice, and keyboard input speed.

[2071] Step 8:

[2072] server:

[2073] An emotion engine is activated to analyze the user's emotional state, and the emotion engine identifies emotions such as joy, sadness, anger, etc. based on the biosignal data provided by the user.

[2074] Step 9:

[2075] server:

[2076] The image analysis results, the user's emotional data, and the user's specified tastes and keywords are provided to a multimodal generation AI, which then generates text based on the given information, in addition to the specified tastes, that corresponds to the user's emotional state.

[2077] Step 10:

[2078] server:

[2079] The generated text is temporarily saved and prepared for presentation to the user.

[2080] Step 11:

[2081] User:

[2082] Review the generated text in the interface and, if necessary, edit parts of the text to suit your preferences.

[2083] Step 12:

[2084] Device:

[2085] The text edited by the user is sent to the server.

[2086] Step 13:

[2087] server:

[2088] User edits are saved and used for future personalization, and also as feedback data to improve the quality of generated text.

[2089] Step 14:

[2090] User:

[2091] To save your completed essay, click the Save button on the interface. You can also share your essay on social media by clicking the Share button.

[2092] Step 15:

[2093] Device:

[2094] When the save button is pressed, the final text is downloaded to the user's device. When the SNS share button is pressed, the SNS sharing API is called and the text is posted.

[2095] Step 16:

[2096] server:

[2097] The final text is stored in a database for future reference and reuse, and when the user returns, the previously stored data is used to provide personalized results.

[2098] Through these steps, users can easily generate novel-like texts based on images, and then edit, save, and share them with personalized content that reflects the user's emotional state. This enriches the user's creative experience, improves their writing skills, and helps with creative activities.

[2099] Example 2

[2100] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2101] While conventional text generation systems can generate text that takes into account user-specified tastes and keywords, they face the problem of difficulty in generating personalized text that reflects the user's emotional state. Furthermore, they lack the functionality to edit and adjust generated text to suit the user's preferences, limiting the user's creative experience. Furthermore, the complicated process of sharing and saving generated text reduces user convenience.

[2102] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2103] In this invention, the server includes: means for inputting an image; means for analyzing the input image and extracting key elements and their features within the image; means for inputting text tastes and keywords specified by the user; means for recognizing the user's emotional state; means for generating and adjusting text based on the recognized user's emotional state; means for presenting the generated text to the user and making it editable; means for analyzing biosignal data such as the user's facial expression, voice, and keyboard input speed; means for generating text based on the user's emotional state based on given information using a multimodal generative AI model; means for providing means for sharing the generated text on various platforms and for invoking a sharing API based on the user's instructions. This allows users to generate personalized text based on information extracted from the image and their own emotional state, enabling a richer creative experience through editing and sharing.

[2104] "Means for inputting images" is a function that allows a user to upload an image file selected from their own terminal to the system.

[2105] "Means for extracting the main elements and their features in an image" refers to a function that identifies the main elements, such as people, animals, and scenery, from the analyzed image and extracts these features as data.

[2106] The "means for inputting text style and keywords" is an input interface that allows the user to specify to the system the style of text to be generated and the content to be included.

[2107] The "means for recognizing the user's emotional state" is a function for analyzing biosignal data such as the user's facial expression, voice, keyboard input speed, etc., and identifying the user's current emotional state.

[2108] "Means for generating and adjusting sentences" refers to a function that uses specific algorithms and AI technology based on collected data to generate sentences that conform to specified tastes and keywords, and further adjusts them according to the user's emotional state.

[2109] The "means for presenting the generated text and making it editable" is an interface that displays the generated text to the user and allows the user to modify and edit parts of the text as necessary.

[2110] A "multimodal generative AI model" is an advanced artificial intelligence technology that integrates and processes information from multiple data sources (e.g., images, text, and emotional data) to generate natural language sentences.

[2111] The "means for analyzing biosignal data" is a function for analyzing biosignals such as facial expressions and voice obtained from the user and determining the emotional state from that data.

[2112] "Sharing API" means an application program interface used to share generated text across various platforms.

[2113] "Means for calling a sharing API based on user instructions" refers to a function that automatically calls the appropriate sharing API and posts text to a specified platform when a user performs an operation such as pressing a share button on the interface.

[2114] MODE FOR CARRYING OUT THE INVENTION

[2115] This system extracts key elements from an image provided by a user and generates text based on the user's preferences and keywords. It also has the ability to personalize, edit, save, and share text based on the user's emotional state. This system is implemented using the following hardware and software:

[2116] 1. Upload an image

[2117] User:

[2118] Users simply select an image file from their device and press the upload button through the system interface. For example, users can upload a photo of their pet at home or a landscape from a travel destination.

[2119] Device:

[2120] The device sends the selected image file as an HTTP request to the server, where the image data is temporarily stored in memory and sent to the server in the appropriate encoding format.

[2121] server:

[2122] The server temporarily stores the received image file in a secure storage location, and after storing it, assigns a unique identifier to the image file and uses that identifier for further processing.

[2123] 2. Select a style

[2124] User:

[2125] Users select the style of their writing (e.g., "fantasy novel-style") from the options provided on the interface, and then enter the keywords and themes they want to include in the writing.

[2126] Device:

[2127] The device sends the tastes and keywords selected by the user to the server as form data, where the input data is encoded in JSON format or similar.

[2128] server:

[2129] The server stores the received tastes and keywords in a database and then launches an image analysis module, which reads pre-saved image files, identifies key elements (e.g., people, animals, landscapes), and extracts their features.

[2130] 3. Emotional Recognition

[2131] User:

[2132] Users provide their emotional state to the system using a webcam or microphone, which is collected through facial expression recognition, voice tone analysis, keyboard input speed, etc.

[2133] server:

[2134] The server then analyzes the received biometric data using an emotion engine to identify specific emotions such as joy, sadness, anger, etc. The emotion analysis results are then used in the subsequent sentence generation process.

[2135] 4. Sentence Generation

[2136] server:

[2137] The server provides data to a multimodal generative AI model based on the results of image analysis, user-specified tastes and keywords, and the analyzed emotional state.

[2138] The generative AI model generates sentences based on the given information, in addition to the specified taste, according to the user's emotional state.

[2139] Examples:

[2140] If the user selects the keywords "fantasy novel-style" and "adventure" and the emotional state is recognized as "joy," the following sentence will be generated:

[2141] "In a meadow bathed in dazzling light, the little adventurer Leah stood up with a smile on her face. A fun-filled adventure with new friends awaited her."

[2142] 5. Check and edit the text

[2143] User:

[2144] The user can review the generated text in the interface and, if necessary, edit parts of the text to suit their preferences.

[2145] Device:

[2146] The terminal sends the text edited by the user to the server, again in JSON format or similar.

[2147] server:

[2148] The server stores the text edited by the user and uses it as feedback data for future personalization and system improvements.

[2149] 6. Save and share your writing

[2150] User:

[2151] Users can use the save and share buttons on the interface to save their completed writing and share it on various platforms.

[2152] Device:

[2153] When the save button is pressed, the device downloads the final text to the user's device. When the SNS share button is pressed, the device calls the SNS sharing API and posts the text.

[2154] server:

[2155] The server stores the final text in a database for future reference and reuse, and when the user returns, it provides personalized results based on the previously saved data.

[2156] Prompt Sentence Examples

[2157] Below are some example prompts to input to a generative AI model:

[2158] "Right now, your emotional state is 'joy.' Create a fantasy-style adventure story based on a landscape photo. The keywords are 'light,' 'adventure,' and 'friends.'"

[2159] The system described above allows users to easily create novel-like texts based on images, and then edit, save, and share them with personalized content that reflects their emotional state. This will further enrich users' creative experiences, improve their writing skills, and help them with creative activities.

[2160] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2161] Step 1: Upload an image

[2162] input:

[2163] Users simply select any image file from their device and press the upload button through the system interface.

[2164] Specific behavior:

[2165] The user selects the image file of their choice through the specified interface and performs the upload operation, at which point the image file is temporarily stored on the user's device.

[2166] output:

[2167] The device sends the selected image file to the server as an HTTP request.

[2168] Step 2: Receive and save the image file

[2169] input:

[2170] Image files sent from the device

[2171] Specific behavior:

[2172] The server temporarily stores the received image file in storage and assigns a unique identifier (ID) to the file.

[2173] output:

[2174] The image file is given a unique identifier and stored in storage, which is used in the next processing step.

[2175] Step 3: Selecting Styles and Keywords

[2176] input:

[2177] The user specifies the style and keywords of the text.

[2178] Specific behavior:

[2179] The user selects the style of the text (e.g., "fantasy novel-style") from the options provided on the interface and enters keywords and themes.

[2180] output:

[2181] The device sends the selected taste and entered keywords as form data to the server. The sent data is encoded in JSON format.

[2182] Step 4: Begin image analysis

[2183] input:

[2184] Taste and keyword data, as well as identifiers of saved image files

[2185] Specific behavior:

[2186] The server receives the taste and keyword data and launches an image analysis module, which reads the saved image file, identifies key elements (e.g., people, animals, landscapes, etc.), and extracts their features.

[2187] output:

[2188] Data on the extracted image elements and their features are generated and used in the next processing step.

[2189] Step 5: Recognizing your emotional state

[2190] input:

[2191] Biosignal data collected from users (facial expressions, voice, keyboard typing speed, etc.)

[2192] Specific behavior:

[2193] Users provide their emotional state to the system using a webcam, microphone, etc. Emotional data is collected.

[2194] output:

[2195] The terminal transmits the collected biosignal data to the server.

[2196] Step 6: Analyze the sentiment data

[2197] input:

[2198] Collected biosignal data

[2199] Specific behavior:

[2200] The server analyzes the received biosignal data using an emotion engine to identify the user's emotional state (joy, sadness, anger, etc.).

[2201] output:

[2202] Parsed emotional state data is generated and used in the next processing step.

[2203] Step 7: Sentence generation

[2204] input:

[2205] Image analysis results, tastes and keywords, analyzed emotional states

[2206] Specific behavior:

[2207] The server provides the above data to a multimodal generative AI model, which uses the information to generate text that reflects the user's emotional state and the user's specified taste.

[2208] output:

[2209] The generated text is temporarily stored on the server.

[2210] Examples:

[2211] If the user sets the keywords "fantasy novel style" and "adventure" and the emotional state is recognized as "joy," the following sentence will be generated: "In a meadow bathed in dazzling light, the little adventurer Leah stood up with a smile on her face. A fun adventure with new friends awaited her."

[2212] Step 8: Presenting and editing the generated text

[2213] input:

[2214] Generated sentence data

[2215] Specific behavior:

[2216] The server sends the generated text to the interface, allowing the user to check the text content.

[2217] output:

[2218] The generated text is presented on the interface, and the user can view and edit it.

[2219] Step 9: Save your edits

[2220] input:

[2221] Text data edited by the user

[2222] Specific behavior:

[2223] The user can check the generated text on the interface, make corrections as necessary, and then click the save button when editing is complete.

[2224] output:

[2225] The terminal transmits the text data edited by the user to the server.

[2226] Step 10: Save edits and feedback

[2227] input:

[2228] Edited text data

[2229] Specific behavior:

[2230] The server stores the text data edited by the user and records it as feedback data to be used for future personalization and system improvement.

[2231] output:

[2232] The edited text data is stored in a database.

[2233] Step 11: Save and share your writing

[2234] input:

[2235] Completed sentence data

[2236] Specific behavior:

[2237] The user saves or shares the text using the save or share buttons on the interface.

[2238] output:

[2239] When the save button is pressed, the completed sentence is downloaded to the user's device. When the share button is pressed, the appropriate SNS sharing API is called and the sentence is posted.

[2240] Step 12: Save the final data

[2241] input:

[2242] Shared or stored text data

[2243] Specific behavior:

[2244] The server stores the final text in a database for future reference and reuse.

[2245] output:

[2246] The final text data is securely stored in a database and can be reused or referenced as needed.

[2247] In this way, this system allows users to easily create novel-like texts based on images, and then edit, save, and share them with personalized content that reflects their emotional state. This will further enrich the user's creative experience, improve their writing skills, and help with creative activities.

[2248] (Application example 2)

[2249] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2250] Conventional text generation systems generate text without considering the user's emotional state, which can result in text that does not match the user's emotions. This makes it difficult to generate text that the user can truly empathize with. Furthermore, conventional systems do not fully utilize the information in the image, limiting the quality of the generated text. This limits the user's creative experience.

[2251] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2252] In this invention, the server includes means for inputting an image, means for analyzing the input image and extracting the main elements and features of the image, means for inputting the taste and keywords of a sentence specified by the user, means for recognizing the analyzed emotional state of the user, means for generating a sentence using the extracted elements and features of the image based on the taste and keywords and the analyzed emotional state, and means for presenting the generated sentence to the user and making it editable. This makes it possible to generate sentences that correspond to the emotional state of the user and to generate high-quality sentences that the user can empathize with.

[2253] The "means for inputting images" refers to a method by which a user can provide a digital image to the system, and is a means by which image files can be uploaded using a smartphone or computer.

[2254] "Means for analyzing an input image and extracting the main elements and their features within the image" refers to a method that uses image recognition technology to identify the main elements of people, objects, landscapes, etc. that appear in an image and extract their features as data.

[2255] The "means for inputting the style and keywords of the text specified by the user" is an interface that allows the user to input the style of the text to be generated, the themes they wish to include, and keywords into the system.

[2256] "Means for recognizing an analyzed emotional state of a user" refers to a method that utilizes technology to identify a user's current emotional state through facial expression or voice analysis of the user.

[2257] "Means for generating sentences using extracted image elements and features based on the tastes and keywords and the analyzed emotional state" refers to a technology for generating sentences by incorporating tastes and keywords specified by the user and the recognized emotional state, and utilizing elements and features extracted from the image.

[2258] "Means for presenting the generated text to the user and making it editable" refers to an interface that displays the generated text to the user and allows the user to freely modify and edit the text.

[2259] The present invention relates to a system for generating sentences in accordance with the emotional state of a user, and a specific implementation method thereof will be described below.

[2260] Hardware and Software Configuration

[2261] This system includes an image input means, an image analysis means, a taste and keyword input means, an emotional state recognition means, a sentence generation means, and a means for making the generated sentences editable. The main hardware and software used are as follows:

[2262] Smartphones: iPhone, Android devices, etc.

[2263] Image recognition APIs: Google Cloud Vision, Amazon Rekognition.

[2264] Emotion analysis API: Microsoft Azure Emotion API, Affectiva.

[2265] Generative AI model: OpenAI GPT-4.

[2266] Database: Firebase, SQLite.

[2267] Processing flow

[2268] 1. Upload an image

[2269] Users select any image from their smartphone and use the app's interface to upload it to the server, which temporarily stores the received image for subsequent analysis.

[2270] 2. Selecting the theme and keywords

[2271] The user selects or inputs the text's style and keywords on the interface. For example, options such as "moving story," "pet," and "support" are displayed. This information is sent to the server as form data.

[2272] 3. Recognizing the user's emotional state

[2273] The user's real-time emotional state is analyzed through the smartphone's camera and microphone. An emotion analysis API is used to identify emotions such as joy, sadness, and anger from facial expressions and voice. This emotional state is then sent to the server.

[2274] 4. Sentence Generation

[2275] The server sends prompts to a generative AI model (OpenAI GPT-4) using the image analysis results, user-selected tastes and keywords, and the analyzed emotional state. The generative AI model generates high-quality novel-like text based on the given information.

[2276] For example, if a user uploads a picture of their pet, selects "Inspirational Stories," and the emotional state is recognized as "Joy," the prompt might look like this:

[2277] Image: A pet dog smiling and holding a ball

[2278] Taste: Inspirational story

[2279] Emotional state: Joy

[2280] Prompt: "This dog has supported me through some tough times. When he brought me the ball that day with a smile on his face..."

[2281] 5. Edit and save your text

[2282] The generated sentences are presented to the user on the interface. The user can review the generated sentences and edit them as necessary. The edited sentences are sent to the server for storage and used as feedback for future personalized generation results.

[2283] This allows users to generate, edit and save personalized text based on images and emotional states, enhancing their creativity.

[2284] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2285] Step 1:

[2286] Uploading an image

[2287] A user selects an image file from their smartphone and uploads it to the server through a specified interface. The device sends the selected image file to the server as an HTTP request, and the server temporarily stores the received image file. The input here is the image file selected by the user, and the output is the image file temporarily stored on the server.

[2288] Step 2:

[2289] Selecting styles and keywords

[2290] The user inputs the taste of the text (e.g., "moving story," "suspense," etc.) and keywords (e.g., "pet," "adventure," etc.) on the interface. The terminal sends the selected taste and keywords to the server as form data, which the server receives and stores. The input here is the taste and keywords selected by the user, and the output is the taste and keyword data stored on the server.

[2291] Step 3:

[2292] Recognizing the user's emotional state

[2293] The user's real-time emotional state is analyzed through the smartphone's camera and microphone. The device sends the acquired biosignal data to an emotion analysis API, which identifies the emotional state (e.g., joy, sadness, anger). The server receives and stores this emotional state data. The input here is the user's biosignal data (images, voice, etc.), and the output is analyzed emotional state data.

[2294] Step 4:

[2295] Image analysis

[2296] The server uses an image recognition API (e.g., Google Cloud Vision, Amazon Rekognition) to analyze the image uploaded by the user. The main elements in the image (e.g., people, animals, landscapes) and their features are extracted and saved as text data. The input here is the uploaded image file, and the output is the text data of the extracted main elements and features.

[2297] Step 5:

[2298] Sentence generation

[2299] The server sends prompts to a generative AI model (e.g., OpenAI GPT-4) based on the image analysis results, the tastes and keywords selected by the user, and the analyzed emotional state, to generate a sentence. The generated sentence is stored on the server. The inputs here are the image analysis results, tastes and keywords, and the emotional state, and the output is the text data of the generated sentence.

[2300] Step 6:

[2301] Editing and saving the generated text

[2302] The user can check the generated text on the interface and edit it as needed. The terminal sends the edited text to the server, which stores it. The input is the generated text and the user's edits, and the output is the final edited text that has been saved.

[2303] Step 7:

[2304] Sharing and saving text

[2305] Users use the save and share buttons on the interface to save their completed writing and share it on various platforms. When the save button is pressed, the device downloads the final writing to the user's device, and when the SNS share button is pressed, it calls the SNS's sharing API to post the writing. The server saves the final writing in a database for future reference and reuse. The input here is the user's save and share instructions, and the output is the saved writing and shared content.

[2306] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2307] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2308] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2309] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2310] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2311] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2312] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2313] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2314] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2315] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2316] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2317] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2318] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2319] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2320] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2321] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2322] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2323] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2324] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2325] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2326] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2327] The following is further disclosed regarding the above embodiment.

[2328] (Claim 1)

[2329] A means for inputting an image;

[2330] A means for analyzing an input image and extracting key elements and their features from the image;

[2331] A means for inputting the taste and keywords of a sentence designated by a user;

[2332] A means for generating a sentence using the extracted image elements and features based on the taste and keywords;

[2333] The system includes a means for presenting the generated text to a user and enabling the user to edit the text.

[2334] (Claim 2)

[2335] 2. The system according to claim 1, further comprising means for personalizing the generated text based on the tastes and keywords, and tuning it in accordance with the user's preferences.

[2336] (Claim 3)

[2337] 2. The system according to claim 1, wherein the image analysis means includes means for identifying key elements in an image, such as people, animals, and scenery, using image recognition technology.

[2338] "Example 1"

[2339] (Claim 1)

[2340] A means for inputting an image;

[2341] A means for analyzing an input image and extracting key elements and their features from the image;

[2342] A means for inputting the taste and keywords of a sentence designated by a user;

[2343] a means for using a multimodal generative model to generate text using extracted image elements and features based on the tastes and keywords;

[2344] means for presenting the generated text to a user and enabling the user to edit the text;

[2345] A way to save the generated text and share it on various platforms,

[2346] A system including:

[2347] (Claim 2)

[2348] 2. The system according to claim 1, further comprising means for personalizing the generated text based on the tastes and keywords, and tuning it in accordance with the user's preferences.

[2349] (Claim 3)

[2350] 2. The system according to claim 1, wherein the image analysis means includes means for identifying key elements in an image, such as people, animals, and scenery, using image recognition technology.

[2351] "Application Example 1"

[2352] (Claim 1)

[2353] A means for inputting an image;

[2354] A means for analyzing an input image and extracting key elements and their features from the image;

[2355] A means for inputting the taste and keywords of a sentence designated by a user;

[2356] A means for generating a sentence using the extracted image elements and features based on the taste and keywords;

[2357] means for presenting the generated text to a user and enabling the user to edit the text;

[2358] A system that includes a means to save the generated text and share it across various platforms.

[2359] (Claim 2)

[2360] 2. The system according to claim 1, further comprising means for personalizing the generated text based on the tastes and keywords, and tuning it in accordance with the user's preferences.

[2361] (Claim 3)

[2362] 2. The system according to claim 1, wherein the image analysis means includes means for identifying key elements in an image, such as people, animals, and scenery, using image recognition technology.

[2363] "Example 2: Combining Emotion Engines"

[2364] (Claim 1)

[2365] A means for inputting an image;

[2366] A means for analyzing an input image and extracting key elements and their features from the image;

[2367] A means for inputting the taste and keywords of a sentence designated by a user;

[2368] A means for generating a sentence using the extracted image elements and features based on the taste and keywords;

[2369] means for recognizing the emotional state of a user;

[2370] means for generating and adjusting sentences based on the recognized emotional state of the user;

[2371] The system includes a means for presenting the generated text to a user and enabling the user to edit the text.

[2372] (Claim 2)

[2373] 2. The system according to claim 1, further comprising means for personalizing the generated text based on the tastes and keywords, and tuning it in accordance with the user's preferences.

[2374] (Claim 3)

[2375] 2. The system according to claim 1, wherein the image analysis means includes means for identifying key elements in an image, such as people, animals, and scenery, using image recognition technology.

[2376] (Claim 4)

[2377] 2. The system of claim 1, wherein the means for recognizing the user's emotional state includes means for analyzing biosignal data such as the user's facial expression, voice, keyboard input speed, etc.

[2378] (Claim 5)

[2379] 2. The system according to claim 1, wherein the sentence generation means includes means for generating sentences according to the emotional state of the user based on given information using a multimodal generative AI model.

[2380] (Claim 6)

[2381] The system according to claim 1, further comprising means for providing an interface for allow...

Claims

1. A means for inputting an image; A means for analyzing an input image and extracting key elements and their features from the image; A means for inputting the taste and keywords of a sentence designated by a user; A means for generating a sentence using the extracted image elements and features based on the taste and keywords; The system includes a means for presenting the generated text to a user and enabling the user to edit the text.

2. 2. The system according to claim 1, further comprising means for personalizing the generated text based on the taste and keywords, and tuning it in accordance with the user's preferences.

3. The system according to claim 1 , wherein the image analysis means includes means for identifying key elements in the image, such as people, animals, and scenery, using image recognition technology.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A