system

The system addresses the challenge of maintaining engaging social media content by converting voice input into text, analyzing, and automating the posting process, reducing user effort and enhancing content diversity.

JP2026016232APending Publication Date: 2026-02-03SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024117322
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Users face challenges in maintaining frequent and diverse social media posts, as recording daily events and compiling them into engaging content is time-consuming and requires effort, often leading to monotonous content.

Method used

A system that converts user voice input into text data, analyzes it, generates manuscripts for social media, and automates the posting process, including tagging, keyword extraction, and scheduling.

Benefits of technology

Reduces user burden and enhances content diversity by automating the social media posting process from voice input to posting, allowing users to easily record and share daily events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016232000001_ABST
    Figure 2026016232000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for inputting a daily event spoken by a user to a terminal as voice data; means for converting the voice data into text data; means for storing the text data in a database; means for analyzing the text data in the database and generating a manuscript for SNS contribution; means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit the manuscript; and means for contributing the manuscript confirmed and edited by the user to a designated SNS platform.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, the use of social networking services (SNS) has become a part of everyday activities. However, maintaining frequent posts and diverse content can be a burden for many users. In particular, keeping a record of daily events and compiling them into engaging posts amid busy lives requires time and effort. Users also need a convenient way to record their daily lives as a personal record to look back on later. Furthermore, posts tend to be monotonous, and it is difficult to keep the content attractive and diverse. [Means for solving the problem]

[0005] To solve the above-mentioned problems, the present invention provides a system that includes a means for inputting everyday events spoken by a user into a terminal as voice data and converting the voice data into text data, a means for saving the converted text data in a database, a means for analyzing the saved text data and generating a manuscript for posting to an SNS, a means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, and a means for posting the confirmed and edited manuscript to a specified SNS platform.Furthermore, by including a means for tagging and extracting keywords from the everyday events input by the user and a means for setting a posting schedule for the generated manuscript at a time specified by the user, the system can reduce the burden on the user and increase the diversity of content.

[0006] "User" refers to the end user who operates the system and inputs daily events.

[0007] "Terminal" refers to a device that is directly operated by the user and that performs voice input, sends and receives voice data, displays text data, etc.

[0008] "Voice data" refers to data recorded on a device as audio signals containing what a user says.

[0009] "Text data" refers to data that has been converted from audio data into text format.

[0010] A "voice recognition engine" refers to software or hardware that analyzes voice data and converts it into text data.

[0011] "Database" refers to a system for storing and managing converted text data.

[0012] "Tagging" refers to the process of adding related keywords and category information to text data.

[0013] "Keyword extraction" refers to the process of extracting important words and phrases from text data.

[0014] "Generative AI" refers to artificial intelligence that analyzes text data in a database and generates manuscripts for posting on social media.

[0015] "Manuscript" refers to text content generated for posting on social media.

[0016] "SNS Platform" refers to online services and applications that provide social networking services.

[0017] "API" stands for Application Programming Interface and refers to an interface for communication between software programs.

[0018] "Posting schedule" refers to the function that allows users to set the timing and frequency of social media posts.

[0019] "Confirmation and editing interface" refers to the screen and operation means that allow the user to confirm the generated manuscript and edit its contents. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0022] First, the terms used in the following description will be explained.

[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0028] [First embodiment]

[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0041] This invention is a system that allows users to talk about everyday events, accumulates them in a database, and analyzes each piece of data to generate drafts for posting to social media. This system automates the entire process from user input to posting to social media, significantly reducing the burden on users.

[0042] System configuration and operation

[0043] 1. User voice input

[0044] The user talks to the device about everyday events, such as, "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner."

[0045] 2. Audio data conversion

[0046] The device records the user's voice and sends the voice data to the server. The server uses a speech recognition engine to convert the voice data into text data. The converted text data is sent back to the device and displayed on the screen. The user can check the text and make corrections if necessary.

[0047] 3. Saving to the database

[0048] The device sends the text data to the server, which then adds a date and category (e.g., work, eating out) to the text data and stores it in a database. The stored data is then organized using natural language processing such as tagging and keyword extraction.

[0049] 4. Creating a manuscript for posting on social media

[0050] When the posting timing specified by the user (e.g., every day, or when a specific event occurs) arrives, the server uses generative AI to generate a script for posting to SNS based on the text data in the database. This script is sent to the device in a format that the user can easily understand.

[0051] 5. User confirmation and editing

[0052] The device displays the generated script for posting to social media to the user and provides an interface for reviewing and editing as necessary. For example, it could be displayed in the form of "Yesterday, my big presentation was a success! Afterwards, I enjoyed a delicious Italian meal with friends. It was a fulfilling day."

[0053] 6. Posting to social media

[0054] Once the user has finished checking and editing, they confirm the post. The device then receives this and sends the confirmed manuscript to the server. The server then calls the API of each SNS platform and executes the post. The success or failure of the post is again notified to the device and conveyed to the user.

[0055] Specific examples

[0056] For example, a user might say: "Today I went on a picnic with my family and had a great time. Then I saw a new movie."

[0057] 1. The user speaks.

[0058] 2. The device records the audio and sends it to the server.

[0059] 3. The server uses a speech recognition engine to convert this into text data: "Today I went on a picnic with my family and had a great time. Then I saw a new movie."

[0060] 4. Display the text data on the terminal.

[0061] 5. The user checks, corrects, and approves.

[0062] 6. The device sends the text data to the server, where it is tagged, keywords are extracted, and the data is stored in a database.

[0063] 7. At the specified time, the server generates a script to post to social media: "Yesterday, I went on a picnic with my family. It was a lot of fun. Afterwards, I also saw a new movie."

[0064] 8. The terminal displays the generated manuscript, which the user can check and edit.

[0065] 9. The user submits the finalized manuscript.

[0066] 10. The server posts to the SNS platform and notifies the device of the results.

[0067] This system allows users to easily and efficiently record everyday events and post them to social media.

[0068] The processing flow will be explained below.

[0069] Step 1:

[0070] The user talks about everyday events into the device, inputting content such as "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner" through the voice input interface.

[0071] Step 2:

[0072] The device records the user's voice and sends the captured voice data to the server via the voice input interface.

[0073] Step 3:

[0074] The server receives the voice data, sends it to the voice recognition engine, and converts it into text data.

[0075] Step 4:

[0076] The server returns the converted text data to the device. The returned text data is displayed as "I had a big presentation at work today, and it was a success. For dinner, I went out to eat Italian food with a friend."

[0077] Step 5:

[0078] The terminal displays the text data to the user and asks for confirmation, after which the user can check the text data and correct it if necessary.

[0079] Step 6:

[0080] The user confirms and modifies the text data by pressing the "Confirm" button.

[0081] Step 7:

[0082] The device sends the text data that has been confirmed to the server, which then adds the date and category (e.g., work, eating out) to the text data and stores it in a database.

[0083] Step 8:

[0084] The server performs natural language processing on the stored data, including tagging and keyword extraction.

[0085] Step 9:

[0086] When the designated posting time arrives (e.g., every day, when a specific event occurs), the server uses a generation AI based on the data in the database to generate a script for posting on social media.

[0087] Step 10:

[0088] The server sends the generated manuscript to the terminal, which displays it to the user and prompts them to check and edit it.

[0089] Step 11:

[0090] The user can check and edit the displayed manuscript. If necessary, press the "Submit" button to finalize the manuscript.

[0091] Step 12:

[0092] The device sends the finalized manuscript to the server, which then calls the API of each SNS platform and posts the manuscript.

[0093] Step 13:

[0094] The server checks whether the posting was successful or not, and sends the result back to the device to notify the user (e.g., "Posting was successful!").

[0095] Example 1

[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0097] In modern society, posting to social networking sites has become a part of users' daily lives, but posting easily can be difficult given their busy schedules. The effort of manually entering text and the burden of thinking up appropriate phrases can be stressful for users. Furthermore, managing the consistency and timing of social networking posts can be difficult, making effective communication difficult. The present invention aims to solve these problems by providing a system that automates social networking posts and significantly reduces the burden on users.

[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0099] In this invention, the server includes means for inputting daily events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for storing the text data in a database, means for analyzing the text data in the database and generating a manuscript for posting to an SNS, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, means for posting the manuscript confirmed and edited by the user to a specified SNS platform, and means for creating prompt sentences using a generative AI model to generate post content in an automated process, thereby enabling users to easily record daily events by voice and automatically post them to an SNS.

[0100] A "user" is someone who uses the system to input everyday events by voice and post them on social media.

[0101] A "terminal" is a device used by a user to perform voice input, such as a smartphone or tablet.

[0102] "Voice data" refers to digital data that is a recording of what a user says to a device.

[0103] "Text data" refers to data of character information converted from voice data using voice recognition technology.

[0104] A "database" is an information system for storing and managing converted text data.

[0105] A "generative AI model" is a machine learning model that uses natural language processing technology to generate specific output (in this case, a manuscript for posting on social media) from input data.

[0106] A "prompt" is an instruction given to a generative AI model, a piece of text that acts as a guide to achieving a specific output.

[0107] "Tagging" is the process of assigning relevant keywords and categories to text data.

[0108] "Keyword extraction" is the process of automatically identifying and extracting important words and phrases from text data.

[0109] "SNS Platform" means a website or application that provides social networking services and allows users to post content thereon.

[0110] A "posting schedule" is a plan to automatically post the generated manuscript for posting to social media based on a specific date, time, or event.

[0111] This invention is a system that allows users to talk about everyday events into a device, accumulates the information in a database, analyzes it, and generates drafts for posting to SNS. This system automates the entire process from user voice input to posting to SNS, significantly reducing the burden on users.

[0112] System configuration

[0113] The system of the present invention mainly comprises the following components:

[0114] 1. User's Device

[0115] 2. Server

[0116] 3. Database

[0117] 4. Speech Recognition Engine

[0118] 5. Generative AI Models

[0119] Hardware and software used

[0120] User devices: Electronic devices such as smartphones and tablets are used, allowing users to record daily events by voice.

[0121] Server: A computer system that can be operated in the cloud or on-premise. It processes voice data and manages databases.

[0122] Database: Use a relational database such as MySQL or PostgreSQL to store and manage text data and its metadata.

[0123] Speech recognition engine: Uses speech recognition services such as Google Cloud Speech-to-Text API or AWS Transcribe to convert voice data into text data.

[0124] Generative AI model: Uses generative AI technologies such as OpenAI GPT-3 and Google BERT to generate natural-sounding sentences based on prompts.

[0125] System operation example

[0126] Voice input from the user

[0127] The user speaks about everyday events into the smartphone's microphone, for example, "Today I went to a cafe with a friend and tried a new cake."

[0128] Audio data conversion and display

[0129] The device records this voice data and sends it to the server. The server uses a speech recognition engine to convert this voice data into text data. The converted text is displayed as "Today I went to a cafe with a friend and tried a new cake," and is sent back to the device.

[0130] Data storage and analysis

[0131] The user checks the displayed text and makes corrections as necessary. Once the process is complete, the text data is sent back to the server and saved in a database. The server adds a date and category (e.g., eating out, friends) to the data and stores it in the database. The server also tags the data and extracts keywords.

[0132] Manuscript generation

[0133] At a specified time (e.g., 7 p.m. every day), the server retrieves the saved text data and uses the generative AI model to generate a script for posting on social media. The generative AI model generates a script based on the prompt: "User's voice input: Today I went to a cafe with a friend and tried the new cake. Please generate a script for posting on social media."

[0134] Posts and Notifications

[0135] The generated message is displayed on the device in the form of "Yesterday I went to a cafe with a friend and tried the new cake. It was delicious." Once the user confirms and edits it, the device sends the finalized data to the server. The server posts the message using the API of the SNS platform and notifies the device of the result.

[0136] This system allows users to easily record their daily events using voice and automatically post them to social media, saving users time and effort and enabling consistent social media posting.

[0137] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0138] Step 1:

[0139] The user speaks about everyday events into the device. The user's voice input is recorded on the device through the microphone of the smartphone or tablet. This generates voice data. Input: User's voice / Output: Recorded voice data.

[0140] Step 2:

[0141] The device sends the recorded audio data to the server. The device securely uploads the audio data to the server using the HTTPS protocol. Input: Audio data / Output: Audio data transferred to the server.

[0142] Step 3:

[0143] The server uses a speech recognition engine to convert the voice data into text data. Here, we use the Google Cloud Speech-to-Text API or AWS Transcribe. The server sends the voice data to the cloud and receives text data as a response. Input: Voice data / Output: Text data.

[0144] Step 4:

[0145] The server sends the converted text data back to the terminal, which displays it on the screen. The user can check the content and correct it if necessary. Input: Text data / Output: Text data displayed on the terminal.

[0146] Step 5:

[0147] The user checks and edits the displayed text data, and confirms the text data once the corrections are complete. Input: Displayed text data / Output: Corrected and confirmed text data.

[0148] Step 6:

[0149] The device sends the confirmed text data to the server. The server receives this data, adds dates and categories (e.g., eating out, friends), and stores it in a database. The server also performs tagging and keyword extraction. Input: Confirmed text data / Output: Data stored in the database.

[0150] Step 7:

[0151] At the specified timing, the server retrieves the text data from the database and generates a prompt sentence for the generative AI model. A generative AI model (e.g., OpenAI GPT-3) is used to create natural-sounding sentences based on the input data. Input: Text data / Output: Generated manuscript for posting on social media.

[0152] Step 8:

[0153] The server sends the generated manuscript for posting to the SNS to the terminal. The terminal displays the manuscript to the user and prompts the user to confirm and edit it. Input: Generated manuscript / Output: Manuscript displayed on the terminal.

[0154] Step 9:

[0155] The user reviews and edits the generated manuscript and decides on the final submission. Input: Displayed manuscript / Output: Reviewed and edited final manuscript.

[0156] Step 10:

[0157] The device sends the final manuscript to the server. The server posts the manuscript using the SNS platform's API. Once posting is complete, the result is notified to the device. Input: Final manuscript / Output: Content posted to SNS and notified result.

[0158] This system allows users to easily record everyday events and automatically post them to social media through a series of steps.

[0159] (Application example 1)

[0160] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0161] In the past, users had to write down their daily events and post them manually on social media or blogs, which was time-consuming and difficult to automate. Furthermore, there was a lack of a way to store records over the long term and organize them systematically, making it difficult to effectively manage content. This made it difficult to use, especially for busy users, and prevented them from efficiently recording and sharing their daily events.

[0162] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0163] In this invention, the server includes means for inputting daily events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for saving the text data in a database, means for analyzing the text data in the database and generating a blog post manuscript, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, and means for posting the confirmed and edited manuscript to a designated blog platform. This enables users to easily record daily events, save and manage them as text data, and automatically post them as blog posts.

[0164] A "device" is an electronic device that a user uses to communicate everyday events.

[0165] "Voice data" refers to data that records what the user has said as voice.

[0166] "Text data" is character information obtained by analyzing voice data.

[0167] A "database" is a digital recording system for storing text data for later analysis and retrieval.

[0168] "Analysis" is the process of processing the text data in the database and extracting specific information or patterns.

[0169] A "blog article manuscript" is a document for blog posting that is generated from analyzed text data and has not yet been reviewed or edited by a user.

[0170] "Checking and editing" refers to the process in which the user looks at the generated manuscript, checks the content, and makes corrections as necessary.

[0171] A "blog platform" is an online service that allows users to post and publish blog articles.

[0172] "Tagging" is a method of assigning specific keywords or categories to text data to make it easier to search and organize.

[0173] "Keyword extraction" is a method of extracting important words from text data and using them for data analysis and search.

[0174] "Set posting schedule" is a function that allows users to set the timing for automatically posting blog articles at a specified date and time.

[0175] This invention provides a system that can effectively record daily events that users talk about to a terminal and automatically generate and post blog articles. Specifically, it is configured as follows.

[0176] Users speak into a device such as a smartphone to input voice data about everyday events. The device records the user's voice and sends the voice data to a server. The server uses a voice recognition engine to convert the voice data into text data, which is then sent back to the device and displayed on the screen. The user can check the text data and make corrections as necessary.

[0177] Once the text data has been corrected, it is sent back to the server. The server then tags the text data, extracts keywords, and stores them in a database. The text data stored in the database is then analyzed using a generative AI model to generate a draft blog post. This draft is then sent back to the device in a format that is easy for the user to understand.

[0178] The device displays the generated blog post draft to the user and provides an interface that prompts the user to review and edit it. Once the user has reviewed the draft and finished editing, they confirm the posting. The server receives the confirmed draft and posts it to the blog platform at the specified time. The success or failure of the posting is notified to the device and conveyed to the user.

[0179] In this system, the following hardware and software are used:

[0180] Hardware: Smartphone (device).

[0181] software:

[0182] For voice recognition, the speech_recognition library is used, utilizing Google's voice recognition service.

[0183] For text generation, we use the transformers library and the GPT-3 model.

[0184] Data can be stored using a local file system or a cloud database (e.g., Amazon S3).

[0185] For social media and blog posts, the Python requests library is used to call designated platform APIs.

[0186] As a concrete example, consider the case where a user says, "I went to a new cafe today. I had some really good coffee." This speech is converted into text and stored on a server. It is then analyzed using a generative AI model, and a blog post like the one below is generated.

[0187] Example prompt sentence:

[0188] "Generate a blog post based on the following text: I went to a new cafe today and had some really good coffee."

[0189] The generated blog post is displayed on the device, and once the user has confirmed and edited it, it is automatically posted to the blog platform. By using this system, users can easily and hassle-freely share their everyday events as blog posts.

[0190] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0191] Processing steps of the system that realizes the application example

[0192] Step 1:

[0193] The user inputs voice data

[0194] A user speaks into a device such as a smartphone. This voice data is input as an audio file of what the user said. For example, "I went to a new cafe today. I had some really good coffee." This voice data is recorded on the device.

[0195] Step 2:

[0196] Converting audio data into text data

[0197] The device sends the recorded voice data to the server. The server uses a voice recognition engine (for example, Google's voice recognition service) to convert the voice data into text data. In this process, the voice data (input) is converted into text information (output). The converted text data, "I went to a new cafe today. I had some very delicious coffee," is sent back and displayed on the device's screen.

[0198] Step 3:

[0199] Check and correct text data

[0200] The user checks the text data displayed on the terminal and makes corrections as necessary. For example, "I drank a very delicious coffee" is changed to "I drank a very delicious espresso." In this procedure, the user looks at the text data (input), makes the necessary corrections (data processing), and obtains the corrected text data (output).

[0201] Step 4:

[0202] Storing text data in a database

[0203] The text data confirmed and corrected by the user is sent back to the server from the device. The server tags the text data and extracts keywords (for example, adding tags such as "cafe" or "coffee"), and stores it in a database. In this process, the text data (input) is organized and stored as tagged text data (output).

[0204] Step 5:

[0205] Generate a blog post draft

[0206] At a specified time (for example, a time set by the user each night), the server uses a generative AI model (for example, GPT-3) to generate a draft blog post based on the tagged text data in the database. In this process, the stored text data (input) is passed through a natural language generation algorithm to generate a draft blog post (output).

[0207] Step 6:

[0208] View, check and edit the generated manuscript

[0209] The generated draft of the blog post is sent to the user's device and displayed on the screen. The user can review this draft and make any necessary corrections. For example, the draft of a blog post might read, "I went to a new cafe today. I enjoyed a delicious espresso." In this step, the generated draft (input) is reviewed and edited by the user (data processing) before becoming the final draft (output).

[0210] Step 7:

[0211] Post your blog post to your preferred blogging platform

[0212] After the user has reviewed and edited the blog post, it is sent from the device to the server, which then calls the API of the designated blog platform to post it. For example, an article such as "I went to a new cafe today. I enjoyed a delicious espresso" is posted to the blog. The success or failure of the post is again reported to the device and communicated to the user. Here, the final draft (input) is posted to the blog platform (data calculation and manipulation), and the posting result (output) is returned to the user.

[0213] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0214] This system allows users to talk about everyday events, stores the content in a database, analyzes it, and generates a draft for posting to an SNS. It also recognizes the user's emotions and reflects them in the generated text. This system automates the entire process from user input to posting to an SNS, not only reducing the burden on the user but also enabling posts that take emotions into consideration.

[0215] System configuration and operation

[0216] 1. User voice input

[0217] The user talks to the device about everyday events. For example, "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner."

[0218] 2. Audio data conversion

[0219] The device records the user's voice and sends the voice data to the server. The server uses a speech recognition engine to convert the voice data into text data. The converted text data is sent back to the device and displayed on the screen. The user can check the text and make corrections if necessary.

[0220] 3. Emotional Recognition

[0221] The server sends the voice data to the emotion engine, which then recognizes the emotion from the user's voice. The emotion engine then analyzes the user's emotional state based on the tone of voice and vocabulary selection, and adds the results to the text data.

[0222] 4. Saving to the database

[0223] The device sends the text data it has confirmed to a server. The server then adds a date and category (e.g., work, eating out) to the text data and stores it in a database. The stored data is then organized through natural language processing, including tagging and keyword extraction. Emotional information added through emotion recognition is also stored.

[0224] 5. Creating a manuscript for posting on social media

[0225] When the posting timing specified by the user (e.g., daily, when a specific event occurs) arrives, the server uses generative AI to generate a script for posting to social media based on the text data in the database. The generated script takes emotional information into account. For example, if positive emotions are detected, positive expressions will be emphasized.

[0226] 6. User confirmation and editing

[0227] The device displays the generated script for posting to social media to the user and provides an interface for reviewing and editing as necessary. For example, it could be displayed in the form of "Yesterday, my big presentation was a success! Afterwards, I enjoyed a delicious Italian meal with friends. It was a fulfilling day."

[0228] 7. Posting to social media

[0229] Once the user has finished checking and editing, they confirm the post. The device then receives this and sends the confirmed manuscript to the server. The server then calls the API of each SNS platform and executes the post. The success or failure of the post is again notified to the device and conveyed to the user.

[0230] Specific examples

[0231] For example, consider the case where a user says, "Today I went on a picnic with my family and it was so much fun. Then I saw a new movie."

[0232] 1. The user speaks.

[0233] 2. The device records the audio and sends it to the server.

[0234] 3. The server uses a speech recognition engine to convert this into text data: "Today I went on a picnic with my family and had a great time. Then I saw a new movie."

[0235] 4. The server uses an emotion engine to recognize positive emotions from the voice and tag the text data as "joy."

[0236] 5. Display the text data on the terminal.

[0237] 6. The user checks, corrects, and approves.

[0238] 7. The device sends the text data to the server, where it is tagged, keywords are extracted, and the data is stored in a database.

[0239] 8. Based on the posting timing, the server generates a script for posting to social media: "Yesterday I went on a picnic with my family. It was so much fun. Afterwards, I also saw a new movie. It was a great day."

[0240] 9. The terminal displays the generated manuscript, which the user can check and edit.

[0241] 10. The user submits the finalized manuscript.

[0242] 11. The server posts to the SNS platform and notifies the device of the results.

[0243] This system allows users to easily record and post everyday events, and also generates engaging content that reflects their emotions.

[0244] The processing flow will be explained below.

[0245] Step 1:

[0246] The user talks to the device about everyday events. For example, "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner."

[0247] Step 2:

[0248] The device will record the user's voice and the recorded voice data will be temporarily stored on the device.

[0249] Step 3:

[0250] The device sends the recorded audio data to the server, where it is converted into an appropriate format and sent.

[0251] Step 4:

[0252] The server receives the voice data, which is then input into a voice recognition engine and converted into text data.

[0253] Step 5:

[0254] The server inputs the converted text data into the emotion engine, which analyzes the user's emotions based on the tone and content of the voice and adds emotional information.

[0255] Step 6:

[0256] The server sends text data and emotional information back to the device. The text data, such as "I had a big presentation at work today, and it went well. For dinner, I went out to eat Italian food with a friend," is tagged with emotional information.

[0257] Step 7:

[0258] The device displays the text data and emotional information to the user, who can then review and correct it if necessary.

[0259] Step 8:

[0260] The user checks and corrects the text data and presses the "Confirm" or "Send" button.

[0261] Step 9:

[0262] The device sends the confirmed and corrected text data to the server, which adds date and category information to the text data and stores it in a database.

[0263] Step 10:

[0264] The server then performs tagging and keyword extraction on the data stored in the database, including emotional information.

[0265] Step 11:

[0266] When the designated posting time (e.g., every day or when a specific event occurs) arrives, the server generates a script for posting to social media based on the data in the database. Emotional information is also taken into account, with positive emotions emphasized.

[0267] Step 12:

[0268] The server sends the generated script to the terminal, generating a script in the format "Yesterday's presentation was a success! Afterwards, I enjoyed some delicious Italian food with my friends. It was a fulfilling day."

[0269] Step 13:

[0270] The terminal displays the generated manuscript to the user and provides an interface that prompts them to check and edit it.

[0271] Step 14:

[0272] The user checks the manuscript, edits it if necessary, and presses the "Submit" button.

[0273] Step 15:

[0274] The device sends the finalized manuscript to the server, which then calls the API of each SNS platform and posts the generated manuscript.

[0275] Step 16:

[0276] The server checks whether the posting was successful or not and returns the result to the device.

[0277] Step 17:

[0278] The device notifies the user of the posting result (e.g., "Posting successful!").

[0279] This series of processes allows users to easily record everyday events and post engaging content that reflects their emotions on social media.

[0280] Example 2

[0281] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0282] Conventional SNS posting systems have the problem that users must manually input daily events and create posts that reflect their emotions, which requires a great deal of time and effort. Furthermore, the technical means for reading and appropriately reflecting emotions are insufficient, making it difficult to properly express the user's intended message. Furthermore, the ability to adjust the timing of posts is limited, reducing user convenience. The purpose of this invention is to solve these problems and provide a system that automatically posts a variety of emotionally reflective posts while reducing the burden on users.

[0283] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0284] In this invention, the server includes means for inputting daily events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for saving the text data in a database, means for organizing the saved text data by adding date and category information, means for recognizing emotions from the voice data and adding the results to the text data, means for analyzing the text data in the database and generating a manuscript for posting to an SNS using generation technology, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, means for posting the manuscript confirmed and edited by the user to a specified SNS platform, and means for notifying the user of the success or failure of the posting. This enables users to easily record daily events and automatically generate and post sentences that reflect their emotions.

[0285] "Voice data" refers to data that records what a user says to a terminal as a voice signal.

[0286] "Text data" refers to textual information data obtained by analyzing and converting voice data.

[0287] A "terminal" is a device that a user uses to input voice, and that has the function of recording and displaying voice.

[0288] A "server" is a device or system that processes voice data sent from a terminal, converts it into text data, stores it in a database, analyzes it, and so on.

[0289] A "database" is a data management system for organizing and storing text data and associated information (e.g., dates, categories, and emotion tags).

[0290] "Emotion recognition" is a technology that analyzes and extracts a user's emotional state from voice and text data.

[0291] "Generation technology" refers to technology for generating specific formats and content based on information in a database, and is primarily used to generate manuscripts for posting on social media.

[0292] "SNS Platform" means a system that provides social networking services that enable users to communicate with other users over the Internet.

[0293] "Posting" refers to the act of sending the generated manuscript for posting to an SNS platform on the Internet and making it public.

[0294] "Notification" refers to the act of the server sending specific information (e.g., posting success or failure) to the terminal to inform the user.

[0295] This system allows users to talk about everyday events into a device, accumulates the content in a database, analyzes it, and generates a manuscript for posting to SNS. It also recognizes the user's emotions and reflects them in the generated text. This system automates the entire process from user input to posting to SNS, reducing the burden on the user and enabling posts that take emotions into consideration.

[0296] First, the user talks about everyday events into a dedicated device. The device uses a built-in microphone to record the user's voice and save it as audio data. The recorded audio data is then sent to a server via the Internet.

[0297] The server uses the Google Cloud Speech-to-Text API to convert the received voice data into text data, which is then sent back to the device via the Internet and displayed for the user to review. The user can then review the text data and make corrections as necessary.

[0298] Next, the server sends the voice data to the Microsoft Azure Cognitive Services emotion recognition API for emotion recognition. The emotion recognition API analyzes the voice tone and vocabulary choice to extract the user's emotional state. The analysis results are added to the text data as emotion tags. The text data is then stored in a database. When storing the data, the server uses a natural language processing library (e.g., NLTK) to add date and category information to the text data and organize it.

[0299] When the user's designated posting time arrives, the server retrieves the target text data from the database and generates a script for posting to social media using a generative AI model such as OpenAI's GPT-4. At this time, the server inputs the following prompt to the generative AI model:

[0300] "Generate a social media post based on the following text and sentiment tags. Emphasize positive sentiment and create content that will interest your readers."

[0301] For example, if the server generates a script based on the text data "I had a big presentation at work today, which was a great success. For dinner, I went out to eat Italian food with friends," and the emotion tag "joy," the resulting script would read, "Yesterday's big presentation was a success! Afterwards, I enjoyed some delicious Italian food with friends. It was a fulfilling day."

[0302] The generated manuscript is sent to the terminal and displayed to the user. The user checks the manuscript and makes edits as necessary. When the user has finally finished checking and editing, the terminal sends the manuscript to the server.

[0303] The server calls the API of various social media platforms (e.g., Twitter API, Facebook Graph API) and posts the finalized manuscript on the Internet. The success or failure of this posting is again notified to the device and conveyed to the user.

[0304] The above process allows users to easily record their daily events and effectively post them on social media, reflecting their emotions. This system significantly reduces the burden on users and is an effective way to provide engaging content that takes emotions into consideration.

[0305] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0306] Step 1:

[0307] Users talk to the device about everyday events.

[0308] Specifically, the device's built-in microphone records the user's voice.

[0309] Input: User's voice

[0310] Output: Audio data (recording file)

[0311] Step 2:

[0312] The device sends the recorded audio data to the server.

[0313] Specifically, the audio file is uploaded to a server via the Internet.

[0314] Input: Audio data

[0315] Output: Audio data file on the server

[0316] Step 3:

[0317] The server uses the Google Cloud Speech-to-Text API to convert the audio data into text data.

[0318] Specifically, audio data is sent to the API and the returned text data is obtained.

[0319] Input: Audio data

[0320] Output: Text data (converted character information)

[0321] Step 4:

[0322] The server retransmits the converted text data to the terminal.

[0323] Specifically, text data is downloaded to the terminal via the Internet.

[0324] Input: Text data

[0325] Output: Text data displayed on the terminal

[0326] Step 5:

[0327] The user checks the text data displayed on the terminal and corrects it if necessary.

[0328] Specifically, use a text editor to correct any errors or unnecessary parts.

[0329] Input: Text data

[0330] Output: Text data that the user has confirmed and corrected

[0331] Step 6:

[0332] The terminal transmits the text data corrected by the user to the server again.

[0333] Specifically, the corrected text data is uploaded to a server via the Internet.

[0334] Input: Corrected text data

[0335] Output: Modified text data on the server

[0336] Step 7:

[0337] The server sends the corrected text data to the Microsoft Azure Cognitive Services emotion recognition API for emotional analysis.

[0338] Specifically, text data is sent to the API and the sentiment analysis results are obtained.

[0339] Input: Text data

[0340] Output: Emotion-tagged text data

[0341] Step 8:

[0342] The server adds date and category information to the text data and stores it in a database.

[0343] Specifically, the text data is inserted into the database as a new record.

[0344] Input: emotion-tagged text data

[0345] Output: Text data stored in the database

[0346] Step 9:

[0347] When the specified posting time arrives, the server uses OpenAI's generative AI model to generate a manuscript for posting on social media.

[0348] Specifically, text data is retrieved from a database and prompt sentences are input into a generative AI model to generate a manuscript.

[0349] Input: Text data in the database, prompt statements

[0350] Output: Generated manuscript for posting to social media

[0351] Step 10:

[0352] The generated manuscript is sent to the terminal and displayed to the user.

[0353] Specifically, the generated manuscript is downloaded to the terminal via the Internet.

[0354] Input: Generated SNS posting manuscript

[0355] Output: Manuscript for posting to social media displayed on the device

[0356] Step 11:

[0357] The user reviews the manuscript and edits it if necessary.

[0358] Specifically, the text editor is used again to make corrections and additions to the manuscript.

[0359] Input: Manuscript to post on social media

[0360] Output: User-confirmed and edited manuscript for posting on social media

[0361] Step 12:

[0362] The terminal transmits the final manuscript to the server.

[0363] Specifically, the final manuscript is uploaded to a server via the Internet.

[0364] Input: Confirmed and edited manuscript for posting to social media

[0365] Output: Final manuscript on the server

[0366] Step 13:

[0367] The server calls the APIs of various social media platforms and posts the finalized manuscript.

[0368] Specifically, you submit your manuscript using the posting API of each SNS and retrieve the results.

[0369] Input: Final manuscript

[0370] Output: Posting results to social media

[0371] Step 14:

[0372] The terminal is notified of the success or failure of the posting.

[0373] Specifically, a notification message is sent to the terminal and displayed to the user.

[0374] Input: Post results

[0375] Output: Notification message displayed on the terminal

[0376] (Application example 2)

[0377] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0378] In conventional customer service, it has been difficult to record and analyze customer interactions in real time and use the results as feedback or reviews. In addition, there has been a lack of a way to efficiently generate reviews and social media posts that reflect customer sentiment, so an effective tool is needed to improve customer satisfaction.

[0379] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting everyday events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for saving the text data in a database, means for analyzing the text data in the database and generating a manuscript for posting on an SNS, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, means for posting the manuscript confirmed and edited by the user to a specified SNS platform, and means for recognizing emotions from the voice data and reflecting the recognized emotional information in the generated text. This makes it possible to record the content of conversations with customers in real time and generate feedback and reviews that reflect the emotional information.

[0380] "Terminal" refers to a device that allows users to input voice data and check and edit text data and emotion recognition results.

[0381] "Voice data" refers to information recorded in audio format about everyday events spoken by a user.

[0382] "Text data" refers to information that has been analyzed and converted into text form from audio data.

[0383] The "database" is a data repository for storing generated text data and emotion recognition results.

[0384] "Analysis" refers to the process of understanding and extracting meaning from the text data in a database.

[0385] A "script for posting on social media" is a piece of text created using a generative AI model based on text data, intended for posting on social media.

[0386] "Emotion recognition" is the process of analyzing a user's emotional state from their voice data and obtaining the results.

[0387] A "generative AI model" is an artificial intelligence technology that generates sentences in a specific format based on input data.

[0388] A "prompt sentence" is the initial input sentence that a generative AI model uses when generating a sentence.

[0389] A "wearable device" is an information terminal that is worn by the user.

[0390] The system is designed to enable users to record customer interactions through wearable devices such as smart glasses and automatically generate transcripts for reviews and social media posts. The system consists of the following main components:

[0391] Hardware and Software Configuration

[0392] 1. Device:

[0393] Wearable devices such as smart glasses.

[0394] It is equipped with a microphone and has the ability to receive voice input.

[0395] Record customer interactions in real time.

[0396] 2. Server:

[0397] It runs a speech recognition engine, a sentiment analysis engine, and a generative AI model.

[0398] Convert the voice data into text data and analyze emotions.

[0399] The text data is stored in a database for later analysis and tagging.

[0400] System processing flow

[0401] 1. Voice input:

[0402] The device records the user's voice in real time and sends it to the server as audio data.

[0403] 2. Audio data conversion:

[0404] The server uses a speech recognition engine to convert the voice data into text data, which is then sent back to the device and displayed to the user.

[0405] 3. Emotion Recognition:

[0406] The server recognizes emotions from the recorded voice data and adds emotional information to it using an emotion analysis engine.

[0407] 4. Save to database:

[0408] The text data and sentiment information are stored in a database, organized by date and category (e.g., product reviews, service ratings).

[0409] 5. Creating a manuscript for posting on social media:

[0410] At the specified timing, the server extracts text data from the database and uses the generative AI model to generate a script for posting on social media. This script reflects emotional information. For example, if "positive emotions" are detected, positive expressions will be emphasized.

[0411] 6. User review and editing:

[0412] The manuscript is displayed on the user's terminal, and the user can check it and edit it if necessary.

[0413] 7. Posting to social media:

[0414] The terminal sends the manuscript that the user has finalized to the server, and the server posts it via the API of each SNS platform.

[0415] Specific examples

[0416] For example, in a physical store, imagine a scenario where a salesperson asks a customer, "Today, you tried this new perfume. What do you think?", and the customer replies, "It smells great! I'd like to come again." The system records this conversation and processes it as follows:

[0417] Voice input: The store clerk's smart glasses record the conversation.

[0418] Conversion of voice data: The server converts this into text data such as "You tried this new perfume today. What did you think?" and "It smelled great! I'd like to come again."

[0419] Emotion recognition: Positive emotions are detected.

[0420] Using a generative AI model: Generate a review using the prompt "Review when emotions are positive: The scent was amazing! I'd love to come back. How can I recreate such a wonderful day?"

[0421] User review and editing: Store associates can review and edit generated reviews using smart glasses.

[0422] Posting to social media: After the review is verified and edited, it will be posted to each social media platform.

[0423] Based on these example prompts, the system can quickly generate engaging, emotionally sensitive content, contributing to improved customer satisfaction.

[0424] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0425] Step 1:

[0426] The device records the user's voice in real time and sends the voice data to the server. The input is the user's voice, and the output is the voice data sent to the server. Specifically, the voice input is captured using the microphone in the smart glasses, and the data is compressed and sent to the server.

[0427] Step 2:

[0428] The server uses a speech recognition engine to convert the voice data into text data. The input is the voice data sent in step 1, and the output is the converted text data. Specifically, Google's speech recognition API is used to convert the voice data into Japanese text.

[0429] Step 3:

[0430] The terminal receives the converted text data and displays it to the user. The input is the text data sent from the server, and the output is the text data displayed on the terminal's display. Specifically, the text is displayed on the display of the smart glasses and confirmed by the user.

[0431] Step 4:

[0432] The server recognizes emotions from voice data and adds that emotional information to text data. The input is text data and voice data, and the output is text data with emotional information. Specifically, it uses Hugging Face's emotion analysis engine to analyze emotions from voice tone and keywords, and adds that information to the text data.

[0433] Step 5:

[0434] The server stores text data and emotion information in a database. The input is text data with emotion information, and the output is the data stored in the database. Specifically, data organized by category is stored in the database.

[0435] Step 6:

[0436] At the specified time, the server extracts text data from the database and uses the generative AI model to generate a manuscript for posting on social media. The input is the text data in the database, and the output is the generated manuscript for posting on social media. Specifically, the generative AI model is input with a prompt sentence: "Review when emotions are positive: The scent was amazing! I'd like to come again. How can I recreate such a good day again?" and outputs the generated text.

[0437] Step 7:

[0438] The device receives the generated SNS post manuscript and displays it to the user. The input is the SNS post manuscript sent from the server, and the output is the manuscript displayed on the device's display. Specifically, it is displayed on the smart glasses display for the user to check and edit.

[0439] Step 8:

[0440] The device sends the manuscript edited by the user to be posted to the server, and the server executes the posting via the API of each SNS platform. The input is the edited manuscript to be posted to the SNS, and the output is the result of posting to each SNS platform. Specifically, the SNS API is called to post, and the device is notified of the success or failure of the posting.

[0441] Step 9:

[0442] The device notifies the user of the results of the post to the SNS. The input is the post result sent from the server, and the output is the result notification displayed on the device's display. Specifically, a message such as "Posting successful" is displayed on the smart glasses' display.

[0443] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0444] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0445] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0446] [Second embodiment]

[0447] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0448] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0449] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0450] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0451] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0452] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0453] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0454] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0455] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0456] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0457] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0458] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0459] This invention is a system that allows users to talk about everyday events, accumulates them in a database, and analyzes each piece of data to generate drafts for posting to social media. This system automates the entire process from user input to posting to social media, significantly reducing the burden on users.

[0460] System configuration and operation

[0461] 1. User voice input

[0462] The user talks to the device about everyday events, such as, "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner."

[0463] 2. Audio data conversion

[0464] The device records the user's voice and sends the voice data to the server. The server uses a speech recognition engine to convert the voice data into text data. The converted text data is sent back to the device and displayed on the screen. The user can check the text and make corrections if necessary.

[0465] 3. Saving to the database

[0466] The device sends the text data to the server, which then adds a date and category (e.g., work, eating out) to the text data and stores it in a database. The stored data is then organized using natural language processing such as tagging and keyword extraction.

[0467] 4. Creating a manuscript for posting on social media

[0468] When the posting timing specified by the user (e.g., every day, or when a specific event occurs) arrives, the server uses generative AI to generate a script for posting to SNS based on the text data in the database. This script is sent to the device in a format that the user can easily understand.

[0469] 5. User confirmation and editing

[0470] The device displays the generated script for posting to social media to the user and provides an interface for reviewing and editing as necessary. For example, it could be displayed in the form of "Yesterday, my big presentation was a success! Afterwards, I enjoyed a delicious Italian meal with friends. It was a fulfilling day."

[0471] 6. Posting to social media

[0472] Once the user has finished checking and editing, they confirm the post. The device then receives this and sends the confirmed manuscript to the server. The server then calls the API of each SNS platform and executes the post. The success or failure of the post is again notified to the device and conveyed to the user.

[0473] Specific examples

[0474] For example, a user might say: "Today I went on a picnic with my family and had a great time. Then I saw a new movie."

[0475] 1. The user speaks.

[0476] 2. The device records the audio and sends it to the server.

[0477] 3. The server uses a speech recognition engine to convert this into text data: "Today I went on a picnic with my family and had a great time. Then I saw a new movie."

[0478] 4. Display the text data on the terminal.

[0479] 5. The user checks, corrects, and approves.

[0480] 6. The device sends the text data to the server, where it is tagged, keywords are extracted, and the data is stored in a database.

[0481] 7. At the specified time, the server generates a script to post to social media: "Yesterday, I went on a picnic with my family. It was a lot of fun. Afterwards, I also saw a new movie."

[0482] 8. The terminal displays the generated manuscript, which the user can check and edit.

[0483] 9. The user submits the finalized manuscript.

[0484] 10. The server posts to the SNS platform and notifies the device of the results.

[0485] This system allows users to easily and efficiently record everyday events and post them to social media.

[0486] The processing flow will be explained below.

[0487] Step 1:

[0488] The user talks about everyday events into the device, inputting content such as "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner" through the voice input interface.

[0489] Step 2:

[0490] The device records the user's voice and sends the captured voice data to the server via the voice input interface.

[0491] Step 3:

[0492] The server receives the voice data, sends it to the voice recognition engine, and converts it into text data.

[0493] Step 4:

[0494] The server returns the converted text data to the device. The returned text data is displayed as "I had a big presentation at work today, and it was a success. For dinner, I went out to eat Italian food with a friend."

[0495] Step 5:

[0496] The terminal displays the text data to the user and asks for confirmation, after which the user can check the text data and correct it if necessary.

[0497] Step 6:

[0498] The user confirms and modifies the text data by pressing the "Confirm" button.

[0499] Step 7:

[0500] The device sends the text data that has been confirmed to the server, which then adds the date and category (e.g., work, eating out) to the text data and stores it in a database.

[0501] Step 8:

[0502] The server performs natural language processing on the stored data, including tagging and keyword extraction.

[0503] Step 9:

[0504] When the designated posting time arrives (e.g., every day, when a specific event occurs), the server uses a generation AI based on the data in the database to generate a script for posting on social media.

[0505] Step 10:

[0506] The server sends the generated manuscript to the terminal, which displays it to the user and prompts them to check and edit it.

[0507] Step 11:

[0508] The user can check and edit the displayed manuscript. If necessary, press the "Submit" button to finalize the manuscript.

[0509] Step 12:

[0510] The device sends the finalized manuscript to the server, which then calls the API of each SNS platform and posts the manuscript.

[0511] Step 13:

[0512] The server checks whether the posting was successful or not, and sends the result back to the device to notify the user (e.g., "Posting was successful!").

[0513] Example 1

[0514] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0515] In modern society, posting to social networking sites has become a part of users' daily lives, but posting easily can be difficult given their busy schedules. The effort of manually entering text and the burden of thinking up appropriate phrases can be stressful for users. Furthermore, managing the consistency and timing of social networking posts can be difficult, making effective communication difficult. The present invention aims to solve these problems by providing a system that automates social networking posts and significantly reduces the burden on users.

[0516] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0517] In this invention, the server includes means for inputting daily events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for storing the text data in a database, means for analyzing the text data in the database and generating a manuscript for posting to an SNS, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, means for posting the manuscript confirmed and edited by the user to a specified SNS platform, and means for creating prompt sentences using a generative AI model to generate post content in an automated process, thereby enabling users to easily record daily events by voice and automatically post them to an SNS.

[0518] A "user" is someone who uses the system to input everyday events by voice and post them on social media.

[0519] A "terminal" is a device used by a user to perform voice input, such as a smartphone or tablet.

[0520] "Voice data" refers to digital data that is a recording of what a user says to a device.

[0521] "Text data" refers to data of character information converted from voice data using voice recognition technology.

[0522] A "database" is an information system for storing and managing converted text data.

[0523] A "generative AI model" is a machine learning model that uses natural language processing technology to generate specific output (in this case, a manuscript for posting on social media) from input data.

[0524] A "prompt" is an instruction given to a generative AI model, a piece of text that acts as a guide to achieving a specific output.

[0525] "Tagging" is the process of assigning relevant keywords and categories to text data.

[0526] "Keyword extraction" is the process of automatically identifying and extracting important words and phrases from text data.

[0527] "SNS Platform" means a website or application that provides social networking services and allows users to post content thereon.

[0528] A "posting schedule" is a plan to automatically post the generated manuscript for posting to social media based on a specific date, time, or event.

[0529] This invention is a system that allows users to talk about everyday events into a device, accumulates the information in a database, analyzes it, and generates drafts for posting to SNS. This system automates the entire process from user voice input to posting to SNS, significantly reducing the burden on users.

[0530] System configuration

[0531] The system of the present invention mainly comprises the following components:

[0532] 1. User's Device

[0533] 2. Server

[0534] 3. Database

[0535] 4. Speech Recognition Engine

[0536] 5. Generative AI Models

[0537] Hardware and software used

[0538] User devices: Electronic devices such as smartphones and tablets are used, allowing users to record daily events by voice.

[0539] Server: A computer system that can be operated in the cloud or on-premise. It processes voice data and manages databases.

[0540] Database: Use a relational database such as MySQL or PostgreSQL to store and manage text data and its metadata.

[0541] Speech recognition engine: Uses speech recognition services such as Google Cloud Speech-to-Text API or AWS Transcribe to convert voice data into text data.

[0542] Generative AI model: Uses generative AI technologies such as OpenAI GPT-3 and Google BERT to generate natural-sounding sentences based on prompts.

[0543] System operation example

[0544] Voice input from the user

[0545] The user speaks about everyday events into the smartphone's microphone, for example, "Today I went to a cafe with a friend and tried a new cake."

[0546] Audio data conversion and display

[0547] The device records this voice data and sends it to the server. The server uses a speech recognition engine to convert this voice data into text data. The converted text is displayed as "Today I went to a cafe with a friend and tried a new cake," and is sent back to the device.

[0548] Data storage and analysis

[0549] The user checks the displayed text and makes corrections as necessary. Once the process is complete, the text data is sent back to the server and saved in a database. The server adds a date and category (e.g., eating out, friends) to the data and stores it in the database. The server also tags the data and extracts keywords.

[0550] Manuscript generation

[0551] At a specified time (e.g., 7 p.m. every day), the server retrieves the saved text data and uses the generative AI model to generate a script for posting on social media. The generative AI model generates a script based on the prompt: "User's voice input: Today I went to a cafe with a friend and tried the new cake. Please generate a script for posting on social media."

[0552] Posts and Notifications

[0553] The generated message is displayed on the device in the form of "Yesterday I went to a cafe with a friend and tried the new cake. It was delicious." Once the user confirms and edits it, the device sends the finalized data to the server. The server posts the message using the API of the SNS platform and notifies the device of the result.

[0554] This system allows users to easily record their daily events using voice and automatically post them to social media, saving users time and effort and enabling consistent social media posting.

[0555] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0556] Step 1:

[0557] The user speaks about everyday events into the device. The user's voice input is recorded on the device through the microphone of the smartphone or tablet. This generates voice data. Input: User's voice / Output: Recorded voice data.

[0558] Step 2:

[0559] The device sends the recorded audio data to the server. The device securely uploads the audio data to the server using the HTTPS protocol. Input: Audio data / Output: Audio data transferred to the server.

[0560] Step 3:

[0561] The server uses a speech recognition engine to convert the voice data into text data. Here, we use the Google Cloud Speech-to-Text API or AWS Transcribe. The server sends the voice data to the cloud and receives text data as a response. Input: Voice data / Output: Text data.

[0562] Step 4:

[0563] The server sends the converted text data back to the terminal, which displays it on the screen. The user can check the content and correct it if necessary. Input: Text data / Output: Text data displayed on the terminal.

[0564] Step 5:

[0565] The user checks and edits the displayed text data, and confirms the text data once the corrections are complete. Input: Displayed text data / Output: Corrected and confirmed text data.

[0566] Step 6:

[0567] The device sends the confirmed text data to the server. The server receives this data, adds dates and categories (e.g., eating out, friends), and stores it in a database. The server also performs tagging and keyword extraction. Input: Confirmed text data / Output: Data stored in the database.

[0568] Step 7:

[0569] At the specified timing, the server retrieves the text data from the database and generates a prompt sentence for the generative AI model. A generative AI model (e.g., OpenAI GPT-3) is used to create natural-sounding sentences based on the input data. Input: Text data / Output: Generated manuscript for posting on social media.

[0570] Step 8:

[0571] The server sends the generated manuscript for posting to the SNS to the terminal. The terminal displays the manuscript to the user and prompts the user to confirm and edit it. Input: Generated manuscript / Output: Manuscript displayed on the terminal.

[0572] Step 9:

[0573] The user reviews and edits the generated manuscript and decides on the final submission. Input: Displayed manuscript / Output: Reviewed and edited final manuscript.

[0574] Step 10:

[0575] The device sends the final manuscript to the server. The server posts the manuscript using the SNS platform's API. Once posting is complete, the result is notified to the device. Input: Final manuscript / Output: Content posted to SNS and notified result.

[0576] This system allows users to easily record everyday events and automatically post them to social media through a series of steps.

[0577] (Application example 1)

[0578] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0579] In the past, users had to write down their daily events and post them manually on social media or blogs, which was time-consuming and difficult to automate. Furthermore, there was a lack of a way to store records over the long term and organize them systematically, making it difficult to effectively manage content. This made it difficult to use, especially for busy users, and prevented them from efficiently recording and sharing their daily events.

[0580] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0581] In this invention, the server includes means for inputting daily events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for saving the text data in a database, means for analyzing the text data in the database and generating a blog post manuscript, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, and means for posting the confirmed and edited manuscript to a designated blog platform. This enables users to easily record daily events, save and manage them as text data, and automatically post them as blog posts.

[0582] A "device" is an electronic device that a user uses to communicate everyday events.

[0583] "Voice data" refers to data that records what the user has said as voice.

[0584] "Text data" is character information obtained by analyzing voice data.

[0585] A "database" is a digital recording system for storing text data for later analysis and retrieval.

[0586] "Analysis" is the process of processing the text data in the database and extracting specific information or patterns.

[0587] A "blog article manuscript" is a document for blog posting that is generated from analyzed text data and has not yet been reviewed or edited by a user.

[0588] "Checking and editing" refers to the process in which the user looks at the generated manuscript, checks the content, and makes corrections as necessary.

[0589] A "blog platform" is an online service that allows users to post and publish blog articles.

[0590] "Tagging" is a method of assigning specific keywords or categories to text data to make it easier to search and organize.

[0591] "Keyword extraction" is a method of extracting important words from text data and using them for data analysis and search.

[0592] "Set posting schedule" is a function that allows users to set the timing for automatically posting blog articles at a specified date and time.

[0593] This invention provides a system that can effectively record daily events that users talk about to a terminal and automatically generate and post blog articles. Specifically, it is configured as follows.

[0594] Users speak into a device such as a smartphone to input voice data about everyday events. The device records the user's voice and sends the voice data to a server. The server uses a voice recognition engine to convert the voice data into text data, which is then sent back to the device and displayed on the screen. The user can check the text data and make corrections as necessary.

[0595] Once the text data has been corrected, it is sent back to the server. The server then tags the text data, extracts keywords, and stores them in a database. The text data stored in the database is then analyzed using a generative AI model to generate a draft blog post. This draft is then sent back to the device in a format that is easy for the user to understand.

[0596] The device displays the generated blog post draft to the user and provides an interface that prompts the user to review and edit it. Once the user has reviewed the draft and finished editing, they confirm the posting. The server receives the confirmed draft and posts it to the blog platform at the specified time. The success or failure of the posting is notified to the device and conveyed to the user.

[0597] In this system, the following hardware and software are used:

[0598] Hardware: Smartphone (device).

[0599] software:

[0600] For voice recognition, the speech_recognition library is used, utilizing Google's voice recognition service.

[0601] For text generation, we use the transformers library and the GPT-3 model.

[0602] Data can be stored using a local file system or a cloud database (e.g., Amazon S3).

[0603] For social media and blog posts, the Python requests library is used to call designated platform APIs.

[0604] As a concrete example, consider the case where a user says, "I went to a new cafe today. I had some really good coffee." This speech is converted into text and stored on a server. It is then analyzed using a generative AI model, and a blog post like the one below is generated.

[0605] Example prompt sentence:

[0606] "Generate a blog post based on the following text: I went to a new cafe today and had some really good coffee."

[0607] The generated blog post is displayed on the device, and once the user has confirmed and edited it, it is automatically posted to the blog platform. By using this system, users can easily and hassle-freely share their everyday events as blog posts.

[0608] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0609] Processing steps of the system that realizes the application example

[0610] Step 1:

[0611] The user inputs voice data

[0612] A user speaks into a device such as a smartphone. This voice data is input as an audio file of what the user said. For example, "I went to a new cafe today. I had some really good coffee." This voice data is recorded on the device.

[0613] Step 2:

[0614] Converting audio data into text data

[0615] The device sends the recorded voice data to the server. The server uses a voice recognition engine (for example, Google's voice recognition service) to convert the voice data into text data. In this process, the voice data (input) is converted into text information (output). The converted text data, "I went to a new cafe today. I had some very delicious coffee," is sent back and displayed on the device's screen.

[0616] Step 3:

[0617] Check and correct text data

[0618] The user checks the text data displayed on the terminal and makes corrections as necessary. For example, "I drank a very delicious coffee" is changed to "I drank a very delicious espresso." In this procedure, the user looks at the text data (input), makes the necessary corrections (data processing), and obtains the corrected text data (output).

[0619] Step 4:

[0620] Storing text data in a database

[0621] The text data confirmed and corrected by the user is sent back to the server from the device. The server tags the text data and extracts keywords (for example, adding tags such as "cafe" or "coffee"), and stores it in a database. In this process, the text data (input) is organized and stored as tagged text data (output).

[0622] Step 5:

[0623] Generate a blog post draft

[0624] At a specified time (for example, a time set by the user each night), the server uses a generative AI model (for example, GPT-3) to generate a draft blog post based on the tagged text data in the database. In this process, the stored text data (input) is passed through a natural language generation algorithm to generate a draft blog post (output).

[0625] Step 6:

[0626] View, check and edit the generated manuscript

[0627] The generated draft of the blog post is sent to the user's device and displayed on the screen. The user can review this draft and make any necessary corrections. For example, the draft of a blog post might read, "I went to a new cafe today. I enjoyed a delicious espresso." In this step, the generated draft (input) is reviewed and edited by the user (data processing) before becoming the final draft (output).

[0628] Step 7:

[0629] Post your blog post to your preferred blogging platform

[0630] After the user has reviewed and edited the blog post, it is sent from the device to the server, which then calls the API of the designated blog platform to post it. For example, an article such as "I went to a new cafe today. I enjoyed a delicious espresso" is posted to the blog. The success or failure of the post is again reported to the device and communicated to the user. Here, the final draft (input) is posted to the blog platform (data calculation and manipulation), and the posting result (output) is returned to the user.

[0631] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0632] This system allows users to talk about everyday events, stores the content in a database, analyzes it, and generates a draft for posting to an SNS. It also recognizes the user's emotions and reflects them in the generated text. This system automates the entire process from user input to posting to an SNS, not only reducing the burden on the user but also enabling posts that take emotions into consideration.

[0633] System configuration and operation

[0634] 1. User voice input

[0635] The user talks to the device about everyday events. For example, "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner."

[0636] 2. Audio data conversion

[0637] The device records the user's voice and sends the voice data to the server. The server uses a speech recognition engine to convert the voice data into text data. The converted text data is sent back to the device and displayed on the screen. The user can check the text and make corrections if necessary.

[0638] 3. Emotional Recognition

[0639] The server sends the voice data to the emotion engine, which then recognizes the emotion from the user's voice. The emotion engine then analyzes the user's emotional state based on the tone of voice and vocabulary selection, and adds the results to the text data.

[0640] 4. Saving to the database

[0641] The device sends the text data it has confirmed to a server. The server then adds a date and category (e.g., work, eating out) to the text data and stores it in a database. The stored data is then organized through natural language processing, including tagging and keyword extraction. Emotional information added through emotion recognition is also stored.

[0642] 5. Creating a manuscript for posting on social media

[0643] When the posting timing specified by the user (e.g., daily, when a specific event occurs) arrives, the server uses generative AI to generate a script for posting to social media based on the text data in the database. The generated script takes emotional information into account. For example, if positive emotions are detected, positive expressions will be emphasized.

[0644] 6. User confirmation and editing

[0645] The device displays the generated script for posting to social media to the user and provides an interface for reviewing and editing as necessary. For example, it could be displayed in the form of "Yesterday, my big presentation was a success! Afterwards, I enjoyed a delicious Italian meal with friends. It was a fulfilling day."

[0646] 7. Posting to social media

[0647] Once the user has finished checking and editing, they confirm the post. The device then receives this and sends the confirmed manuscript to the server. The server then calls the API of each SNS platform and executes the post. The success or failure of the post is again notified to the device and conveyed to the user.

[0648] Specific examples

[0649] For example, consider the case where a user says, "Today I went on a picnic with my family and it was so much fun. Then I saw a new movie."

[0650] 1. The user speaks.

[0651] 2. The device records the audio and sends it to the server.

[0652] 3. The server uses a speech recognition engine to convert this into text data: "Today I went on a picnic with my family and had a great time. Then I saw a new movie."

[0653] 4. The server uses an emotion engine to recognize positive emotions from the voice and tag the text data as "joy."

[0654] 5. Display the text data on the terminal.

[0655] 6. The user checks, corrects, and approves.

[0656] 7. The device sends the text data to the server, where it is tagged, keywords are extracted, and the data is stored in a database.

[0657] 8. Based on the posting timing, the server generates a script for posting to social media: "Yesterday I went on a picnic with my family. It was so much fun. Afterwards, I also saw a new movie. It was a great day."

[0658] 9. The terminal displays the generated manuscript, which the user can check and edit.

[0659] 10. The user submits the finalized manuscript.

[0660] 11. The server posts to the SNS platform and notifies the device of the results.

[0661] This system allows users to easily record and post everyday events, and also generates engaging content that reflects their emotions.

[0662] The processing flow will be explained below.

[0663] Step 1:

[0664] The user talks to the device about everyday events. For example, "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner."

[0665] Step 2:

[0666] The device will record the user's voice and the recorded voice data will be temporarily stored on the device.

[0667] Step 3:

[0668] The device sends the recorded audio data to the server, where it is converted into an appropriate format and sent.

[0669] Step 4:

[0670] The server receives the voice data, which is then input into a voice recognition engine and converted into text data.

[0671] Step 5:

[0672] The server inputs the converted text data into the emotion engine, which analyzes the user's emotions based on the tone and content of the voice and adds emotional information.

[0673] Step 6:

[0674] The server sends text data and emotional information back to the device. The text data, such as "I had a big presentation at work today, and it went well. For dinner, I went out to eat Italian food with a friend," is tagged with emotional information.

[0675] Step 7:

[0676] The device displays the text data and emotional information to the user, who can then review and correct it if necessary.

[0677] Step 8:

[0678] The user checks and corrects the text data and presses the "Confirm" or "Send" button.

[0679] Step 9:

[0680] The device sends the confirmed and corrected text data to the server, which adds date and category information to the text data and stores it in a database.

[0681] Step 10:

[0682] The server then performs tagging and keyword extraction on the data stored in the database, including emotional information.

[0683] Step 11:

[0684] When the designated posting time (e.g., every day or when a specific event occurs) arrives, the server generates a script for posting to social media based on the data in the database. Emotional information is also taken into account, with positive emotions emphasized.

[0685] Step 12:

[0686] The server sends the generated script to the terminal, generating a script in the format "Yesterday's presentation was a success! Afterwards, I enjoyed some delicious Italian food with my friends. It was a fulfilling day."

[0687] Step 13:

[0688] The terminal displays the generated manuscript to the user and provides an interface that prompts them to check and edit it.

[0689] Step 14:

[0690] The user checks the manuscript, edits it if necessary, and presses the "Submit" button.

[0691] Step 15:

[0692] The device sends the finalized manuscript to the server, which then calls the API of each SNS platform and posts the generated manuscript.

[0693] Step 16:

[0694] The server checks whether the posting was successful or not and returns the result to the device.

[0695] Step 17:

[0696] The device notifies the user of the posting result (e.g., "Posting successful!").

[0697] This series of processes allows users to easily record everyday events and post engaging content that reflects their emotions on social media.

[0698] Example 2

[0699] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0700] Conventional SNS posting systems have the problem that users must manually input daily events and create posts that reflect their emotions, which requires a great deal of time and effort. Furthermore, the technical means for reading and appropriately reflecting emotions are insufficient, making it difficult to properly express the user's intended message. Furthermore, the ability to adjust the timing of posts is limited, reducing user convenience. The purpose of this invention is to solve these problems and provide a system that automatically posts a variety of emotionally reflective posts while reducing the burden on users.

[0701] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0702] In this invention, the server includes means for inputting daily events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for saving the text data in a database, means for organizing the saved text data by adding date and category information, means for recognizing emotions from the voice data and adding the results to the text data, means for analyzing the text data in the database and generating a manuscript for posting to an SNS using generation technology, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, means for posting the manuscript confirmed and edited by the user to a specified SNS platform, and means for notifying the user of the success or failure of the posting. This enables users to easily record daily events and automatically generate and post sentences that reflect their emotions.

[0703] "Voice data" refers to data that records what a user says to a terminal as a voice signal.

[0704] "Text data" refers to textual information data obtained by analyzing and converting voice data.

[0705] A "terminal" is a device that a user uses to input voice, and that has the function of recording and displaying voice.

[0706] A "server" is a device or system that processes voice data sent from a terminal, converts it into text data, stores it in a database, analyzes it, and so on.

[0707] A "database" is a data management system for organizing and storing text data and associated information (e.g., dates, categories, and emotion tags).

[0708] "Emotion recognition" is a technology that analyzes and extracts a user's emotional state from voice and text data.

[0709] "Generation technology" refers to technology for generating specific formats and content based on information in a database, and is primarily used to generate manuscripts for posting on social media.

[0710] "SNS Platform" means a system that provides social networking services that enable users to communicate with other users over the Internet.

[0711] "Posting" refers to the act of sending the generated manuscript for posting to an SNS platform on the Internet and making it public.

[0712] "Notification" refers to the act of the server sending specific information (e.g., posting success or failure) to the terminal to inform the user.

[0713] This system allows users to talk about everyday events into a device, accumulates the content in a database, analyzes it, and generates a manuscript for posting to SNS. It also recognizes the user's emotions and reflects them in the generated text. This system automates the entire process from user input to posting to SNS, reducing the burden on the user and enabling posts that take emotions into consideration.

[0714] First, the user talks about everyday events into a dedicated device. The device uses a built-in microphone to record the user's voice and save it as audio data. The recorded audio data is then sent to a server via the Internet.

[0715] The server uses the Google Cloud Speech-to-Text API to convert the received voice data into text data, which is then sent back to the device via the Internet and displayed for the user to review. The user can then review the text data and make corrections as necessary.

[0716] Next, the server sends the voice data to the Microsoft Azure Cognitive Services emotion recognition API for emotion recognition. The emotion recognition API analyzes the voice tone and vocabulary choice to extract the user's emotional state. The analysis results are added to the text data as emotion tags. The text data is then stored in a database. When storing the data, the server uses a natural language processing library (e.g., NLTK) to add date and category information to the text data and organize it.

[0717] When the user's designated posting time arrives, the server retrieves the target text data from the database and generates a script for posting to social media using a generative AI model such as OpenAI's GPT-4. At this time, the server inputs the following prompt to the generative AI model:

[0718] "Generate a social media post based on the following text and sentiment tags. Emphasize positive sentiment and create content that will interest your readers."

[0719] For example, if the server generates a script based on the text data "I had a big presentation at work today, which was a great success. For dinner, I went out to eat Italian food with friends," and the emotion tag "joy," the resulting script would read, "Yesterday's big presentation was a success! Afterwards, I enjoyed some delicious Italian food with friends. It was a fulfilling day."

[0720] The generated manuscript is sent to the terminal and displayed to the user. The user checks the manuscript and makes edits as necessary. When the user has finally finished checking and editing, the terminal sends the manuscript to the server.

[0721] The server calls the API of various social media platforms (e.g., Twitter API, Facebook Graph API) and posts the finalized manuscript on the Internet. The success or failure of this posting is again notified to the device and conveyed to the user.

[0722] The above process allows users to easily record their daily events and effectively post them on social media, reflecting their emotions. This system significantly reduces the burden on users and is an effective way to provide engaging content that takes emotions into consideration.

[0723] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0724] Step 1:

[0725] Users talk to the device about everyday events.

[0726] Specifically, the device's built-in microphone records the user's voice.

[0727] Input: User's voice

[0728] Output: Audio data (recording file)

[0729] Step 2:

[0730] The device sends the recorded audio data to the server.

[0731] Specifically, the audio file is uploaded to a server via the Internet.

[0732] Input: Audio data

[0733] Output: Audio data file on the server

[0734] Step 3:

[0735] The server uses the Google Cloud Speech-to-Text API to convert the audio data into text data.

[0736] Specifically, audio data is sent to the API and the returned text data is obtained.

[0737] Input: Audio data

[0738] Output: Text data (converted character information)

[0739] Step 4:

[0740] The server retransmits the converted text data to the terminal.

[0741] Specifically, text data is downloaded to the terminal via the Internet.

[0742] Input: Text data

[0743] Output: Text data displayed on the terminal

[0744] Step 5:

[0745] The user checks the text data displayed on the terminal and corrects it if necessary.

[0746] Specifically, use a text editor to correct any errors or unnecessary parts.

[0747] Input: Text data

[0748] Output: Text data that the user has confirmed and corrected

[0749] Step 6:

[0750] The terminal transmits the text data corrected by the user to the server again.

[0751] Specifically, the corrected text data is uploaded to a server via the Internet.

[0752] Input: Corrected text data

[0753] Output: Modified text data on the server

[0754] Step 7:

[0755] The server sends the corrected text data to the Microsoft Azure Cognitive Services emotion recognition API for emotional analysis.

[0756] Specifically, text data is sent to the API and the sentiment analysis results are obtained.

[0757] Input: Text data

[0758] Output: Emotion-tagged text data

[0759] Step 8:

[0760] The server adds date and category information to the text data and stores it in a database.

[0761] Specifically, the text data is inserted into the database as a new record.

[0762] Input: emotion-tagged text data

[0763] Output: Text data stored in the database

[0764] Step 9:

[0765] When the specified posting time arrives, the server uses OpenAI's generative AI model to generate a manuscript for posting on social media.

[0766] Specifically, text data is retrieved from a database and prompt sentences are input into a generative AI model to generate a manuscript.

[0767] Input: Text data in the database, prompt statements

[0768] Output: Generated manuscript for posting to social media

[0769] Step 10:

[0770] The generated manuscript is sent to the terminal and displayed to the user.

[0771] Specifically, the generated manuscript is downloaded to the terminal via the Internet.

[0772] Input: Generated SNS posting manuscript

[0773] Output: Manuscript for posting to social media displayed on the device

[0774] Step 11:

[0775] The user reviews the manuscript and edits it if necessary.

[0776] Specifically, the text editor is used again to make corrections and additions to the manuscript.

[0777] Input: Manuscript to post on social media

[0778] Output: User-confirmed and edited manuscript for posting on social media

[0779] Step 12:

[0780] The terminal transmits the final manuscript to the server.

[0781] Specifically, the final manuscript is uploaded to a server via the Internet.

[0782] Input: Confirmed and edited manuscript for posting to social media

[0783] Output: Final manuscript on the server

[0784] Step 13:

[0785] The server calls the APIs of various social media platforms and posts the finalized manuscript.

[0786] Specifically, you submit your manuscript using the posting API of each SNS and retrieve the results.

[0787] Input: Final manuscript

[0788] Output: Posting results to social media

[0789] Step 14:

[0790] The terminal is notified of the success or failure of the posting.

[0791] Specifically, a notification message is sent to the terminal and displayed to the user.

[0792] Input: Post results

[0793] Output: Notification message displayed on the terminal

[0794] (Application example 2)

[0795] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0796] In conventional customer service, it has been difficult to record and analyze customer interactions in real time and use the results as feedback or reviews. In addition, there has been a lack of a way to efficiently generate reviews and social media posts that reflect customer sentiment, so an effective tool is needed to improve customer satisfaction.

[0797] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting everyday events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for saving the text data in a database, means for analyzing the text data in the database and generating a manuscript for posting on an SNS, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, means for posting the manuscript confirmed and edited by the user to a specified SNS platform, and means for recognizing emotions from the voice data and reflecting the recognized emotional information in the generated text. This makes it possible to record the content of conversations with customers in real time and generate feedback and reviews that reflect the emotional information.

[0798] "Terminal" refers to a device that allows users to input voice data and check and edit text data and emotion recognition results.

[0799] "Voice data" refers to information recorded in audio format about everyday events spoken by a user.

[0800] "Text data" refers to information that has been analyzed and converted into text form from audio data.

[0801] The "database" is a data repository for storing generated text data and emotion recognition results.

[0802] "Analysis" refers to the process of understanding and extracting meaning from the text data in a database.

[0803] A "script for posting on social media" is a piece of text created using a generative AI model based on text data, intended for posting on social media.

[0804] "Emotion recognition" is the process of analyzing a user's emotional state from their voice data and obtaining the results.

[0805] A "generative AI model" is an artificial intelligence technology that generates sentences in a specific format based on input data.

[0806] A "prompt sentence" is the initial input sentence that a generative AI model uses when generating a sentence.

[0807] A "wearable device" is an information terminal that is worn by the user.

[0808] The system is designed to enable users to record customer interactions through wearable devices such as smart glasses and automatically generate transcripts for reviews and social media posts. The system consists of the following main components:

[0809] Hardware and Software Configuration

[0810] 1. Device:

[0811] Wearable devices such as smart glasses.

[0812] It is equipped with a microphone and has the ability to receive voice input.

[0813] Record customer interactions in real time.

[0814] 2. Server:

[0815] It runs a speech recognition engine, a sentiment analysis engine, and a generative AI model.

[0816] Convert the voice data into text data and analyze emotions.

[0817] The text data is stored in a database for later analysis and tagging.

[0818] System processing flow

[0819] 1. Voice input:

[0820] The device records the user's voice in real time and sends it to the server as audio data.

[0821] 2. Audio data conversion:

[0822] The server uses a speech recognition engine to convert the voice data into text data, which is then sent back to the device and displayed to the user.

[0823] 3. Emotion Recognition:

[0824] The server recognizes emotions from the recorded voice data and adds emotional information to it using an emotion analysis engine.

[0825] 4. Save to database:

[0826] The text data and sentiment information are stored in a database, organized by date and category (e.g., product reviews, service ratings).

[0827] 5. Creating a manuscript for posting on social media:

[0828] At the specified timing, the server extracts text data from the database and uses the generative AI model to generate a script for posting on social media. This script reflects emotional information. For example, if "positive emotions" are detected, positive expressions will be emphasized.

[0829] 6. User review and editing:

[0830] The manuscript is displayed on the user's terminal, and the user can check it and edit it if necessary.

[0831] 7. Posting to social media:

[0832] The terminal sends the manuscript that the user has finalized to the server, and the server posts it via the API of each SNS platform.

[0833] Specific examples

[0834] For example, in a physical store, imagine a scenario where a salesperson asks a customer, "Today, you tried this new perfume. What do you think?", and the customer replies, "It smells great! I'd like to come again." The system records this conversation and processes it as follows:

[0835] Voice input: The store clerk's smart glasses record the conversation.

[0836] Conversion of voice data: The server converts this into text data such as "You tried this new perfume today. What did you think?" and "It smelled great! I'd like to come again."

[0837] Emotion recognition: Positive emotions are detected.

[0838] Using a generative AI model: Generate a review using the prompt "Review when emotions are positive: The scent was amazing! I'd love to come back. How can I recreate such a wonderful day?"

[0839] User review and editing: Store associates can review and edit generated reviews using smart glasses.

[0840] Posting to social media: After the review is verified and edited, it will be posted to each social media platform.

[0841] Based on these example prompts, the system can quickly generate engaging, emotionally sensitive content, contributing to improved customer satisfaction.

[0842] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0843] Step 1:

[0844] The device records the user's voice in real time and sends the voice data to the server. The input is the user's voice, and the output is the voice data sent to the server. Specifically, the voice input is captured using the microphone in the smart glasses, and the data is compressed and sent to the server.

[0845] Step 2:

[0846] The server uses a speech recognition engine to convert the voice data into text data. The input is the voice data sent in step 1, and the output is the converted text data. Specifically, Google's speech recognition API is used to convert the voice data into Japanese text.

[0847] Step 3:

[0848] The terminal receives the converted text data and displays it to the user. The input is the text data sent from the server, and the output is the text data displayed on the terminal's display. Specifically, the text is displayed on the display of the smart glasses and confirmed by the user.

[0849] Step 4:

[0850] The server recognizes emotions from voice data and adds that emotional information to text data. The input is text data and voice data, and the output is text data with emotional information. Specifically, it uses Hugging Face's emotion analysis engine to analyze emotions from voice tone and keywords, and adds that information to the text data.

[0851] Step 5:

[0852] The server stores text data and emotion information in a database. The input is text data with emotion information, and the output is the data stored in the database. Specifically, data organized by category is stored in the database.

[0853] Step 6:

[0854] At the specified time, the server extracts text data from the database and uses the generative AI model to generate a manuscript for posting on social media. The input is the text data in the database, and the output is the generated manuscript for posting on social media. Specifically, the generative AI model is input with a prompt sentence: "Review when emotions are positive: The scent was amazing! I'd like to come again. How can I recreate such a good day again?" and outputs the generated text.

[0855] Step 7:

[0856] The device receives the generated SNS post manuscript and displays it to the user. The input is the SNS post manuscript sent from the server, and the output is the manuscript displayed on the device's display. Specifically, it is displayed on the smart glasses display for the user to check and edit.

[0857] Step 8:

[0858] The device sends the manuscript edited by the user to be posted to the server, and the server executes the posting via the API of each SNS platform. The input is the edited manuscript to be posted to the SNS, and the output is the result of posting to each SNS platform. Specifically, the SNS API is called to post, and the device is notified of the success or failure of the posting.

[0859] Step 9:

[0860] The device notifies the user of the results of the post to the SNS. The input is the post result sent from the server, and the output is the result notification displayed on the device's display. Specifically, a message such as "Posting successful" is displayed on the smart glasses' display.

[0861] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0862] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0863] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0864] [Third embodiment]

[0865] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0866] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0867] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0868] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0869] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0870] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0871] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0872] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0873] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0874] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0875] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0876] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0877] This invention is a system that allows users to talk about everyday events, accumulates them in a database, and analyzes each piece of data to generate drafts for posting to social media. This system automates the entire process from user input to posting to social media, significantly reducing the burden on users.

[0878] System configuration and operation

[0879] 1. User voice input

[0880] The user talks to the device about everyday events, such as, "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner."

[0881] 2. Audio data conversion

[0882] The device records the user's voice and sends the voice data to the server. The server uses a speech recognition engine to convert the voice data into text data. The converted text data is sent back to the device and displayed on the screen. The user can check the text and make corrections if necessary.

[0883] 3. Saving to the database

[0884] The device sends the text data to the server, which then adds a date and category (e.g., work, eating out) to the text data and stores it in a database. The stored data is then organized using natural language processing such as tagging and keyword extraction.

[0885] 4. Creating a manuscript for posting on social media

[0886] When the posting timing specified by the user (e.g., every day, or when a specific event occurs) arrives, the server uses generative AI to generate a script for posting to SNS based on the text data in the database. This script is sent to the device in a format that the user can easily understand.

[0887] 5. User confirmation and editing

[0888] The device displays the generated script for posting to social media to the user and provides an interface for reviewing and editing as necessary. For example, it could be displayed in the form of "Yesterday, my big presentation was a success! Afterwards, I enjoyed a delicious Italian meal with friends. It was a fulfilling day."

[0889] 6. Posting to social media

[0890] Once the user has finished checking and editing, they confirm the post. The device then receives this and sends the confirmed manuscript to the server. The server then calls the API of each SNS platform and executes the post. The success or failure of the post is again notified to the device and conveyed to the user.

[0891] Specific examples

[0892] For example, a user might say: "Today I went on a picnic with my family and had a great time. Then I saw a new movie."

[0893] 1. The user speaks.

[0894] 2. The device records the audio and sends it to the server.

[0895] 3. The server uses a speech recognition engine to convert this into text data: "Today I went on a picnic with my family and had a great time. Then I saw a new movie."

[0896] 4. Display the text data on the terminal.

[0897] 5. The user checks, corrects, and approves.

[0898] 6. The device sends the text data to the server, where it is tagged, keywords are extracted, and the data is stored in a database.

[0899] 7. At the specified time, the server generates a script to post to social media: "Yesterday, I went on a picnic with my family. It was a lot of fun. Afterwards, I also saw a new movie."

[0900] 8. The terminal displays the generated manuscript, which the user can check and edit.

[0901] 9. The user submits the finalized manuscript.

[0902] 10. The server posts to the SNS platform and notifies the device of the results.

[0903] This system allows users to easily and efficiently record everyday events and post them to social media.

[0904] The processing flow will be explained below.

[0905] Step 1:

[0906] The user talks about everyday events into the device, inputting content such as "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner" through the voice input interface.

[0907] Step 2:

[0908] The device records the user's voice and sends the captured voice data to the server via the voice input interface.

[0909] Step 3:

[0910] The server receives the voice data, sends it to the voice recognition engine, and converts it into text data.

[0911] Step 4:

[0912] The server returns the converted text data to the device. The returned text data is displayed as "I had a big presentation at work today, and it was a success. For dinner, I went out to eat Italian food with a friend."

[0913] Step 5:

[0914] The terminal displays the text data to the user and asks for confirmation, after which the user can check the text data and correct it if necessary.

[0915] Step 6:

[0916] The user confirms and modifies the text data by pressing the "Confirm" button.

[0917] Step 7:

[0918] The device sends the text data that has been confirmed to the server, which then adds the date and category (e.g., work, eating out) to the text data and stores it in a database.

[0919] Step 8:

[0920] The server performs natural language processing on the stored data, including tagging and keyword extraction.

[0921] Step 9:

[0922] When the designated posting time arrives (e.g., every day, when a specific event occurs), the server uses a generation AI based on the data in the database to generate a script for posting on social media.

[0923] Step 10:

[0924] The server sends the generated manuscript to the terminal, which displays it to the user and prompts them to check and edit it.

[0925] Step 11:

[0926] The user can check and edit the displayed manuscript. If necessary, press the "Submit" button to finalize the manuscript.

[0927] Step 12:

[0928] The device sends the finalized manuscript to the server, which then calls the API of each SNS platform and posts the manuscript.

[0929] Step 13:

[0930] The server checks whether the posting was successful or not, and sends the result back to the device to notify the user (e.g., "Posting was successful!").

[0931] Example 1

[0932] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0933] In modern society, posting to social networking sites has become a part of users' daily lives, but posting easily can be difficult given their busy schedules. The effort of manually entering text and the burden of thinking up appropriate phrases can be stressful for users. Furthermore, managing the consistency and timing of social networking posts can be difficult, making effective communication difficult. The present invention aims to solve these problems by providing a system that automates social networking posts and significantly reduces the burden on users.

[0934] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0935] In this invention, the server includes means for inputting daily events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for storing the text data in a database, means for analyzing the text data in the database and generating a manuscript for posting to an SNS, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, means for posting the manuscript confirmed and edited by the user to a specified SNS platform, and means for creating prompt sentences using a generative AI model to generate post content in an automated process, thereby enabling users to easily record daily events by voice and automatically post them to an SNS.

[0936] A "user" is someone who uses the system to input everyday events by voice and post them on social media.

[0937] A "terminal" is a device used by a user to perform voice input, such as a smartphone or tablet.

[0938] "Voice data" refers to digital data that is a recording of what a user says to a device.

[0939] "Text data" refers to data of character information converted from voice data using voice recognition technology.

[0940] A "database" is an information system for storing and managing converted text data.

[0941] A "generative AI model" is a machine learning model that uses natural language processing technology to generate specific output (in this case, a manuscript for posting on social media) from input data.

[0942] A "prompt" is an instruction given to a generative AI model, a piece of text that acts as a guide to achieving a specific output.

[0943] "Tagging" is the process of assigning relevant keywords and categories to text data.

[0944] "Keyword extraction" is the process of automatically identifying and extracting important words and phrases from text data.

[0945] "SNS Platform" means a website or application that provides social networking services and allows users to post content thereon.

[0946] A "posting schedule" is a plan to automatically post the generated manuscript for posting to social media based on a specific date, time, or event.

[0947] This invention is a system that allows users to talk about everyday events into a device, accumulates the information in a database, analyzes it, and generates drafts for posting to SNS. This system automates the entire process from user voice input to posting to SNS, significantly reducing the burden on users.

[0948] System configuration

[0949] The system of the present invention mainly comprises the following components:

[0950] 1. User's Device

[0951] 2. Server

[0952] 3. Database

[0953] 4. Speech Recognition Engine

[0954] 5. Generative AI Models

[0955] Hardware and software used

[0956] User devices: Electronic devices such as smartphones and tablets are used, allowing users to record daily events by voice.

[0957] Server: A computer system that can be operated in the cloud or on-premise. It processes voice data and manages databases.

[0958] Database: Use a relational database such as MySQL or PostgreSQL to store and manage text data and its metadata.

[0959] Speech recognition engine: Uses speech recognition services such as Google Cloud Speech-to-Text API or AWS Transcribe to convert voice data into text data.

[0960] Generative AI model: Uses generative AI technologies such as OpenAI GPT-3 and Google BERT to generate natural-sounding sentences based on prompts.

[0961] System operation example

[0962] Voice input from the user

[0963] The user speaks about everyday events into the smartphone's microphone, for example, "Today I went to a cafe with a friend and tried a new cake."

[0964] Audio data conversion and display

[0965] The device records this voice data and sends it to the server. The server uses a speech recognition engine to convert this voice data into text data. The converted text is displayed as "Today I went to a cafe with a friend and tried a new cake," and is sent back to the device.

[0966] Data storage and analysis

[0967] The user checks the displayed text and makes corrections as necessary. Once the process is complete, the text data is sent back to the server and saved in a database. The server adds a date and category (e.g., eating out, friends) to the data and stores it in the database. The server also tags the data and extracts keywords.

[0968] Manuscript generation

[0969] At a specified time (e.g., 7 p.m. every day), the server retrieves the saved text data and uses the generative AI model to generate a script for posting on social media. The generative AI model generates a script based on the prompt: "User's voice input: Today I went to a cafe with a friend and tried the new cake. Please generate a script for posting on social media."

[0970] Posts and Notifications

[0971] The generated message is displayed on the device in the form of "Yesterday I went to a cafe with a friend and tried the new cake. It was delicious." Once the user confirms and edits it, the device sends the finalized data to the server. The server posts the message using the API of the SNS platform and notifies the device of the result.

[0972] This system allows users to easily record their daily events using voice and automatically post them to social media, saving users time and effort and enabling consistent social media posting.

[0973] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0974] Step 1:

[0975] The user speaks about everyday events into the device. The user's voice input is recorded on the device through the microphone of the smartphone or tablet. This generates voice data. Input: User's voice / Output: Recorded voice data.

[0976] Step 2:

[0977] The device sends the recorded audio data to the server. The device securely uploads the audio data to the server using the HTTPS protocol. Input: Audio data / Output: Audio data transferred to the server.

[0978] Step 3:

[0979] The server uses a speech recognition engine to convert the voice data into text data. Here, we use the Google Cloud Speech-to-Text API or AWS Transcribe. The server sends the voice data to the cloud and receives text data as a response. Input: Voice data / Output: Text data.

[0980] Step 4:

[0981] The server sends the converted text data back to the terminal, which displays it on the screen. The user can check the content and correct it if necessary. Input: Text data / Output: Text data displayed on the terminal.

[0982] Step 5:

[0983] The user checks and edits the displayed text data, and confirms the text data once the corrections are complete. Input: Displayed text data / Output: Corrected and confirmed text data.

[0984] Step 6:

[0985] The device sends the confirmed text data to the server. The server receives this data, adds dates and categories (e.g., eating out, friends), and stores it in a database. The server also performs tagging and keyword extraction. Input: Confirmed text data / Output: Data stored in the database.

[0986] Step 7:

[0987] At the specified timing, the server retrieves the text data from the database and generates a prompt sentence for the generative AI model. A generative AI model (e.g., OpenAI GPT-3) is used to create natural-sounding sentences based on the input data. Input: Text data / Output: Generated manuscript for posting on social media.

[0988] Step 8:

[0989] The server sends the generated manuscript for posting to the SNS to the terminal. The terminal displays the manuscript to the user and prompts the user to confirm and edit it. Input: Generated manuscript / Output: Manuscript displayed on the terminal.

[0990] Step 9:

[0991] The user reviews and edits the generated manuscript and decides on the final submission. Input: Displayed manuscript / Output: Reviewed and edited final manuscript.

[0992] Step 10:

[0993] The device sends the final manuscript to the server. The server posts the manuscript using the SNS platform's API. Once posting is complete, the result is notified to the device. Input: Final manuscript / Output: Content posted to SNS and notified result.

[0994] This system allows users to easily record everyday events and automatically post them to social media through a series of steps.

[0995] (Application example 1)

[0996] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0997] In the past, users had to write down their daily events and post them manually on social media or blogs, which was time-consuming and difficult to automate. Furthermore, there was a lack of a way to store records over the long term and organize them systematically, making it difficult to effectively manage content. This made it difficult to use, especially for busy users, and prevented them from efficiently recording and sharing their daily events.

[0998] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0999] In this invention, the server includes means for inputting daily events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for saving the text data in a database, means for analyzing the text data in the database and generating a blog post manuscript, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, and means for posting the confirmed and edited manuscript to a designated blog platform. This enables users to easily record daily events, save and manage them as text data, and automatically post them as blog posts.

[1000] A "device" is an electronic device that a user uses to communicate everyday events.

[1001] "Voice data" refers to data that records what the user has said as voice.

[1002] "Text data" is character information obtained by analyzing voice data.

[1003] A "database" is a digital recording system for storing text data for later analysis and retrieval.

[1004] "Analysis" is the process of processing the text data in the database and extracting specific information or patterns.

[1005] A "blog article manuscript" is a document for blog posting that is generated from analyzed text data and has not yet been reviewed or edited by a user.

[1006] "Checking and editing" refers to the process in which the user looks at the generated manuscript, checks the content, and makes corrections as necessary.

[1007] A "blog platform" is an online service that allows users to post and publish blog articles.

[1008] "Tagging" is a method of assigning specific keywords or categories to text data to make it easier to search and organize.

[1009] "Keyword extraction" is a method of extracting important words from text data and using them for data analysis and search.

[1010] "Set posting schedule" is a function that allows users to set the timing for automatically posting blog articles at a specified date and time.

[1011] This invention provides a system that can effectively record daily events that users talk about to a terminal and automatically generate and post blog articles. Specifically, it is configured as follows.

[1012] Users speak into a device such as a smartphone to input voice data about everyday events. The device records the user's voice and sends the voice data to a server. The server uses a voice recognition engine to convert the voice data into text data, which is then sent back to the device and displayed on the screen. The user can check the text data and make corrections as necessary.

[1013] Once the text data has been corrected, it is sent back to the server. The server then tags the text data, extracts keywords, and stores them in a database. The text data stored in the database is then analyzed using a generative AI model to generate a draft blog post. This draft is then sent back to the device in a format that is easy for the user to understand.

[1014] The device displays the generated blog post draft to the user and provides an interface that prompts the user to review and edit it. Once the user has reviewed the draft and finished editing, they confirm the posting. The server receives the confirmed draft and posts it to the blog platform at the specified time. The success or failure of the posting is notified to the device and conveyed to the user.

[1015] In this system, the following hardware and software are used:

[1016] Hardware: Smartphone (device).

[1017] software:

[1018] For voice recognition, the speech_recognition library is used, utilizing Google's voice recognition service.

[1019] For text generation, we use the transformers library and the GPT-3 model.

[1020] Data can be stored using a local file system or a cloud database (e.g., Amazon S3).

[1021] For social media and blog posts, the Python requests library is used to call designated platform APIs.

[1022] As a concrete example, consider the case where a user says, "I went to a new cafe today. I had some really good coffee." This speech is converted into text and stored on a server. It is then analyzed using a generative AI model, and a blog post like the one below is generated.

[1023] Example prompt sentence:

[1024] "Generate a blog post based on the following text: I went to a new cafe today and had some really good coffee."

[1025] The generated blog post is displayed on the device, and once the user has confirmed and edited it, it is automatically posted to the blog platform. By using this system, users can easily and hassle-freely share their everyday events as blog posts.

[1026] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1027] Processing steps of the system that realizes the application example

[1028] Step 1:

[1029] The user inputs voice data

[1030] A user speaks into a device such as a smartphone. This voice data is input as an audio file of what the user said. For example, "I went to a new cafe today. I had some really good coffee." This voice data is recorded on the device.

[1031] Step 2:

[1032] Converting audio data into text data

[1033] The device sends the recorded voice data to the server. The server uses a voice recognition engine (for example, Google's voice recognition service) to convert the voice data into text data. In this process, the voice data (input) is converted into text information (output). The converted text data, "I went to a new cafe today. I had some very delicious coffee," is sent back and displayed on the device's screen.

[1034] Step 3:

[1035] Check and correct text data

[1036] The user checks the text data displayed on the terminal and makes corrections as necessary. For example, "I drank a very delicious coffee" is changed to "I drank a very delicious espresso." In this procedure, the user looks at the text data (input), makes the necessary corrections (data processing), and obtains the corrected text data (output).

[1037] Step 4:

[1038] Storing text data in a database

[1039] The text data confirmed and corrected by the user is sent back to the server from the device. The server tags the text data and extracts keywords (for example, adding tags such as "cafe" or "coffee"), and stores it in a database. In this process, the text data (input) is organized and stored as tagged text data (output).

[1040] Step 5:

[1041] Generate a blog post draft

[1042] At a specified time (for example, a time set by the user each night), the server uses a generative AI model (for example, GPT-3) to generate a draft blog post based on the tagged text data in the database. In this process, the stored text data (input) is passed through a natural language generation algorithm to generate a draft blog post (output).

[1043] Step 6:

[1044] View, check and edit the generated manuscript

[1045] The generated draft of the blog post is sent to the user's device and displayed on the screen. The user can review this draft and make any necessary corrections. For example, the draft of a blog post might read, "I went to a new cafe today. I enjoyed a delicious espresso." In this step, the generated draft (input) is reviewed and edited by the user (data processing) before becoming the final draft (output).

[1046] Step 7:

[1047] Post your blog post to your preferred blogging platform

[1048] After the user has reviewed and edited the blog post, it is sent from the device to the server, which then calls the API of the designated blog platform to post it. For example, an article such as "I went to a new cafe today. I enjoyed a delicious espresso" is posted to the blog. The success or failure of the post is again reported to the device and communicated to the user. Here, the final draft (input) is posted to the blog platform (data calculation and manipulation), and the posting result (output) is returned to the user.

[1049] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1050] This system allows users to talk about everyday events, stores the content in a database, analyzes it, and generates a draft for posting to an SNS. It also recognizes the user's emotions and reflects them in the generated text. This system automates the entire process from user input to posting to an SNS, not only reducing the burden on the user but also enabling posts that take emotions into consideration.

[1051] System configuration and operation

[1052] 1. User voice input

[1053] The user talks to the device about everyday events. For example, "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner."

[1054] 2. Audio data conversion

[1055] The device records the user's voice and sends the voice data to the server. The server uses a speech recognition engine to convert the voice data into text data. The converted text data is sent back to the device and displayed on the screen. The user can check the text and make corrections if necessary.

[1056] 3. Emotional Recognition

[1057] The server sends the voice data to the emotion engine, which then recognizes the emotion from the user's voice. The emotion engine then analyzes the user's emotional state based on the tone of voice and vocabulary selection, and adds the results to the text data.

[1058] 4. Saving to the database

[1059] The device sends the text data it has confirmed to a server. The server then adds a date and category (e.g., work, eating out) to the text data and stores it in a database. The stored data is then organized through natural language processing, including tagging and keyword extraction. Emotional information added through emotion recognition is also stored.

[1060] 5. Creating a manuscript for posting on social media

[1061] When the posting timing specified by the user (e.g., daily, when a specific event occurs) arrives, the server uses generative AI to generate a script for posting to social media based on the text data in the database. The generated script takes emotional information into account. For example, if positive emotions are detected, positive expressions will be emphasized.

[1062] 6. User confirmation and editing

[1063] The device displays the generated script for posting to social media to the user and provides an interface for reviewing and editing as necessary. For example, it could be displayed in the form of "Yesterday, my big presentation was a success! Afterwards, I enjoyed a delicious Italian meal with friends. It was a fulfilling day."

[1064] 7. Posting to social media

[1065] Once the user has finished checking and editing, they confirm the post. The device then receives this and sends the confirmed manuscript to the server. The server then calls the API of each SNS platform and executes the post. The success or failure of the post is again notified to the device and conveyed to the user.

[1066] Specific examples

[1067] For example, consider the case where a user says, "Today I went on a picnic with my family and it was so much fun. Then I saw a new movie."

[1068] 1. The user speaks.

[1069] 2. The device records the audio and sends it to the server.

[1070] 3. The server uses a speech recognition engine to convert this into text data: "Today I went on a picnic with my family and had a great time. Then I saw a new movie."

[1071] 4. The server uses an emotion engine to recognize positive emotions from the voice and tag the text data as "joy."

[1072] 5. Display the text data on the terminal.

[1073] 6. The user checks, corrects, and approves.

[1074] 7. The device sends the text data to the server, where it is tagged, keywords are extracted, and the data is stored in a database.

[1075] 8. Based on the posting timing, the server generates a script for posting to social media: "Yesterday I went on a picnic with my family. It was so much fun. Afterwards, I also saw a new movie. It was a great day."

[1076] 9. The terminal displays the generated manuscript, which the user can check and edit.

[1077] 10. The user submits the finalized manuscript.

[1078] 11. The server posts to the SNS platform and notifies the device of the results.

[1079] This system allows users to easily record and post everyday events, and also generates engaging content that reflects their emotions.

[1080] The processing flow will be explained below.

[1081] Step 1:

[1082] The user talks to the device about everyday events. For example, "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner."

[1083] Step 2:

[1084] The device will record the user's voice and the recorded voice data will be temporarily stored on the device.

[1085] Step 3:

[1086] The device sends the recorded audio data to the server, where it is converted into an appropriate format and sent.

[1087] Step 4:

[1088] The server receives the voice data, which is then input into a voice recognition engine and converted into text data.

[1089] Step 5:

[1090] The server inputs the converted text data into the emotion engine, which analyzes the user's emotions based on the tone and content of the voice and adds emotional information.

[1091] Step 6:

[1092] The server sends text data and emotional information back to the device. The text data, such as "I had a big presentation at work today, and it went well. For dinner, I went out to eat Italian food with a friend," is tagged with emotional information.

[1093] Step 7:

[1094] The device displays the text data and emotional information to the user, who can then review and correct it if necessary.

[1095] Step 8:

[1096] The user checks and corrects the text data and presses the "Confirm" or "Send" button.

[1097] Step 9:

[1098] The device sends the confirmed and corrected text data to the server, which adds date and category information to the text data and stores it in a database.

[1099] Step 10:

[1100] The server then performs tagging and keyword extraction on the data stored in the database, including emotional information.

[1101] Step 11:

[1102] When the designated posting time (e.g., every day or when a specific event occurs) arrives, the server generates a script for posting to social media based on the data in the database. Emotional information is also taken into account, with positive emotions emphasized.

[1103] Step 12:

[1104] The server sends the generated script to the terminal, generating a script in the format "Yesterday's presentation was a success! Afterwards, I enjoyed some delicious Italian food with my friends. It was a fulfilling day."

[1105] Step 13:

[1106] The terminal displays the generated manuscript to the user and provides an interface that prompts them to check and edit it.

[1107] Step 14:

[1108] The user checks the manuscript, edits it if necessary, and presses the "Submit" button.

[1109] Step 15:

[1110] The device sends the finalized manuscript to the server, which then calls the API of each SNS platform and posts the generated manuscript.

[1111] Step 16:

[1112] The server checks whether the posting was successful or not and returns the result to the device.

[1113] Step 17:

[1114] The device notifies the user of the posting result (e.g., "Posting successful!").

[1115] This series of processes allows users to easily record everyday events and post engaging content that reflects their emotions on social media.

[1116] Example 2

[1117] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1118] Conventional SNS posting systems have the problem that users must manually input daily events and create posts that reflect their emotions, which requires a great deal of time and effort. Furthermore, the technical means for reading and appropriately reflecting emotions are insufficient, making it difficult to properly express the user's intended message. Furthermore, the ability to adjust the timing of posts is limited, reducing user convenience. The purpose of this invention is to solve these problems and provide a system that automatically posts a variety of emotionally reflective posts while reducing the burden on users.

[1119] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1120] In this invention, the server includes means for inputting daily events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for saving the text data in a database, means for organizing the saved text data by adding date and category information, means for recognizing emotions from the voice data and adding the results to the text data, means for analyzing the text data in the database and generating a manuscript for posting to an SNS using generation technology, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, means for posting the manuscript confirmed and edited by the user to a specified SNS platform, and means for notifying the user of the success or failure of the posting. This enables users to easily record daily events and automatically generate and post sentences that reflect their emotions.

[1121] "Voice data" refers to data that records what a user says to a terminal as a voice signal.

[1122] "Text data" refers to textual information data obtained by analyzing and converting voice data.

[1123] A "terminal" is a device that a user uses to input voice, and that has the function of recording and displaying voice.

[1124] A "server" is a device or system that processes voice data sent from a terminal, converts it into text data, stores it in a database, analyzes it, and so on.

[1125] A "database" is a data management system for organizing and storing text data and associated information (e.g., dates, categories, and emotion tags).

[1126] "Emotion recognition" is a technology that analyzes and extracts a user's emotional state from voice and text data.

[1127] "Generation technology" refers to technology for generating specific formats and content based on information in a database, and is primarily used to generate manuscripts for posting on social media.

[1128] "SNS Platform" means a system that provides social networking services that enable users to communicate with other users over the Internet.

[1129] "Posting" refers to the act of sending the generated manuscript for posting to an SNS platform on the Internet and making it public.

[1130] "Notification" refers to the act of the server sending specific information (e.g., posting success or failure) to the terminal to inform the user.

[1131] This system allows users to talk about everyday events into a device, accumulates the content in a database, analyzes it, and generates a manuscript for posting to SNS. It also recognizes the user's emotions and reflects them in the generated text. This system automates the entire process from user input to posting to SNS, reducing the burden on the user and enabling posts that take emotions into consideration.

[1132] First, the user talks about everyday events into a dedicated device. The device uses a built-in microphone to record the user's voice and save it as audio data. The recorded audio data is then sent to a server via the Internet.

[1133] The server uses the Google Cloud Speech-to-Text API to convert the received voice data into text data, which is then sent back to the device via the Internet and displayed for the user to review. The user can then review the text data and make corrections as necessary.

[1134] Next, the server sends the voice data to the Microsoft Azure Cognitive Services emotion recognition API for emotion recognition. The emotion recognition API analyzes the voice tone and vocabulary choice to extract the user's emotional state. The analysis results are added to the text data as emotion tags. The text data is then stored in a database. When storing the data, the server uses a natural language processing library (e.g., NLTK) to add date and category information to the text data and organize it.

[1135] When the user's designated posting time arrives, the server retrieves the target text data from the database and generates a script for posting to social media using a generative AI model such as OpenAI's GPT-4. At this time, the server inputs the following prompt to the generative AI model:

[1136] "Generate a social media post based on the following text and sentiment tags. Emphasize positive sentiment and create content that will interest your readers."

[1137] For example, if the server generates a script based on the text data "I had a big presentation at work today, which was a great success. For dinner, I went out to eat Italian food with friends," and the emotion tag "joy," the resulting script would read, "Yesterday's big presentation was a success! Afterwards, I enjoyed some delicious Italian food with friends. It was a fulfilling day."

[1138] The generated manuscript is sent to the terminal and displayed to the user. The user checks the manuscript and makes edits as necessary. When the user has finally finished checking and editing, the terminal sends the manuscript to the server.

[1139] The server calls the API of various social media platforms (e.g., Twitter API, Facebook Graph API) and posts the finalized manuscript on the Internet. The success or failure of this posting is again notified to the device and conveyed to the user.

[1140] The above process allows users to easily record their daily events and effectively post them on social media, reflecting their emotions. This system significantly reduces the burden on users and is an effective way to provide engaging content that takes emotions into consideration.

[1141] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1142] Step 1:

[1143] Users talk to the device about everyday events.

[1144] Specifically, the device's built-in microphone records the user's voice.

[1145] Input: User's voice

[1146] Output: Audio data (recording file)

[1147] Step 2:

[1148] The device sends the recorded audio data to the server.

[1149] Specifically, the audio file is uploaded to a server via the Internet.

[1150] Input: Audio data

[1151] Output: Audio data file on the server

[1152] Step 3:

[1153] The server uses the Google Cloud Speech-to-Text API to convert the audio data into text data.

[1154] Specifically, audio data is sent to the API and the returned text data is obtained.

[1155] Input: Audio data

[1156] Output: Text data (converted character information)

[1157] Step 4:

[1158] The server retransmits the converted text data to the terminal.

[1159] Specifically, text data is downloaded to the terminal via the Internet.

[1160] Input: Text data

[1161] Output: Text data displayed on the terminal

[1162] Step 5:

[1163] The user checks the text data displayed on the terminal and corrects it if necessary.

[1164] Specifically, use a text editor to correct any errors or unnecessary parts.

[1165] Input: Text data

[1166] Output: Text data that the user has confirmed and corrected

[1167] Step 6:

[1168] The terminal transmits the text data corrected by the user to the server again.

[1169] Specifically, the corrected text data is uploaded to a server via the Internet.

[1170] Input: Corrected text data

[1171] Output: Modified text data on the server

[1172] Step 7:

[1173] The server sends the corrected text data to the Microsoft Azure Cognitive Services emotion recognition API for emotional analysis.

[1174] Specifically, text data is sent to the API and the sentiment analysis results are obtained.

[1175] Input: Text data

[1176] Output: Emotion-tagged text data

[1177] Step 8:

[1178] The server adds date and category information to the text data and stores it in a database.

[1179] Specifically, the text data is inserted into the database as a new record.

[1180] Input: emotion-tagged text data

[1181] Output: Text data stored in the database

[1182] Step 9:

[1183] When the specified posting time arrives, the server uses OpenAI's generative AI model to generate a manuscript for posting on social media.

[1184] Specifically, text data is retrieved from a database and prompt sentences are input into a generative AI model to generate a manuscript.

[1185] Input: Text data in the database, prompt statements

[1186] Output: Generated manuscript for posting to social media

[1187] Step 10:

[1188] The generated manuscript is sent to the terminal and displayed to the user.

[1189] Specifically, the generated manuscript is downloaded to the terminal via the Internet.

[1190] Input: Generated SNS posting manuscript

[1191] Output: Manuscript for posting to social media displayed on the device

[1192] Step 11:

[1193] The user reviews the manuscript and edits it if necessary.

[1194] Specifically, the text editor is used again to make corrections and additions to the manuscript.

[1195] Input: Manuscript to post on social media

[1196] Output: User-confirmed and edited manuscript for posting on social media

[1197] Step 12:

[1198] The terminal transmits the final manuscript to the server.

[1199] Specifically, the final manuscript is uploaded to a server via the Internet.

[1200] Input: Confirmed and edited manuscript for posting to social media

[1201] Output: Final manuscript on the server

[1202] Step 13:

[1203] The server calls the APIs of various social media platforms and posts the finalized manuscript.

[1204] Specifically, you submit your manuscript using the posting API of each SNS and retrieve the results.

[1205] Input: Final manuscript

[1206] Output: Posting results to social media

[1207] Step 14:

[1208] The terminal is notified of the success or failure of the posting.

[1209] Specifically, a notification message is sent to the terminal and displayed to the user.

[1210] Input: Post results

[1211] Output: Notification message displayed on the terminal

[1212] (Application example 2)

[1213] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1214] In conventional customer service, it has been difficult to record and analyze customer interactions in real time and use the results as feedback or reviews. In addition, there has been a lack of a way to efficiently generate reviews and social media posts that reflect customer sentiment, so an effective tool is needed to improve customer satisfaction.

[1215] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting everyday events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for saving the text data in a database, means for analyzing the text data in the database and generating a manuscript for posting on an SNS, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, means for posting the manuscript confirmed and edited by the user to a specified SNS platform, and means for recognizing emotions from the voice data and reflecting the recognized emotional information in the generated text. This makes it possible to record the content of conversations with customers in real time and generate feedback and reviews that reflect the emotional information.

[1216] "Terminal" refers to a device that allows users to input voice data and check and edit text data and emotion recognition results.

[1217] "Voice data" refers to information recorded in audio format about everyday events spoken by a user.

[1218] "Text data" refers to information that has been analyzed and converted into text form from audio data.

[1219] The "database" is a data repository for storing generated text data and emotion recognition results.

[1220] "Analysis" refers to the process of understanding and extracting meaning from the text data in a database.

[1221] A "script for posting on social media" is a piece of text created using a generative AI model based on text data, intended for posting on social media.

[1222] "Emotion recognition" is the process of analyzing a user's emotional state from their voice data and obtaining the results.

[1223] A "generative AI model" is an artificial intelligence technology that generates sentences in a specific format based on input data.

[1224] A "prompt sentence" is the initial input sentence that a generative AI model uses when generating a sentence.

[1225] A "wearable device" is an information terminal that is worn by the user.

[1226] The system is designed to enable users to record customer interactions through wearable devices such as smart glasses and automatically generate transcripts for reviews and social media posts. The system consists of the following main components:

[1227] Hardware and Software Configuration

[1228] 1. Device:

[1229] Wearable devices such as smart glasses.

[1230] It is equipped with a microphone and has the ability to receive voice input.

[1231] Record customer interactions in real time.

[1232] 2. Server:

[1233] It runs a speech recognition engine, a sentiment analysis engine, and a generative AI model.

[1234] Convert the voice data into text data and analyze emotions.

[1235] The text data is stored in a database for later analysis and tagging.

[1236] System processing flow

[1237] 1. Voice input:

[1238] The device records the user's voice in real time and sends it to the server as audio data.

[1239] 2. Audio data conversion:

[1240] The server uses a speech recognition engine to convert the voice data into text data, which is then sent back to the device and displayed to the user.

[1241] 3. Emotion Recognition:

[1242] The server recognizes emotions from the recorded voice data and adds emotional information to it using an emotion analysis engine.

[1243] 4. Save to database:

[1244] The text data and sentiment information are stored in a database, organized by date and category (e.g., product reviews, service ratings).

[1245] 5. Creating a manuscript for posting on social media:

[1246] At the specified timing, the server extracts text data from the database and uses the generative AI model to generate a script for posting on social media. This script reflects emotional information. For example, if "positive emotions" are detected, positive expressions will be emphasized.

[1247] 6. User review and editing:

[1248] The manuscript is displayed on the user's terminal, and the user can check it and edit it if necessary.

[1249] 7. Posting to social media:

[1250] The terminal sends the manuscript that the user has finalized to the server, and the server posts it via the API of each SNS platform.

[1251] Specific examples

[1252] For example, in a physical store, imagine a scenario where a salesperson asks a customer, "Today, you tried this new perfume. What do you think?", and the customer replies, "It smells great! I'd like to come again." The system records this conversation and processes it as follows:

[1253] Voice input: The store clerk's smart glasses record the conversation.

[1254] Conversion of voice data: The server converts this into text data such as "You tried this new perfume today. What did you think?" and "It smelled great! I'd like to come again."

[1255] Emotion recognition: Positive emotions are detected.

[1256] Using a generative AI model: Generate a review using the prompt "Review when emotions are positive: The scent was amazing! I'd love to come back. How can I recreate such a wonderful day?"

[1257] User review and editing: Store associates can review and edit generated reviews using smart glasses.

[1258] Posting to social media: After the review is verified and edited, it will be posted to each social media platform.

[1259] Based on these example prompts, the system can quickly generate engaging, emotionally sensitive content, contributing to improved customer satisfaction.

[1260] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1261] Step 1:

[1262] The device records the user's voice in real time and sends the voice data to the server. The input is the user's voice, and the output is the voice data sent to the server. Specifically, the voice input is captured using the microphone in the smart glasses, and the data is compressed and sent to the server.

[1263] Step 2:

[1264] The server uses a speech recognition engine to convert the voice data into text data. The input is the voice data sent in step 1, and the output is the converted text data. Specifically, Google's speech recognition API is used to convert the voice data into Japanese text.

[1265] Step 3:

[1266] The terminal receives the converted text data and displays it to the user. The input is the text data sent from the server, and the output is the text data displayed on the terminal's display. Specifically, the text is displayed on the display of the smart glasses and confirmed by the user.

[1267] Step 4:

[1268] The server recognizes emotions from voice data and adds that emotional information to text data. The input is text data and voice data, and the output is text data with emotional information. Specifically, it uses Hugging Face's emotion analysis engine to analyze emotions from voice tone and keywords, and adds that information to the text data.

[1269] Step 5:

[1270] The server stores text data and emotion information in a database. The input is text data with emotion information, and the output is the data stored in the database. Specifically, data organized by category is stored in the database.

[1271] Step 6:

[1272] At the specified time, the server extracts text data from the database and uses the generative AI model to generate a manuscript for posting on social media. The input is the text data in the database, and the output is the generated manuscript for posting on social media. Specifically, the generative AI model is input with a prompt sentence: "Review when emotions are positive: The scent was amazing! I'd like to come again. How can I recreate such a good day again?" and outputs the generated text.

[1273] Step 7:

[1274] The device receives the generated SNS post manuscript and displays it to the user. The input is the SNS post manuscript sent from the server, and the output is the manuscript displayed on the device's display. Specifically, it is displayed on the smart glasses display for the user to check and edit.

[1275] Step 8:

[1276] The device sends the manuscript edited by the user to be posted to the server, and the server executes the posting via the API of each SNS platform. The input is the edited manuscript to be posted to the SNS, and the output is the result of posting to each SNS platform. Specifically, the SNS API is called to post, and the device is notified of the success or failure of the posting.

[1277] Step 9:

[1278] The device notifies the user of the results of the post to the SNS. The input is the post result sent from the server, and the output is the result notification displayed on the device's display. Specifically, a message such as "Posting successful" is displayed on the smart glasses' display.

[1279] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1280] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1281] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1282] [Fourth embodiment]

[1283] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1284] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1285] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1286] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1287] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1288] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1289] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1290] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1291] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1292] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1293] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1294] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1295] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1296] This invention is a system that allows users to talk about everyday events, accumulates them in a database, and analyzes each piece of data to generate drafts for posting to social media. This system automates the entire process from user input to posting to social media, significantly reducing the burden on users.

[1297] System configuration and operation

[1298] 1. User voice input

[1299] The user talks to the device about everyday events, such as, "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner."

[1300] 2. Audio data conversion

[1301] The device records the user's voice and sends the voice data to the server. The server uses a speech recognition engine to convert the voice data into text data. The converted text data is sent back to the device and displayed on the screen. The user can check the text and make corrections if necessary.

[1302] 3. Saving to the database

[1303] The device sends the text data to the server, which then adds a date and category (e.g., work, eating out) to the text data and stores it in a database. The stored data is then organized using natural language processing such as tagging and keyword extraction.

[1304] 4. Creating a manuscript for posting on social media

[1305] When the posting timing specified by the user (e.g., every day, or when a specific event occurs) arrives, the server uses generative AI to generate a script for posting to SNS based on the text data in the database. This script is sent to the device in a format that the user can easily understand.

[1306] 5. User confirmation and editing

[1307] The device displays the generated script for posting to social media to the user and provides an interface for reviewing and editing as necessary. For example, it could be displayed in the form of "Yesterday, my big presentation was a success! Afterwards, I enjoyed a delicious Italian meal with friends. It was a fulfilling day."

[1308] 6. Posting to social media

[1309] Once the user has finished checking and editing, they confirm the post. The device then receives this and sends the confirmed manuscript to the server. The server then calls the API of each SNS platform and executes the post. The success or failure of the post is again notified to the device and conveyed to the user.

[1310] Specific examples

[1311] For example, a user might say: "Today I went on a picnic with my family and had a great time. Then I saw a new movie."

[1312] 1. The user speaks.

[1313] 2. The device records the audio and sends it to the server.

[1314] 3. The server uses a speech recognition engine to convert this into text data: "Today I went on a picnic with my family and had a great time. Then I saw a new movie."

[1315] 4. Display the text data on the terminal.

[1316] 5. The user checks, corrects, and approves.

[1317] 6. The device sends the text data to the server, where it is tagged, keywords are extracted, and the data is stored in a database.

[1318] 7. At the specified time, the server generates a script to post to social media: "Yesterday, I went on a picnic with my family. It was a lot of fun. Afterwards, I also saw a new movie."

[1319] 8. The terminal displays the generated manuscript, which the user can check and edit.

[1320] 9. The user submits the finalized manuscript.

[1321] 10. The server posts to the SNS platform and notifies the device of the results.

[1322] This system allows users to easily and efficiently record everyday events and post them to social media.

[1323] The processing flow will be explained below.

[1324] Step 1:

[1325] The user talks about everyday events into the device, inputting content such as "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner" through the voice input interface.

[1326] Step 2:

[1327] The device records the user's voice and sends the captured voice data to the server via the voice input interface.

[1328] Step 3:

[1329] The server receives the voice data, sends it to the voice recognition engine, and converts it into text data.

[1330] Step 4:

[1331] The server returns the converted text data to the device. The returned text data is displayed as "I had a big presentation at work today, and it was a success. For dinner, I went out to eat Italian food with a friend."

[1332] Step 5:

[1333] The terminal displays the text data to the user and asks for confirmation, after which the user can check the text data and correct it if necessary.

[1334] Step 6:

[1335] The user confirms and modifies the text data by pressing the "Confirm" button.

[1336] Step 7:

[1337] The device sends the text data that has been confirmed to the server, which then adds the date and category (e.g., work, eating out) to the text data and stores it in a database.

[1338] Step 8:

[1339] The server performs natural language processing on the stored data, including tagging and keyword extraction.

[1340] Step 9:

[1341] When the designated posting time arrives (e.g., every day, when a specific event occurs), the server uses a generation AI based on the data in the database to generate a script for posting on social media.

[1342] Step 10:

[1343] The server sends the generated manuscript to the terminal, which displays it to the user and prompts them to check and edit it.

[1344] Step 11:

[1345] The user can check and edit the displayed manuscript. If necessary, press the "Submit" button to finalize the manuscript.

[1346] Step 12:

[1347] The device sends the finalized manuscript to the server, which then calls the API of each SNS platform and posts the manuscript.

[1348] Step 13:

[1349] The server checks whether the posting was successful or not, and sends the result back to the device to notify the user (e.g., "Posting was successful!").

[1350] Example 1

[1351] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1352] In modern society, posting to social networking sites has become a part of users' daily lives, but posting easily can be difficult given their busy schedules. The effort of manually entering text and the burden of thinking up appropriate phrases can be stressful for users. Furthermore, managing the consistency and timing of social networking posts can be difficult, making effective communication difficult. The present invention aims to solve these problems by providing a system that automates social networking posts and significantly reduces the burden on users.

[1353] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1354] In this invention, the server includes means for inputting daily events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for storing the text data in a database, means for analyzing the text data in the database and generating a manuscript for posting to an SNS, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, means for posting the manuscript confirmed and edited by the user to a specified SNS platform, and means for creating prompt sentences using a generative AI model to generate post content in an automated process, thereby enabling users to easily record daily events by voice and automatically post them to an SNS.

[1355] A "user" is someone who uses the system to input everyday events by voice and post them on social media.

[1356] A "terminal" is a device used by a user to perform voice input, such as a smartphone or tablet.

[1357] "Voice data" refers to digital data that is a recording of what a user says to a device.

[1358] "Text data" refers to data of character information converted from voice data using voice recognition technology.

[1359] A "database" is an information system for storing and managing converted text data.

[1360] A "generative AI model" is a machine learning model that uses natural language processing technology to generate specific output (in this case, a manuscript for posting on social media) from input data.

[1361] A "prompt" is an instruction given to a generative AI model, a piece of text that acts as a guide to achieving a specific output.

[1362] "Tagging" is the process of assigning relevant keywords and categories to text data.

[1363] "Keyword extraction" is the process of automatically identifying and extracting important words and phrases from text data.

[1364] "SNS Platform" means a website or application that provides social networking services and allows users to post content thereon.

[1365] A "posting schedule" is a plan to automatically post the generated manuscript for posting to social media based on a specific date, time, or event.

[1366] This invention is a system that allows users to talk about everyday events into a device, accumulates the information in a database, analyzes it, and generates drafts for posting to SNS. This system automates the entire process from user voice input to posting to SNS, significantly reducing the burden on users.

[1367] System configuration

[1368] The system of the present invention mainly comprises the following components:

[1369] 1. User's Device

[1370] 2. Server

[1371] 3. Database

[1372] 4. Speech Recognition Engine

[1373] 5. Generative AI Models

[1374] Hardware and software used

[1375] User devices: Electronic devices such as smartphones and tablets are used, allowing users to record daily events by voice.

[1376] Server: A computer system that can be operated in the cloud or on-premise. It processes voice data and manages databases.

[1377] Database: Use a relational database such as MySQL or PostgreSQL to store and manage text data and its metadata.

[1378] Speech recognition engine: Uses speech recognition services such as Google Cloud Speech-to-Text API or AWS Transcribe to convert voice data into text data.

[1379] Generative AI model: Uses generative AI technologies such as OpenAI GPT-3 and Google BERT to generate natural-sounding sentences based on prompts.

[1380] System operation example

[1381] Voice input from the user

[1382] The user speaks about everyday events into the smartphone's microphone, for example, "Today I went to a cafe with a friend and tried a new cake."

[1383] Audio data conversion and display

[1384] The device records this voice data and sends it to the server. The server uses a speech recognition engine to convert this voice data into text data. The converted text is displayed as "Today I went to a cafe with a friend and tried a new cake," and is sent back to the device.

[1385] Data storage and analysis

[1386] The user checks the displayed text and makes corrections as necessary. Once the process is complete, the text data is sent back to the server and saved in a database. The server adds a date and category (e.g., eating out, friends) to the data and stores it in the database. The server also tags the data and extracts keywords.

[1387] Manuscript generation

[1388] At a specified time (e.g., 7 p.m. every day), the server retrieves the saved text data and uses the generative AI model to generate a script for posting on social media. The generative AI model generates a script based on the prompt: "User's voice input: Today I went to a cafe with a friend and tried the new cake. Please generate a script for posting on social media."

[1389] Posts and Notifications

[1390] The generated message is displayed on the device in the form of "Yesterday I went to a cafe with a friend and tried the new cake. It was delicious." Once the user confirms and edits it, the device sends the finalized data to the server. The server posts the message using the API of the SNS platform and notifies the device of the result.

[1391] This system allows users to easily record their daily events using voice and automatically post them to social media, saving users time and effort and enabling consistent social media posting.

[1392] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1393] Step 1:

[1394] The user speaks about everyday events into the device. The user's voice input is recorded on the device through the microphone of the smartphone or tablet. This generates voice data. Input: User's voice / Output: Recorded voice data.

[1395] Step 2:

[1396] The device sends the recorded audio data to the server. The device securely uploads the audio data to the server using the HTTPS protocol. Input: Audio data / Output: Audio data transferred to the server.

[1397] Step 3:

[1398] The server uses a speech recognition engine to convert the voice data into text data. Here, we use the Google Cloud Speech-to-Text API or AWS Transcribe. The server sends the voice data to the cloud and receives text data as a response. Input: Voice data / Output: Text data.

[1399] Step 4:

[1400] The server sends the converted text data back to the terminal, which displays it on the screen. The user can check the content and correct it if necessary. Input: Text data / Output: Text data displayed on the terminal.

[1401] Step 5:

[1402] The user checks and edits the displayed text data, and confirms the text data once the corrections are complete. Input: Displayed text data / Output: Corrected and confirmed text data.

[1403] Step 6:

[1404] The device sends the confirmed text data to the server. The server receives this data, adds dates and categories (e.g., eating out, friends), and stores it in a database. The server also performs tagging and keyword extraction. Input: Confirmed text data / Output: Data stored in the database.

[1405] Step 7:

[1406] At the specified timing, the server retrieves the text data from the database and generates a prompt sentence for the generative AI model. A generative AI model (e.g., OpenAI GPT-3) is used to create natural-sounding sentences based on the input data. Input: Text data / Output: Generated manuscript for posting on social media.

[1407] Step 8:

[1408] The server sends the generated manuscript for posting to the SNS to the terminal. The terminal displays the manuscript to the user and prompts the user to confirm and edit it. Input: Generated manuscript / Output: Manuscript displayed on the terminal.

[1409] Step 9:

[1410] The user reviews and edits the generated manuscript and decides on the final submission. Input: Displayed manuscript / Output: Reviewed and edited final manuscript.

[1411] Step 10:

[1412] The device sends the final manuscript to the server. The server posts the manuscript using the SNS platform's API. Once posting is complete, the result is notified to the device. Input: Final manuscript / Output: Content posted to SNS and notified result.

[1413] This system allows users to easily record everyday events and automatically post them to social media through a series of steps.

[1414] (Application example 1)

[1415] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1416] In the past, users had to write down their daily events and post them manually on social media or blogs, which was time-consuming and difficult to automate. Furthermore, there was a lack of a way to store records over the long term and organize them systematically, making it difficult to effectively manage content. This made it difficult to use, especially for busy users, and prevented them from efficiently recording and sharing their daily events.

[1417] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1418] In this invention, the server includes means for inputting daily events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for saving the text data in a database, means for analyzing the text data in the database and generating a blog post manuscript, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, and means for posting the confirmed and edited manuscript to a designated blog platform. This enables users to easily record daily events, save and manage them as text data, and automatically post them as blog posts.

[1419] A "device" is an electronic device that a user uses to communicate everyday events.

[1420] "Voice data" refers to data that records what the user has said as voice.

[1421] "Text data" is character information obtained by analyzing voice data.

[1422] A "database" is a digital recording system for storing text data for later analysis and retrieval.

[1423] "Analysis" is the process of processing the text data in the database and extracting specific information or patterns.

[1424] A "blog article manuscript" is a document for blog posting that is generated from analyzed text data and has not yet been reviewed or edited by a user.

[1425] "Checking and editing" refers to the process in which the user looks at the generated manuscript, checks the content, and makes corrections as necessary.

[1426] A "blog platform" is an online service that allows users to post and publish blog articles.

[1427] "Tagging" is a method of assigning specific keywords or categories to text data to make it easier to search and organize.

[1428] "Keyword extraction" is a method of extracting important words from text data and using them for data analysis and search.

[1429] "Set posting schedule" is a function that allows users to set the timing for automatically posting blog articles at a specified date and time.

[1430] This invention provides a system that can effectively record daily events that users talk about to a terminal and automatically generate and post blog articles. Specifically, it is configured as follows.

[1431] Users speak into a device such as a smartphone to input voice data about everyday events. The device records the user's voice and sends the voice data to a server. The server uses a voice recognition engine to convert the voice data into text data, which is then sent back to the device and displayed on the screen. The user can check the text data and make corrections as necessary.

[1432] Once the text data has been corrected, it is sent back to the server. The server then tags the text data, extracts keywords, and stores them in a database. The text data stored in the database is then analyzed using a generative AI model to generate a draft blog post. This draft is then sent back to the device in a format that is easy for the user to understand.

[1433] The device displays the generated blog post draft to the user and provides an interface that prompts the user to review and edit it. Once the user has reviewed the draft and finished editing, they confirm the posting. The server receives the confirmed draft and posts it to the blog platform at the specified time. The success or failure of the posting is notified to the device and conveyed to the user.

[1434] In this system, the following hardware and software are used:

[1435] Hardware: Smartphone (device).

[1436] software:

[1437] For voice recognition, the speech_recognition library is used, utilizing Google's voice recognition service.

[1438] For text generation, we use the transformers library and the GPT-3 model.

[1439] Data can be stored using a local file system or a cloud database (e.g., Amazon S3).

[1440] For social media and blog posts, the Python requests library is used to call designated platform APIs.

[1441] As a concrete example, consider the case where a user says, "I went to a new cafe today. I had some really good coffee." This speech is converted into text and stored on a server. It is then analyzed using a generative AI model, and a blog post like the one below is generated.

[1442] Example prompt sentence:

[1443] "Generate a blog post based on the following text: I went to a new cafe today and had some really good coffee."

[1444] The generated blog post is displayed on the device, and once the user has confirmed and edited it, it is automatically posted to the blog platform. By using this system, users can easily and hassle-freely share their everyday events as blog posts.

[1445] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1446] Processing steps of the system that realizes the application example

[1447] Step 1:

[1448] The user inputs voice data

[1449] A user speaks into a device such as a smartphone. This voice data is input as an audio file of what the user said. For example, "I went to a new cafe today. I had some really good coffee." This voice data is recorded on the device.

[1450] Step 2:

[1451] Converting audio data into text data

[1452] The device sends the recorded voice data to the server. The server uses a voice recognition engine (for example, Google's voice recognition service) to convert the voice data into text data. In this process, the voice data (input) is converted into text information (output). The converted text data, "I went to a new cafe today. I had some very delicious coffee," is sent back and displayed on the device's screen.

[1453] Step 3:

[1454] Check and correct text data

[1455] The user checks the text data displayed on the terminal and makes corrections as necessary. For example, "I drank a very delicious coffee" is changed to "I drank a very delicious espresso." In this procedure, the user looks at the text data (input), makes the necessary corrections (data processing), and obtains the corrected text data (output).

[1456] Step 4:

[1457] Storing text data in a database

[1458] The text data confirmed and corrected by the user is sent back to the server from the device. The server tags the text data and extracts keywords (for example, adding tags such as "cafe" or "coffee"), and stores it in a database. In this process, the text data (input) is organized and stored as tagged text data (output).

[1459] Step 5:

[1460] Generate a blog post draft

[1461] At a specified time (for example, a time set by the user each night), the server uses a generative AI model (for example, GPT-3) to generate a draft blog post based on the tagged text data in the database. In this process, the stored text data (input) is passed through a natural language generation algorithm to generate a draft blog post (output).

[1462] Step 6:

[1463] View, check and edit the generated manuscript

[1464] The generated draft of the blog post is sent to the user's device and displayed on the screen. The user can review this draft and make any necessary corrections. For example, the draft of a blog post might read, "I went to a new cafe today. I enjoyed a delicious espresso." In this step, the generated draft (input) is reviewed and edited by the user (data processing) before becoming the final draft (output).

[1465] Step 7:

[1466] Post your blog post to your preferred blogging platform

[1467] After the user has reviewed and edited the blog post, it is sent from the device to the server, which then calls the API of the designated blog platform to post it. For example, an article such as "I went to a new cafe today. I enjoyed a delicious espresso" is posted to the blog. The success or failure of the post is again reported to the device and communicated to the user. Here, the final draft (input) is posted to the blog platform (data calculation and manipulation), and the posting result (output) is returned to the user.

[1468] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1469] This system allows users to talk about everyday events, stores the content in a database, analyzes it, and generates a draft for posting to an SNS. It also recognizes the user's emotions and reflects them in the generated text. This system automates the entire process from user input to posting to an SNS, not only reducing the burden on the user but also enabling posts that take emotions into consideration.

[1470] System configuration and operation

[1471] 1. User voice input

[1472] The user talks to the device about everyday events. For example, "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner."

[1473] 2. Audio data conversion

[1474] The device records the user's voice and sends the voice data to the server. The server uses a speech recognition engine to convert the voice data into text data. The converted text data is sent back to the device and displayed on the screen. The user can check the text and make corrections if necessary.

[1475] 3. Emotional Recognition

[1476] The server sends the voice data to the emotion engine, which then recognizes the emotion from the user's voice. The emotion engine then analyzes the user's emotional state based on the tone of voice and vocabulary selection, and adds the results to the text data.

[1477] 4. Saving to the database

[1478] The device sends the text data it has confirmed to a server. The server then adds a date and category (e.g., work, eating out) to the text data and stores it in a database. The stored data is then organized through natural language processing, including tagging and keyword extraction. Emotional information added through emotion recognition is also stored.

[1479] 5. Creating a manuscript for posting on social media

[1480] When the posting timing specified by the user (e.g., daily, when a specific event occurs) arrives, the server uses generative AI to generate a script for posting to social media based on the text data in the database. The generated script takes emotional information into account. For example, if positive emotions are detected, positive expressions will be emphasized.

[1481] 6. User confirmation and editing

[1482] The device displays the generated script for posting to social media to the user and provides an interface for reviewing and editing as necessary. For example, it could be displayed in the form of "Yesterday, my big presentation was a success! Afterwards, I enjoyed a delicious Italian meal with friends. It was a fulfilling day."

[1483] 7. Posting to social media

[1484] Once the user has finished checking and editing, they confirm the post. The device then receives this and sends the confirmed manuscript to the server. The server then calls the API of each SNS platform and executes the post. The success or failure of the post is again notified to the device and conveyed to the user.

[1485] Specific examples

[1486] For example, consider the case where a user says, "Today I went on a picnic with my family and it was so much fun. Then I saw a new movie."

[1487] 1. The user speaks.

[1488] 2. The device records the audio and sends it to the server.

[1489] 3. The server uses a speech recognition engine to convert this into text data: "Today I went on a picnic with my family and had a great time. Then I saw a new movie."

[1490] 4. The server uses an emotion engine to recognize positive emotions from the voice and tag the text data as "joy."

[1491] 5. Display the text data on the terminal.

[1492] 6. The user checks, corrects, and approves.

[1493] 7. The device sends the text data to the server, where it is tagged, keywords are extracted, and the data is stored in a database.

[1494] 8. Based on the posting timing, the server generates a script for posting to social media: "Yesterday I went on a picnic with my family. It was so much fun. Afterwards, I also saw a new movie. It was a great day."

[1495] 9. The terminal displays the generated manuscript, which the user can check and edit.

[1496] 10. The user submits the finalized manuscript.

[1497] 11. The server posts to the SNS platform and notifies the device of the results.

[1498] This system allows users to easily record and post everyday events, and also generates engaging content that reflects their emotions.

[1499] The processing flow will be explained below.

[1500] Step 1:

[1501] The user talks to the device about everyday events. For example, "I had a big presentation at work today and it went well. I went out to eat Italian food with a friend for dinner."

[1502] Step 2:

[1503] The device will record the user's voice and the recorded voice data will be temporarily stored on the device.

[1504] Step 3:

[1505] The device sends the recorded audio data to the server, where it is converted into an appropriate format and sent.

[1506] Step 4:

[1507] The server receives the voice data, which is then input into a voice recognition engine and converted into text data.

[1508] Step 5:

[1509] The server inputs the converted text data into the emotion engine, which analyzes the user's emotions based on the tone and content of the voice and adds emotional information.

[1510] Step 6:

[1511] The server sends text data and emotional information back to the device. The text data, such as "I had a big presentation at work today, and it went well. For dinner, I went out to eat Italian food with a friend," is tagged with emotional information.

[1512] Step 7:

[1513] The device displays the text data and emotional information to the user, who can then review and correct it if necessary.

[1514] Step 8:

[1515] The user checks and corrects the text data and presses the "Confirm" or "Send" button.

[1516] Step 9:

[1517] The device sends the confirmed and corrected text data to the server, which adds date and category information to the text data and stores it in a database.

[1518] Step 10:

[1519] The server then performs tagging and keyword extraction on the data stored in the database, including emotional information.

[1520] Step 11:

[1521] When the designated posting time (e.g., every day or when a specific event occurs) arrives, the server generates a script for posting to social media based on the data in the database. Emotional information is also taken into account, with positive emotions emphasized.

[1522] Step 12:

[1523] The server sends the generated script to the terminal, generating a script in the format "Yesterday's presentation was a success! Afterwards, I enjoyed some delicious Italian food with my friends. It was a fulfilling day."

[1524] Step 13:

[1525] The terminal displays the generated manuscript to the user and provides an interface that prompts them to check and edit it.

[1526] Step 14:

[1527] The user checks the manuscript, edits it if necessary, and presses the "Submit" button.

[1528] Step 15:

[1529] The device sends the finalized manuscript to the server, which then calls the API of each SNS platform and posts the generated manuscript.

[1530] Step 16:

[1531] The server checks whether the posting was successful or not and returns the result to the device.

[1532] Step 17:

[1533] The device notifies the user of the posting result (e.g., "Posting successful!").

[1534] This series of processes allows users to easily record everyday events and post engaging content that reflects their emotions on social media.

[1535] Example 2

[1536] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1537] Conventional SNS posting systems have the problem that users must manually input daily events and create posts that reflect their emotions, which requires a great deal of time and effort. Furthermore, the technical means for reading and appropriately reflecting emotions are insufficient, making it difficult to properly express the user's intended message. Furthermore, the ability to adjust the timing of posts is limited, reducing user convenience. The purpose of this invention is to solve these problems and provide a system that automatically posts a variety of emotionally reflective posts while reducing the burden on users.

[1538] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1539] In this invention, the server includes means for inputting daily events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for saving the text data in a database, means for organizing the saved text data by adding date and category information, means for recognizing emotions from the voice data and adding the results to the text data, means for analyzing the text data in the database and generating a manuscript for posting to an SNS using generation technology, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, means for posting the manuscript confirmed and edited by the user to a specified SNS platform, and means for notifying the user of the success or failure of the posting. This enables users to easily record daily events and automatically generate and post sentences that reflect their emotions.

[1540] "Voice data" refers to data that records what a user says to a terminal as a voice signal.

[1541] "Text data" refers to textual information data obtained by analyzing and converting voice data.

[1542] A "terminal" is a device that a user uses to input voice, and that has the function of recording and displaying voice.

[1543] A "server" is a device or system that processes voice data sent from a terminal, converts it into text data, stores it in a database, analyzes it, and so on.

[1544] A "database" is a data management system for organizing and storing text data and associated information (e.g., dates, categories, and emotion tags).

[1545] "Emotion recognition" is a technology that analyzes and extracts a user's emotional state from voice and text data.

[1546] "Generation technology" refers to technology for generating specific formats and content based on information in a database, and is primarily used to generate manuscripts for posting on social media.

[1547] "SNS Platform" means a system that provides social networking services that enable users to communicate with other users over the Internet.

[1548] "Posting" refers to the act of sending the generated manuscript for posting to an SNS platform on the Internet and making it public.

[1549] "Notification" refers to the act of the server sending specific information (e.g., posting success or failure) to the terminal to inform the user.

[1550] This system allows users to talk about everyday events into a device, accumulates the content in a database, analyzes it, and generates a manuscript for posting to SNS. It also recognizes the user's emotions and reflects them in the generated text. This system automates the entire process from user input to posting to SNS, reducing the burden on the user and enabling posts that take emotions into consideration.

[1551] First, the user talks about everyday events into a dedicated device. The device uses a built-in microphone to record the user's voice and save it as audio data. The recorded audio data is then sent to a server via the Internet.

[1552] The server uses the Google Cloud Speech-to-Text API to convert the received voice data into text data, which is then sent back to the device via the Internet and displayed for the user to review. The user can then review the text data and make corrections as necessary.

[1553] Next, the server sends the voice data to the Microsoft Azure Cognitive Services emotion recognition API for emotion recognition. The emotion recognition API analyzes the voice tone and vocabulary choice to extract the user's emotional state. The analysis results are added to the text data as emotion tags. The text data is then stored in a database. When storing the data, the server uses a natural language processing library (e.g., NLTK) to add date and category information to the text data and organize it.

[1554] When the user's designated posting time arrives, the server retrieves the target text data from the database and generates a script for posting to social media using a generative AI model such as OpenAI's GPT-4. At this time, the server inputs the following prompt to the generative AI model:

[1555] "Generate a social media post based on the following text and sentiment tags. Emphasize positive sentiment and create content that will interest your readers."

[1556] For example, if the server generates a script based on the text data "I had a big presentation at work today, which was a great success. For dinner, I went out to eat Italian food with friends," and the emotion tag "joy," the resulting script would read, "Yesterday's big presentation was a success! Afterwards, I enjoyed some delicious Italian food with friends. It was a fulfilling day."

[1557] The generated manuscript is sent to the terminal and displayed to the user. The user checks the manuscript and makes edits as necessary. When the user has finally finished checking and editing, the terminal sends the manuscript to the server.

[1558] The server calls the API of various social media platforms (e.g., Twitter API, Facebook Graph API) and posts the finalized manuscript on the Internet. The success or failure of this posting is again notified to the device and conveyed to the user.

[1559] The above process allows users to easily record their daily events and effectively post them on social media, reflecting their emotions. This system significantly reduces the burden on users and is an effective way to provide engaging content that takes emotions into consideration.

[1560] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1561] Step 1:

[1562] Users talk to the device about everyday events.

[1563] Specifically, the device's built-in microphone records the user's voice.

[1564] Input: User's voice

[1565] Output: Audio data (recording file)

[1566] Step 2:

[1567] The device sends the recorded audio data to the server.

[1568] Specifically, the audio file is uploaded to a server via the Internet.

[1569] Input: Audio data

[1570] Output: Audio data file on the server

[1571] Step 3:

[1572] The server uses the Google Cloud Speech-to-Text API to convert the audio data into text data.

[1573] Specifically, audio data is sent to the API and the returned text data is obtained.

[1574] Input: Audio data

[1575] Output: Text data (converted character information)

[1576] Step 4:

[1577] The server retransmits the converted text data to the terminal.

[1578] Specifically, text data is downloaded to the terminal via the Internet.

[1579] Input: Text data

[1580] Output: Text data displayed on the terminal

[1581] Step 5:

[1582] The user checks the text data displayed on the terminal and corrects it if necessary.

[1583] Specifically, use a text editor to correct any errors or unnecessary parts.

[1584] Input: Text data

[1585] Output: Text data that the user has confirmed and corrected

[1586] Step 6:

[1587] The terminal transmits the text data corrected by the user to the server again.

[1588] Specifically, the corrected text data is uploaded to a server via the Internet.

[1589] Input: Corrected text data

[1590] Output: Modified text data on the server

[1591] Step 7:

[1592] The server sends the corrected text data to the Microsoft Azure Cognitive Services emotion recognition API for emotional analysis.

[1593] Specifically, text data is sent to the API and the sentiment analysis results are obtained.

[1594] Input: Text data

[1595] Output: Emotion-tagged text data

[1596] Step 8:

[1597] The server adds date and category information to the text data and stores it in a database.

[1598] Specifically, the text data is inserted into the database as a new record.

[1599] Input: emotion-tagged text data

[1600] Output: Text data stored in the database

[1601] Step 9:

[1602] When the specified posting time arrives, the server uses OpenAI's generative AI model to generate a manuscript for posting on social media.

[1603] Specifically, text data is retrieved from a database and prompt sentences are input into a generative AI model to generate a manuscript.

[1604] Input: Text data in the database, prompt statements

[1605] Output: Generated manuscript for posting to social media

[1606] Step 10:

[1607] The generated manuscript is sent to the terminal and displayed to the user.

[1608] Specifically, the generated manuscript is downloaded to the terminal via the Internet.

[1609] Input: Generated SNS posting manuscript

[1610] Output: Manuscript for posting to social media displayed on the device

[1611] Step 11:

[1612] The user reviews the manuscript and edits it if necessary.

[1613] Specifically, the text editor is used again to make corrections and additions to the manuscript.

[1614] Input: Manuscript to post on social media

[1615] Output: User-confirmed and edited manuscript for posting on social media

[1616] Step 12:

[1617] The terminal transmits the final manuscript to the server.

[1618] Specifically, the final manuscript is uploaded to a server via the Internet.

[1619] Input: Confirmed and edited manuscript for posting to social media

[1620] Output: Final manuscript on the server

[1621] Step 13:

[1622] The server calls the APIs of various social media platforms and posts the finalized manuscript.

[1623] Specifically, you submit your manuscript using the posting API of each SNS and retrieve the results.

[1624] Input: Final manuscript

[1625] Output: Posting results to social media

[1626] Step 14:

[1627] The terminal is notified of the success or failure of the posting.

[1628] Specifically, a notification message is sent to the terminal and displayed to the user.

[1629] Input: Post results

[1630] Output: Notification message displayed on the terminal

[1631] (Application example 2)

[1632] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1633] In conventional customer service, it has been difficult to record and analyze customer interactions in real time and use the results as feedback or reviews. In addition, there has been a lack of a way to efficiently generate reviews and social media posts that reflect customer sentiment, so an effective tool is needed to improve customer satisfaction.

[1634] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting everyday events spoken by a user into a terminal as voice data, means for converting the voice data into text data, means for saving the text data in a database, means for analyzing the text data in the database and generating a manuscript for posting on an SNS, means for displaying the generated manuscript on the user's terminal and prompting the user to confirm and edit it, means for posting the manuscript confirmed and edited by the user to a specified SNS platform, and means for recognizing emotions from the voice data and reflecting the recognized emotional information in the generated text. This makes it possible to record the content of conversations with customers in real time and generate feedback and reviews that reflect the emotional information.

[1635] "Terminal" refers to a device that allows users to input voice data and check and edit text data and emotion recognition results.

[1636] "Voice data" refers to information recorded in audio format about everyday events spoken by a user.

[1637] "Text data" refers to information that has been analyzed and converted into text form from audio data.

[1638] The "database" is a data repository for storing generated text data and emotion recognition results.

[1639] "Analysis" refers to the process of understanding and extracting meaning from the text data in a database.

[1640] A "script for posting on social media" is a piece of text created using a generative AI model based on text data, intended for posting on social media.

[1641] "Emotion recognition" is the process of analyzing a user's emotional state from their voice data and obtaining the results.

[1642] A "generative AI model" is an artificial intelligence technology that generates sentences in a specific format based on input data.

[1643] A "prompt sentence" is the initial input sentence that a generative AI model uses when generating a sentence.

[1644] A "wearable device" is an information terminal that is worn by the user.

[1645] The system is designed to enable users to record customer interactions through wearable devices such as smart glasses and automatically generate transcripts for reviews and social media posts. The system consists of the following main components:

[1646] Hardware and Software Configuration

[1647] 1. Device:

[1648] Wearable devices such as smart glasses.

[1649] It is equipped with a microphone and has the ability to receive voice input.

[1650] Record customer interactions in real time.

[1651] 2. Server:

[1652] It runs a speech recognition engine, a sentiment analysis engine, and a generative AI model.

[1653] Convert the voice data into text data and analyze emotions.

[1654] The text data is stored in a database for later analysis and tagging.

[1655] System processing flow

[1656] 1. Voice input:

[1657] The device records the user's voice in real time and sends it to the server as audio data.

[1658] 2. Audio data conversion:

[1659] The server uses a speech recognition engine to convert the voice data into text data, which is then sent back to the device and displayed to the user.

[1660] 3. Emotion Recognition:

[1661] The server recognizes emotions from the recorded voice data and adds emotional information to it using an emotion analysis engine.

[1662] 4. Save to database:

[1663] The text data and sentiment information are stored in a database, organized by date and category (e.g., product reviews, service ratings).

[1664] 5. Creating a manuscript for posting on social media:

[1665] At the specified timing, the server extracts text data from the database and uses the generative AI model to generate a script for posting on social media. This script reflects emotional information. For example, if "positive emotions" are detected, positive expressions will be emphasized.

[1666] 6. User review and editing:

[1667] The manuscript is displayed on the user's terminal, and the user can check it and edit it if necessary.

[1668] 7. Posting to social media:

[1669] The terminal sends the manuscript that the user has finalized to the server, and the server posts it via the API of each SNS platform.

[1670] Specific examples

[1671] For example, in a physical store, imagine a scenario where a salesperson asks a customer, "Today, you tried this new perfume. What do you think?", and the customer replies, "It smells great! I'd like to come again." The system records this conversation and processes it as follows:

[1672] Voice input: The store clerk's smart glasses record the conversation.

[1673] Conversion of voice data: The server converts this into text data such as "You tried this new perfume today. What did you think?" and "It smelled great! I'd like to come again."

[1674] Emotion recognition: Positive emotions are detected.

[1675] Using a generative AI model: Generate a review using the prompt "Review when emotions are positive: The scent was amazing! I'd love to come back. How can I recreate such a wonderful day?"

[1676] User review and editing: Store associates can review and edit generated reviews using smart glasses.

[1677] Posting to social media: After the review is verified and edited, it will be posted to each social media platform.

[1678] Based on these example prompts, the system can quickly generate engaging, emotionally sensitive content, contributing to improved customer satisfaction.

[1679] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1680] Step 1:

[1681] The device records the user's voice in real time and sends the voice data to the server. The input is the user's voice, and the output is the voice data sent to the server. Specifically, the voice input is captured using the microphone in the smart glasses, and the data is compressed and sent to the server.

[1682] Step 2:

[1683] The server uses a speech recognition engine to convert the voice data into text data. The input is the voice data sent in step 1, and the output is the converted text data. Specifically, Google's speech recognition API is used to convert the voice data into Japanese text.

[1684] Step 3:

[1685] The terminal receives the converted text data and displays it to the user. The input is the text data sent from the server, and the output is the text data displayed on the terminal's display. Specifically, the text is displayed on the display of the smart glasses and confirmed by the user.

[1686] Step 4:

[1687] The server recognizes emotions from voice data and adds that emotional information to text data. The input is text data and voice data, and the output is text data with emotional information. Specifically, it uses Hugging Face's emotion analysis engine to analyze emotions from voice tone and keywords, and adds that information to the text data.

[1688] Step 5:

[1689] The server stores text data and emotion information in a database. The input is text data with emotion information, and the output is the data stored in the database. Specifically, data organized by category is stored in the database.

[1690] Step 6:

[1691] At the specified time, the server extracts text data from the database and uses the generative AI model to generate a manuscript for posting on social media. The input is the text data in the database, and the output is the generated manuscript for posting on social media. Specifically, the generative AI model is input with a prompt sentence: "Review when emotions are positive: The scent was amazing! I'd like to come again. How can I recreate such a good day again?" and outputs the generated text.

[1692] Step 7:

[1693] The device receives the generated SNS post manuscript and displays it to the user. The input is the SNS post manuscript sent from the server, and the output is the manuscript displayed on the device's display. Specifically, it is displayed on the smart glasses display for the user to check and edit.

[1694] Step 8:

[1695] The device sends the manuscript edited by the user to be posted to the server, and the server executes the posting via the API of each SNS platform. The input is the edited manuscript to be posted to the SNS, and the output is the result of posting to each SNS platform. Specifically, the SNS API is called to post, and the device is notified of the success or failure of the posting.

[1696] Step 9:

[1697] The device notifies the user of the results of the post to the SNS. The input is the post result sent from the server, and the output is the result notification displayed on the device's display. Specifically, a message such as "Posting successful" is displayed on the smart glasses' display.

[1698] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1699] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1700] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1701] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1702] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1703] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1704] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1705] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1706] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1707] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1708] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1709] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1710] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1711] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1712] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1713] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1714] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1715] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1716] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1717] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1718] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1719] The following is further disclosed regarding the above embodiment.

[1720] (Claim 1)

[1721] A means for inputting daily events spoken by the user into the terminal as voice data;

[1722] means for converting voice data into text data;

[1723] a means for storing the text data in a database;

[1724] A means of analyzing the text data in the database and generating manuscripts for posting on social media;

[1725] A means to display the generated manuscript on the user's terminal and prompt the user to check and edit it;

[1726] A means for users to post the manuscripts they have reviewed and edited to a designated SNS platform;

[1727] A system including:

[1728] (Claim 2)

[1729] 10. The system of claim 1, further comprising means for tagging and extracting keywords from daily events entered by a user.

[1730] (Claim 3)

[1731] 2. The system according to claim 1, further comprising means for setting a posting schedule for the generated manuscript for posting to SNS at a timing specified by the user.

[1732] "Example 1"

[1733] (Claim 1)

[1734] A means for inputting daily events spoken by the user into the terminal as voice data;

[1735] means for converting voice data into text data;

[1736] a means for storing the text data in a database;

[1737] A means of analyzing the text data in the database and generating manuscripts for posting on social media;

[1738] A means to display the generated manuscript on the user's terminal and prompt the user to check and edit it;

[1739] A means for users to post the manuscripts they have reviewed and edited to a designated SNS platform;

[1740] a means for generating prompts using a generative AI model to generate posts through an automated process;

[1741] A system including:

[1742] (Claim 2)

[1743] 10. The system of claim 1, further comprising means for tagging and extracting keywords from daily events entered by a user.

[1744] (Claim 3)

[1745] 2. The system according to claim 1, further comprising means for setting a posting schedule for the generated manuscript for posting to SNS at a timing specified by the user.

[1746] "Application Example 1"

[1747] (Claim 1)

[1748] A means for inputting daily events spoken by the user into the terminal as voice data;

[1749] means for converting voice data into text data;

[1750] a means for storing the text data in a database;

[1751] A means for analyzing the text data in the database and generating a blog post manuscript;

[1752] A means to display the generated manuscript on the user's terminal and prompt the user to check and edit it;

[1753] A means for users to post the reviewed and edited manuscript to a designated blog platform;

[1754] A system including:

[1755] (Claim 2)

[1756] 10. The system of claim 1, further comprising means for tagging and extracting keywords from daily events entered by a user.

[1757] (Claim 3)

[1758] 2. The system according to claim 1, further comprising means for setting a posting schedule for the generated blog article manuscript at a timing designated by the user.

[1759] "Example 2: Combining Emotion Engines"

[1760] (Claim 1)

[1761] A means for inputting daily events spoken by the user into the terminal as voice data;

[1762] means for converting voice data into text data;

[1763] a means for storing the text data in a database;

[1764] A method for organizing saved text data by adding date and category information,

[1765] A means for recognizing emotions from voice data and adding the results to text data;

[1766] A means for analyzing the text data in the database and generating a manuscript for posting on SNS using generation technology;

[1767] A means to display the generated manuscript on the user's terminal and prompt the user to check and edit it;

[1768] A means for users to post the manuscripts they have reviewed and edited to a designated SNS platform;

[1769] a means of notifying you of the success or failure of your submission;

[1770] A system including:

[1771] (Claim 2)

[1772] 10. The system of claim 1, further comprising means for tagging and extracting keywords from daily events entered by a user.

[1773] (Claim 3)

[1774] 2. The system according to claim 1, further comprising means for setting a posting schedule for the generated manuscript for posting to SNS at a timing specified by the user.

[1775] "Application example 2 when combining emotion engines"

[1776] (Claim 1)

[1777] A means for inputting daily events spoken by the user into the terminal as voice data;

[1778] means for converting voice data into text data;

[1779] a means for storing the text data in a database;

[1780] A means of analyzing the text data in the database and generating manuscripts for posting on social media;

[1781] A means to display the generated manuscript on the user's terminal and prompt the user to check and edit it;

[1782] A means for users to post the manuscripts they have reviewed and edited to a designated SNS platform;

[1783] A means for recognizing emotions from voice data and reflecting the recognized emotional information in text generation;

[1784] A system including:

[1785] (Claim 2)

[1786] 10. The system of claim 1, further comprising means for tagging and extracting keywords from daily events entered by a user.

[1787] (Claim 3)

[1788] 2. The system according to claim 1, further comprising means for setting a posting schedule for the generated manuscript for posting to SNS at a timing specified by the user.

[1789] (Claim 4)

[1790] A means for recording voice data, performing emotion analysis in real time, and generating emotion-based content;

[1791] The system according to claim 1, further comprising means for displaying the generated text data on an in-store terminal or a wearable device and prompting confirmation and editing. [Explanation of symbols]

[1792] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for inputting daily events spoken by the user into the terminal as voice data; means for converting voice data into text data; a means for storing the text data in a database; A means of analyzing the text data in the database and generating manuscripts for posting on social media; A method for displaying the generated manuscript on the user's device and prompting the user to check and edit it; A means for users to post the manuscript they have reviewed and edited on a designated SNS platform, A system including:

2. The system according to claim 1, further comprising means for tagging and extracting keywords from daily events input by the user.

3. The system according to claim 1 , further comprising a means for setting a posting schedule for the generated manuscript for posting to an SNS at a timing designated by the user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A