System
The system addresses the challenge of creating high-quality presentation materials by integrating voice recognition, image input, and automatic generation, facilitating efficient and user-friendly material creation.
Patent Information
- Application Number
- JP2024123926
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Creating high-quality presentation materials is time-consuming and requires specialized knowledge, with few systems effectively converting voice input into text and automatically generating materials, lacking usability and accuracy.
A system equipped with a voice recognition unit, image input unit, generation unit, editing unit, and output unit that converts voice data to text, imports image data, automatically generates presentation materials, allows fine-tuning, and saves them in a specified format.
Enables efficient creation of high-quality presentation materials by integrating voice recognition, image input, and automatic generation, reducing time and effort while supporting user edits.
Smart Images

Figure 2026022409000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Creating presentation materials requires a lot of time and effort, and creating high-quality materials in particular requires specialized know-how. Furthermore, there are few opportunities to effectively evaluate and improve the quality of materials, making it difficult to improve their quality. Furthermore, there are few systems that convert voice input into text and automatically generate materials based on that text, and even those that do exist lack usability and accuracy. To address these issues, we provide a system that utilizes voice recognition and generative AI to semi-automatically create high-quality presentation materials quickly. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems with a system equipped with a voice recognition unit, an image input unit, a generation unit, an editing unit, and an output unit. First, the voice recognition unit converts voice data input by the user into text data. Next, the image input unit imports image data specified by the user. The generation unit automatically creates presentation materials based on this text data and image data. Furthermore, the user can fine-tune the created presentation materials using the editing unit. Finally, the output unit saves or outputs the materials in a specified format. In this way, the process of creating presentation materials is made more efficient and the creation of high-quality materials is supported.
[0006] The "voice recognition means" is a device or software that has the function of acquiring voice data uttered by a user and converting the voice data into text data.
[0007] "Image input means" refers to a device or software that has the function of reading image data designated by the user and storing it in a form that can be processed within the system.
[0008] The "generation means" refers to a device or software that has the function of automatically creating presentation materials based on input text data and image data.
[0009] The "editing means" is a device or software that has the function of allowing the user to make minor corrections to the generated presentation materials.
[0010] "Output means" refers to a device or software that has the function of saving the created or edited presentation materials in a predetermined format or outputting them to an external device. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0012] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0013] First, the terms used in the following description will be explained.
[0014] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0015] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0016] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0017] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0019] [First embodiment]
[0020] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0021] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0022] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0023] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0024] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0026] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0027] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0028] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0029] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0030] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0031] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0032] As a specific embodiment of the present invention, a system including a voice recognition means, an image input means, a generation means, an editing means, and an output means will be described.
[0033] First, this system receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. The user uses a microphone to input the content of their presentation by voice, and the server converts this voice into text data.
[0034] Next, the user uses the terminal to upload the images and graphs they want to use to the server. The image input means reads the uploaded image data and stores it in a format that can be processed within the system. At this time, the user can upload multiple images and graphs at once.
[0035] The server then uses a generation means to combine the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. The generation means uses an algorithm to optimize the arrangement of text and images and the composition of slides.
[0036] The generated presentation materials can then be edited by the user through the editing tool. The user can check the contents of the slides via their terminal and modify the text or rearrange images as necessary. This editing tool provides an intuitive interface and is designed to allow users to make edits easily.
[0037] Finally, after the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means, which has the function of saving the materials in a downloadable format or sending them to a specified email address or cloud storage.
[0038] As a concrete example, consider the following scenario. A user wishes to "create a presentation that explains market growth," and says "This slide will explain market growth" via speech recognition means. The user then uploads an image file called "market growth.png." The server receives this data and uses generation means to automatically generate a title slide and content slides, placing the images in the appropriate positions. The user then checks the generated slides, makes minor edits to the content, and finally saves them in .pptx format via output means.
[0039] In this way, the present invention allows users to create high-quality presentation materials in a short amount of time.
[0040] The processing flow will be explained below.
[0041] Step 1:
[0042] The user starts voice input. When the user speaks into the microphone, the device captures the voice data through the microphone.
[0043] Step 2:
[0044] The terminal transmits the captured voice data to the server, which converts the voice data into text data via a voice recognition means.
[0045] Step 3:
[0046] The server displays the converted text data to the user for confirmation, and the user checks the converted text data and corrects it if necessary.
[0047] Step 4:
[0048] The user uses the device to select the images and graphs they want to use and upload them to the server, where the device sends the image files to the server.
[0049] Step 5:
[0050] The server reads the uploaded image data using the image input means and stores it in a format that can be processed within the system.
[0051] Step 6:
[0052] The server uses a generating means to automatically generate presentation materials based on the text data and image data, and the generating means optimizes the layout of the slides, the arrangement of the text, and the arrangement of the images.
[0053] Step 7:
[0054] The server delivers the generated presentation materials to the user, who then checks the generated slides through their terminal.
[0055] Step 8:
[0056] The user uses the editing means to make necessary minor corrections to the generated presentation material, such as correcting text or rearranging images.
[0057] Step 9:
[0058] Once the editing is complete, the user notifies the server, which then uses the output means to save the final presentation in a predetermined format (e.g., .pptx format).
[0059] Step 10:
[0060] The server provides the saved materials to the user's device in a downloadable format or sends them to a specified email address or cloud storage.
[0061] Example 1
[0062] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0063] Conventional presentation creation systems require users to manually input text, arrange images, and consider the slide layout, which is time-consuming and labor-intensive. Furthermore, the lack of integration of voice recognition and automatic generation functions makes it difficult to efficiently create high-quality presentation materials.
[0064] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0065] In this invention, the server includes a speech recognition means, an image processing means, a text generation means, a presentation generation means, an editing interface means, and an output saving means, which enables the conversion of speech input into text data, the processing of image data, and the automatic generation of presentation materials by combining text and images, which can then be easily edited and saved.
[0066] "Speech recognition means" is a system that has the function of analyzing voice data and converting it into text data.
[0067] The "image processing means" is a system that receives image data uploaded by users, converts it into a processable format, and saves it.
[0068] The "text generation means" is a system that has the function of using text data generated from voice data analyzed by the voice recognition means.
[0069] A "presentation generation means" is a system that combines text data and image data to automatically generate slides and presentation materials.
[0070] The "editing interface means" is a system that has the function of allowing users to easily modify and edit the generated presentation materials.
[0071] The "output saving means" is a system that has the function of saving the final presentation materials in a specified format and generating a link that allows users to download them.
[0072] As a specific embodiment of the present invention, a system including a speech recognition means, an image processing means, a text generation means, a presentation generation means, an editing interface means, and an output storage means will be described.
[0073] This system receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. The user uses a microphone to input the content of their presentation by voice, and the server uses voice recognition software (e.g., a voice recognition API) to convert this voice into text data.
[0074] Next, the user uses the terminal to upload the images and graphs they want to use to the server. The image processing means reads the uploaded image data, converts it into a format that can be processed within the system, and stores it. At this time, the user can upload multiple images and graphs at once. Specifically, the server uses image processing software (e.g., OpenCV) to perform the processing.
[0075] The server then uses a presentation generation means to combine the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. This presentation generation means includes algorithms for determining the optimal layout of text and images and the composition of slides. Specifically, the server uses presentation generation software (e.g., Python-PPTx).
[0076] The generated presentation materials can then be fine-tuned by the user through an editing interface. The user can check the contents of the slides via their terminal and modify text or rearrange images as necessary. This editing interface is designed to provide an intuitive interface that allows users to easily make modifications.
[0077] Finally, after the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output saving means, which has the function of saving the materials in a downloadable format or sending them to a specified email address or cloud storage.
[0078] As a concrete example, consider the following scenario. A user wishes to "create a presentation that explains market growth," and says via speech recognition means, "This slide will explain market growth." The user then uploads an image file called "market growth.png." The server receives this data and uses the presentation generation means to automatically generate a title slide and content slides, placing the images in the appropriate positions. The user then checks the generated slides, makes minor edits to the content, and finally saves them in .pptx format via the output saving means.
[0079] Example prompt sentence:
[0080] "Create a presentation that explains market growth. This slide explains market growth."
[0081] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0082] Step 1:
[0083] The user speaks into the microphone to input the content of the presentation. The server receives the voice data through a voice recognition means. This voice data is the input, and the server sends this voice data to voice recognition software (e.g., a voice recognition API) to obtain text data. The obtained text data is the output.
[0084] Specific behavior:
[0085] The user speaks into the microphone, "This slide explains market growth."
[0086] The server captures voice data in real time and sends it to the voice recognition API for analysis.
[0087] The speech recognition API converts the speech data into text data and returns the text data to the server.
[0088] Step 2:
[0089] A user uses a terminal to upload images or graphs they want to use to the server. This image data is the input, and the server receives this data using image processing means. The image processing means uses image processing software such as OpenCV to convert the data into a processable format and save it within the system. This converted image data is the output.
[0090] Specific behavior:
[0091] The user logs into the device's web interface, selects the image file "Market Growth.png," and clicks the upload button.
[0092] The server processes the received image data using OpenCV and stores it in its internal storage.
[0093] Step 3:
[0094] The server receives the text data obtained from the speech recognition means and the image data obtained from the image processing means, and automatically generates presentation materials using the presentation generation means. At this stage, the input is text data and image data, and the output is the presentation materials. The presentation generation means uses software such as Python-PPTx.
[0095] Specific behavior:
[0096] The server obtains the text data "This slide explains market growth" and the image data "market growth.png".
[0097] Use the Python-PPTx library to create a new presentation file and place text and images on slides.
[0098] The text will be displayed in the title area and the image will be automatically placed in the content area.
[0099] Step 4:
[0100] The user uses the terminal to check the generated slides and make minor corrections through the editing interface means. The input of this step is the generated presentation material, and the output is the corrected presentation material.
[0101] Specific behavior:
[0102] The user previews the generated slides in a web interface.
[0103] Users can click on the text boxes to modify their content or drag the image to change its position.
[0104] Once the user has completed the edits, they click the "Save" button and the edits are sent to the server.
[0105] Step 5:
[0106] The user signals the server that they are finished editing. The server uses the output storage facility to save the final presentation in the specified format (e.g., .pptx) and generate a download link. The input to this step is the modified presentation, and the output is the saved presentation and a download link.
[0107] Specific behavior:
[0108] The user clicks the "Done" button to notify the server that editing is complete.
[0109] The server uses Python-PPTx to export the final presentation in .pptx format.
[0110] The server saves the generated file in the user's download folder, generates a download link, and notifies the user.
[0111] (Application example 1)
[0112] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0113] Current virtual stores lack systems that can quickly and effectively provide product explanations and comparison presentations. This makes it difficult for users to easily obtain detailed information about products, which can hinder purchasing behavior. Furthermore, traditional methods require a lot of time and effort to create presentation materials, making them unsuitable for use in virtual stores, where immediacy is required.
[0114] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0115] In this invention, the server includes a voice recognition means, an image input means, a generation means, an editing means, an output means, a visual output means for automatically generating product descriptions and comparative presentations in the virtual store, and an integration means for creating a presentation based on the voice input and image data, thereby enabling a user to quickly automatically generate product descriptions and comparative presentations in the virtual store based on the voice input and image data.
[0116] A "voice recognition means" is a device or system that has the function of converting voice data into text data.
[0117] "Image input means" refers to a device or system that has the function of receiving and processing images and graphs uploaded by users.
[0118] The "generation means" is a device or system that has an algorithm that combines text data obtained from the voice recognition means and image data obtained from the image input means to automatically create presentation materials.
[0119] The "editing means" is a device or system that has the function of providing an interface that allows the user to make minor edits to the generated presentation materials.
[0120] "Output means" refers to a device or system that has the function of saving edited presentation materials in a predetermined format, making them available for download, or transmitting them to a specified location.
[0121] The "visual output means for automatically generating product explanations and comparative presentations within a virtual store" is a device or system that has the function of visually providing product explanations and comparative presentations to users within a virtual store.
[0122] The "integration means for creating a presentation based on voice input and image data" is a device or system having an algorithm for automatically generating presentation materials by integrating voice input and image data.
[0123] As a specific embodiment of the present invention, a system for automatically generating product explanations and comparison presentations in a virtual store, which is configured by the following steps, will be described.
[0124] 1. Voice recognition means:
[0125] The server receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. For example, when a user uses a microphone to explain the features and advantages of a product, the voice data is sent to the server and converted into text data.
[0126] 2. Image input method:
[0127] Users use their terminals to upload images and graphs related to products to the server. The image input means reads the uploaded image data and stores it in a format that can be processed within the system. Users can upload multiple images and graphs at once.
[0128] 3. Generation means:
[0129] The server combines the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. This generation means uses a generative AI model to optimize the arrangement of text and images and the composition of slides.
[0130] 4. Editing methods:
[0131] The generated presentation materials can then be fine-tuned by the user. The user can check the contents of the slides via their terminal and modify the text or rearrange images as necessary. This editing tool provides an intuitive interface and is designed to allow users to make edits easily.
[0132] 5. Output Method:
[0133] After the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means that can save the materials in a downloadable format or send them to a specified email address or cloud storage.
[0134] Hardware / Software used
[0135] Hardware:
[0136] Microphone (e.g. USB microphone)
[0137] Computer or smartphone
[0138] software:
[0139] Python 3.x
[0140] speech_recognition module
[0141] python-pptx module
[0142] Google Speech Recognition API
[0143] Specific examples
[0144] For example, imagine a user is trying to explain a new electronics product in a virtual store. The user might say into the microphone, "This product has the latest technology and twice the battery life of previous models."
[0145] Next, the user uploads product images and graphs of specifications via their device. The server receives this data and uses a generative AI model to automatically generate a presentation in the form of the following prompt:
[0146] Example prompt sentence:
[0147] "I'm going to create a presentation in a virtual store to explain a new electronics product. Here's what I'm saying: This product is equipped with the latest technology and has twice the battery life of previous models. I've uploaded some voice-recognized text and related product images. Please generate the presentation materials based on this."
[0148] In this way, the present invention enables users to create high-quality presentation materials in a short amount of time.
[0149] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0150] Step 1:
[0151] The user uses a microphone to input a description of the product by voice. The voice input is sent to the server as voice data via the microphone. The server converts this voice data into text data using a voice recognition means. Specifically, the server analyzes the voice data using the Google voice recognition API and generates corresponding text data.
[0152] Input: Audio data of product description
[0153] Data processing: Converting voice data into text data
[0154] Output: Text data
[0155] Step 2:
[0156] Users use their terminals to upload images and graphs related to products. The server receives these image data using image input means and stores them in the system in a processable format. Users can upload multiple images and graphs at once.
[0157] Input: Image file (e.g. product image, specification graph)
[0158] Data processing: Converting image data into a format that can be processed within the system
[0159] Output: Image data in a processable format
[0160] Step 3:
[0161] The server uses the generation means to automatically generate presentation materials by combining the text data obtained from the voice recognition means and the image data obtained from the image input means, and uses the generative AI model to determine the optimal arrangement of text and images and the composition of slides to create the presentation materials.
[0162] Input: Text data, image data
[0163] Data processing: Automatically generate presentation materials by combining text and images
[0164] Output: Presentation materials
[0165] Step 4:
[0166] The user can check the content of the generated presentation materials through the terminal and make any necessary corrections using the editing tool, which provides an intuitive interface and is designed to allow the user to easily modify text and rearrange images.
[0167] Input: Generated presentation materials
[0168] Data processing: Minor edits to presentation content (text edits, image rearrangement)
[0169] Output: Revised presentation
[0170] Step 5:
[0171] After the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means that can save the materials in a downloadable format or send them to a specified email address or cloud storage.
[0172] Input: Revised presentation
[0173] Data processing: Save and output presentation materials in a specified format
[0174] Output: Final presentation file (.pptx format)
[0175] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0176] As a specific embodiment of the present invention, a system including a voice recognition means, an image input means, a generation means, an editing means, an output means, and an emotion engine will be described.
[0177] First, the system receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. The user uses a microphone to input the content of their presentation by voice, and the server converts this voice into text data. At the same time, the emotion engine recognizes the user's emotional state from their voice tone and speaking style.
[0178] Next, the user uses the terminal to upload the images and graphs they want to use to the server. The image input means reads the uploaded image data and stores it in a format that can be processed within the system. At this time, the user can upload multiple images and graphs at once.
[0179] The server then uses a generation means to combine the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. The generation means uses algorithms to optimize the placement of text and images and the composition of slides. Furthermore, an emotion engine adjusts the content and tone of the presentation materials based on the user's emotional state.
[0180] For example, if the user is in a positive emotional state, the generator will suggest recommended expressions and images. As a specific example, if the user requests "I want to create a presentation explaining market growth" and the emotion engine detects the user's excitement, the server will adjust the insertion of emphasized expressions and bright images.
[0181] Conversely, if the user shows a negative emotional state, the server will provide encouraging messages and advice. For example, if the user shows signs of lacking confidence, the server will display a message such as, "This is your area of expertise. Continue your presentation with confidence."
[0182] The generated presentation materials can then be edited by the user through the editing tool. The user can check the generated slides on their device and edit text or rearrange images as needed. This editing tool provides an intuitive interface and is designed to allow users to make edits easily.
[0183] Finally, after the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means, which has the function of saving the materials in a downloadable format or sending them to a specified email address or cloud storage.
[0184] This system enables users to create high-quality, emotionally sensitive presentation materials in a short amount of time, significantly reducing the effort required for creating materials and helping users maximize the effectiveness of their presentations.
[0185] The processing flow will be explained below.
[0186] Step 1:
[0187] The user starts voice input. When the user speaks into the microphone, the device captures the voice data through the microphone.
[0188] Step 2:
[0189] The terminal transmits the captured voice data to the server, which converts the voice data into text data via a voice recognition means.
[0190] Step 3:
[0191] The server uses an emotion engine to analyze the user's emotional state from the voice data, and this analysis determines whether the user has a positive emotion, a negative emotion, or a neutral emotion.
[0192] Step 4:
[0193] The server displays the converted text data to the user for confirmation, and the user checks the converted text data and corrects it if necessary.
[0194] Step 5:
[0195] The user uses the device to select the images and graphs they want to use and upload them to the server, where the device sends the image files to the server.
[0196] Step 6:
[0197] The server reads the uploaded image data using the image input means and stores it in a format that can be processed within the system.
[0198] Step 7:
[0199] The server uses the generation means to automatically generate presentation materials based on the text data and image data. During this generation process, the tone and content of the materials are adjusted based on the analysis results of the emotion engine.
[0200] Step 8:
[0201] If the user is in a positive emotional state, the server adds suggested phrases and upbeat images, such as inserting powerful words and graphs to highlight a slide that explains market growth.
[0202] Step 9:
[0203] If the user is showing a negative emotional state, the server will provide encouraging messages and advice, such as "You've spent a lot of time preparing for your presentation. Go ahead and give it with confidence."
[0204] Step 10:
[0205] The server delivers the generated presentation materials to the user, who then checks the generated slides through their terminal.
[0206] Step 11:
[0207] The user uses the editing means to make necessary minor corrections to the generated presentation material, such as correcting text or rearranging images.
[0208] Step 12:
[0209] Once the editing is complete, the user notifies the server, which then uses the output means to save the final presentation in a predetermined format (e.g., .pptx format).
[0210] Step 13:
[0211] The server provides the saved materials to the user's device in a downloadable format or sends them to a specified email address or cloud storage.
[0212] Example 2
[0213] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0214] Conventional presentation creation systems have the ability to convert audio data into text data and then combine it with image data to automatically generate presentation materials. However, they lack the ability to adjust the content and tone based on the user's emotions, making it difficult to efficiently create presentation materials that take emotions into consideration. This has resulted in the problem of not maximizing the effectiveness of the presentation. Additionally, editing the generated materials is often not intuitive, which increases the user's workload.
[0215] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a voice recognition means, an image input means, a generation means, an emotion analysis means, an editing means, and an output means. This makes it possible to convert voice data to text, import image data, adjust the content of materials based on the emotional state, intuitively edit materials, and save them in a predetermined format.
[0216] The "voice recognition means" is a means having a function of inputting voice data and converting the voice data into text data.
[0217] "Image input means" refers to a means for importing image data and graph data uploaded by users into a format that can be processed within the system.
[0218] The "generation means" is a means for automatically generating presentation materials by combining text data obtained from the voice recognition means and image data obtained from the image input means.
[0219] "Editing means" refers to a means by which users can modify and edit the generated presentation materials through an intuitive interface.
[0220] "Emotion analysis means" is a means for analyzing the user's emotional state from the tone of voice and speaking style, and reflecting that information in the presentation materials.
[0221] The "output means" is a means for saving the final edited presentation materials in a predetermined format and transmitting them to a specified location or in a format that allows the user to download them.
[0222] A "server" is a device or system that receives audio data and image data from a client, processes and analyzes them, and supports the creation and editing of presentation materials.
[0223] The present invention relates to a presentation material creation system that includes a voice recognition unit, an image input unit, a generation unit, an emotion analysis unit, an editing unit, and an output unit.
[0224] First, the system receives voice input from the user via a voice recognition means. The user uses a microphone to input the presentation content by voice. Voice input begins through a dedicated application on the device or a browser screen. This voice data is sent to the server. The server then converts the received voice data into text data in real time using a voice recognition means such as the Google Speech-to-Text API. At this time, an emotion analysis means analyzes the tone and rhythm of the voice to recognize the user's emotional state. This emotional state is then used in the subsequent stage of generating presentation materials.
[0225] Next, the user uses the terminal to upload images and graphs to be used in the presentation to the server. The user selects multiple image files at once from a file selection dialog using the terminal's browser and presses the upload button. The server receives these image data, and the image input means stores them in a format that can be processed within the system.
[0226] The server uses a generation means to combine text data obtained by the voice recognition means with image data obtained by the image input means to automatically generate presentation materials. Libraries such as Apache POI are used in this generation process. The generation means arranges the text data and image data on slides in an optimal layout and adjusts the content and tone of the materials based on the results of the emotion analysis means. If positive emotions are recognized, bright colors and emphasized expressions are used frequently. On the other hand, if negative emotions are recognized, encouraging messages and advice are added.
[0227] For example, if a user enters the following prompt into the system:
[0228] I want to create a presentation that explains the growth of the market. Emotional state: Excitement
[0229] In this case, the server uses sentiment analysis to recognize the user's excitement and generates slides with a positive tone, featuring emphatic expressions and upbeat images.
[0230] The generated presentation materials can be edited by the user through the editing tool. The user can check the generated slides using the terminal and edit the text or rearrange the images as needed. This editing tool provides an intuitive interface and is designed to allow the user to make edits easily.
[0231] Finally, after the user has finished editing, the server uses the output means to save the final presentation in a predetermined format (e.g., .pptx format), which is then made available to the user in a downloadable format or sent to a specified email address or cloud storage.
[0232] This system allows users to create high-quality, emotionally sensitive presentation materials in a short amount of time, significantly reducing the effort required to create materials.
[0233] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0234] Step 1: Accepting voice input
[0235] The user uses the device's microphone to input the presentation content by voice, starting voice input through a dedicated application on the device or through the browser screen.
[0236] Input: User's voice data
[0237] Output: Audio data sent to the server
[0238] Specific behavior:
[0239] The user speaks the content of the presentation into the microphone.
[0240] An application on the device captures the user's voice in real time.
[0241] Send the captured audio data to the server.
[0242] Step 2: Convert audio data to text
[0243] The server converts the received voice data into text data in real time using a voice recognition means.
[0244] Input: Audio data
[0245] Output: Text data
[0246] Specific behavior:
[0247] The server passes the audio data to the Google Speech-to-Text API.
[0248] The API analyzes the audio data and generates corresponding text data.
[0249] The server receives the generated text data.
[0250] Step 3: Recognizing your emotional state
[0251] The server uses an emotion analysis means to analyze the user's emotional state from the text data and voice tone.
[0252] Input: Text data, audio tone
[0253] Output: User's emotional state
[0254] Specific behavior:
[0255] The server passes the text data and voice tone information to the emotion analysis engine.
[0256] The emotion analysis engine analyzes voice characteristics such as intonation and speed to identify the user's emotional state (e.g., positive, negative, excited, etc.).
[0257] The server obtains the user's emotional state as the analysis result.
[0258] Step 4: Upload images and graphs
[0259] The user uses the terminal to upload images and graphs to be used in the presentation to the server.
[0260] Input: Image data, graph data
[0261] Output: Image and graph data sent to the server
[0262] Specific behavior:
[0263] Displays a file selection dialog in the device's browser.
[0264] The user selects the file to upload and presses the upload button.
[0265] The selected file is sent from the terminal to the server.
[0266] Step 5: Join the data and generate a presentation
[0267] The server uses the generating means to combine the text data obtained from the voice recognition means and the image data obtained from the image input means to automatically generate presentation materials.
[0268] Input: Text data, image data, emotional state
[0269] Output: Presentation materials
[0270] Specific behavior:
[0271] The server acquires the text data and the image data.
[0272] A generating means places the text data and image data on the slide.
[0273] Adjust the content and tone of your materials based on your emotional state (use bright colors if you're positive).
[0274] Step 6: Edit your presentation
[0275] The user checks the generated slides via the terminal and makes any necessary corrections using the editing means.
[0276] Input: Presentation materials
[0277] Output: Revised presentation materials
[0278] Specific behavior:
[0279] The presentation materials generated on the terminal are displayed.
[0280] The user modifies text and images through a GUI.
[0281] The changes are updated on the server in real time.
[0282] Step 7: Finalize your presentation
[0283] The server uses the output means to save the final presentation materials in a predetermined format (for example, .pptx format).
[0284] Input: Revised presentation materials
[0285] Output: Presentation material saved in the specified format
[0286] Specific behavior:
[0287] The final presentation materials are generated on the server.
[0288] Export the generated materials in a specified format (such as pptx).
[0289] Generate and provide a URL from which users can download the materials, or send it to a specified email address.
[0290] This process allows users to efficiently create high-quality, emotionally sensitive presentation materials in a short amount of time.
[0291] (Application example 2)
[0292] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0293] Existing customer service systems have difficulty generating appropriate responses to customer questions and requests immediately. Furthermore, they respond without taking into account the emotional state of staff, resulting in inconsistent quality of customer service. Furthermore, it is difficult for staff to gain confidence when dealing with customers, which can lead to a decline in motivation.
[0294] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, an image input means, an emotion analysis means, a generation means, an editing means, and an output means. This allows the staff member to instantly convert voice input into text data through the smart glasses, generate a response based on the emotional state, and display it on the smart glasses' display. This improves the quality of customer service, increases staff motivation, and enables efficient work performance.
[0295] "Speech recognition means" is a means for inputting speech and converting it into text data.
[0296] "Image input means" refers to means for taking in image data and converting it into a format that can be processed within the system.
[0297] The "emotion analysis means" is a means for analyzing the user's emotional state from voice data, image data, and the like.
[0298] The "generation means" is a means for automatically generating presentation materials and responses by combining text data obtained from the speech recognition means with image data and emotional states obtained from the image input means.
[0299] "Editing tools" are tools that allow users to fine-tune and adjust automatically generated materials and responses.
[0300] "Output means" refers to a means for saving the final generated materials and responses in a predetermined format and transmitting them in a downloadable format or to cloud storage as needed.
[0301] As an embodiment of the present invention, a system including a voice recognition means, an image input means, an emotion analysis means, a generation means, an editing means, and an output means will be described.
[0302] First, the voice recognition means has the function of capturing the user's voice through a microphone installed in the smart glasses and converting that voice into text data. For voice recognition, for example, the speech_recognition library is used. This library converts the user's voice into text through Google's voice recognition service.
[0303] Next, the emotion analysis means analyzes the user's emotional state from the voice data. This analysis is performed using the emotion_recognition library. This library analyzes the tone and pitch of the voice data to determine whether the user is friendly or stressed.
[0304] The generation means combines the text data obtained from the speech recognition means with the emotional state obtained from the emotion analysis means to generate an appropriate response or guide message. This generation uses a generative AI model (e.g., GPT-3). The generated response is adjusted based on the user's emotional state.
[0305] For example, if a customer asks, "What's the difference between these products?", the voice is first converted into text data. Next, an emotion analysis tool detects the friendly emotional state of the staff member. Based on this information, the generative AI model generates a response such as, "I'll explain the difference between these products in an easy-to-understand way."
[0306] The editing tool allows users to fine-tune the generated responses and presentation materials. An intuitive user interface (UI) is used for editing, allowing users to easily modify text and rearrange images.
[0307] Finally, the output means saves the generated materials and responses in a predetermined format, displays them on the smart glasses display, or saves them in a downloadable format. This function can also be linked to cloud storage services (e.g., Google Drive or Dropbox).
[0308] As a concrete example, the following prompt sentences may be fed into a generative AI model to generate a response:
[0309] Voice: A customer asks, "What's different about this product?" The staff member's emotional state is friendly. What response would you generate?
[0310] This system will improve the quality of customer service, increase staff motivation, and enable efficient work execution.
[0311] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0312] Step 1:
[0313] The user wears the smart glasses and speaks their questions and requests into the microphone. This voice data is temporarily stored in the smart glasses' internal memory.
[0314] Step 2:
[0315] The voice data is sent to the server, and the server converts the voice data to text data using the speech_recognition library. The voice data is analyzed and output as string data. This step converts voice input to text data.
[0316] Step 3:
[0317] Based on the converted text data, the server performs emotion analysis of the audio data using the emotion_recognition library. It analyzes the tone, pitch, and speed of the audio to detect the user's emotional state. It then outputs the emotional state (e.g., friendly, stressed).
[0318] Step 4:
[0319] The server uses a generative AI model (e.g., GPT-3) to generate an appropriate response based on the input text data and emotional state. In this step, the appropriate response is input as a prompt to the generative AI model, and the generated text response is output.
[0320] Step 5:
[0321] The generated response is displayed for the user to review on the smart glasses display, and the user has an interface to edit the response if necessary. If edits are made, the revised response data is resubmitted to the server and the final version is saved.
[0322] Step 6:
[0323] The server saves the final response data in a specified format and sends it to cloud storage or a specified email address as needed. In this step, data is saved and output.
[0324] Through the above processing steps, the user can instantly generate, display, edit, and save appropriate responses to customer questions and requests.
[0325] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0326] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0327] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0328] [Second embodiment]
[0329] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0330] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0331] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0332] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0333] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0334] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0335] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0336] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0337] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0338] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0339] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0340] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0341] As a specific embodiment of the present invention, a system including a voice recognition means, an image input means, a generation means, an editing means, and an output means will be described.
[0342] First, this system receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. The user uses a microphone to input the content of their presentation by voice, and the server converts this voice into text data.
[0343] Next, the user uses the terminal to upload the images and graphs they want to use to the server. The image input means reads the uploaded image data and stores it in a format that can be processed within the system. At this time, the user can upload multiple images and graphs at once.
[0344] The server then uses a generation means to combine the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. The generation means uses an algorithm to optimize the arrangement of text and images and the composition of slides.
[0345] The generated presentation materials can then be edited by the user through the editing tool. The user can check the contents of the slides via their terminal and modify the text or rearrange images as necessary. This editing tool provides an intuitive interface and is designed to allow users to make edits easily.
[0346] Finally, after the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means, which has the function of saving the materials in a downloadable format or sending them to a specified email address or cloud storage.
[0347] As a concrete example, consider the following scenario. A user wishes to "create a presentation that explains market growth," and says "This slide will explain market growth" via speech recognition means. The user then uploads an image file called "market growth.png." The server receives this data and uses generation means to automatically generate a title slide and content slides, placing the images in the appropriate positions. The user then checks the generated slides, makes minor edits to the content, and finally saves them in .pptx format via output means.
[0348] In this way, the present invention allows users to create high-quality presentation materials in a short amount of time.
[0349] The processing flow will be explained below.
[0350] Step 1:
[0351] The user starts voice input. When the user speaks into the microphone, the device captures the voice data through the microphone.
[0352] Step 2:
[0353] The terminal transmits the captured voice data to the server, which converts the voice data into text data via a voice recognition means.
[0354] Step 3:
[0355] The server displays the converted text data to the user for confirmation, and the user checks the converted text data and corrects it if necessary.
[0356] Step 4:
[0357] The user uses the device to select the images and graphs they want to use and upload them to the server, where the device sends the image files to the server.
[0358] Step 5:
[0359] The server reads the uploaded image data using the image input means and stores it in a format that can be processed within the system.
[0360] Step 6:
[0361] The server uses a generating means to automatically generate presentation materials based on the text data and image data, and the generating means optimizes the layout of the slides, the arrangement of the text, and the arrangement of the images.
[0362] Step 7:
[0363] The server delivers the generated presentation materials to the user, who then checks the generated slides through their terminal.
[0364] Step 8:
[0365] The user uses the editing means to make necessary minor corrections to the generated presentation material, such as correcting text or rearranging images.
[0366] Step 9:
[0367] Once the editing is complete, the user notifies the server, which then uses the output means to save the final presentation in a predetermined format (e.g., .pptx format).
[0368] Step 10:
[0369] The server provides the saved materials to the user's device in a downloadable format or sends them to a specified email address or cloud storage.
[0370] Example 1
[0371] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0372] Conventional presentation creation systems require users to manually input text, arrange images, and consider the slide layout, which is time-consuming and labor-intensive. Furthermore, the lack of integration of voice recognition and automatic generation functions makes it difficult to efficiently create high-quality presentation materials.
[0373] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0374] In this invention, the server includes a speech recognition means, an image processing means, a text generation means, a presentation generation means, an editing interface means, and an output saving means, which enables the conversion of speech input into text data, the processing of image data, and the automatic generation of presentation materials by combining text and images, which can then be easily edited and saved.
[0375] "Speech recognition means" is a system that has the function of analyzing voice data and converting it into text data.
[0376] The "image processing means" is a system that receives image data uploaded by users, converts it into a processable format, and saves it.
[0377] The "text generation means" is a system that has the function of using text data generated from voice data analyzed by the voice recognition means.
[0378] A "presentation generation means" is a system that combines text data and image data to automatically generate slides and presentation materials.
[0379] The "editing interface means" is a system that has the function of allowing users to easily modify and edit the generated presentation materials.
[0380] The "output saving means" is a system that has the function of saving the final presentation materials in a specified format and generating a link that allows users to download them.
[0381] As a specific embodiment of the present invention, a system including a speech recognition means, an image processing means, a text generation means, a presentation generation means, an editing interface means, and an output storage means will be described.
[0382] This system receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. The user uses a microphone to input the content of their presentation by voice, and the server uses voice recognition software (e.g., a voice recognition API) to convert this voice into text data.
[0383] Next, the user uses the terminal to upload the images and graphs they want to use to the server. The image processing means reads the uploaded image data, converts it into a format that can be processed within the system, and stores it. At this time, the user can upload multiple images and graphs at once. Specifically, the server uses image processing software (e.g., OpenCV) to perform the processing.
[0384] The server then uses a presentation generation means to combine the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. This presentation generation means includes algorithms for determining the optimal layout of text and images and the composition of slides. Specifically, the server uses presentation generation software (e.g., Python-PPTx).
[0385] The generated presentation materials can then be fine-tuned by the user through an editing interface. The user can check the contents of the slides via their terminal and modify text or rearrange images as necessary. This editing interface is designed to provide an intuitive interface that allows users to easily make modifications.
[0386] Finally, after the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output saving means, which has the function of saving the materials in a downloadable format or sending them to a specified email address or cloud storage.
[0387] As a concrete example, consider the following scenario. A user wishes to "create a presentation that explains market growth," and says via speech recognition means, "This slide will explain market growth." The user then uploads an image file called "market growth.png." The server receives this data and uses the presentation generation means to automatically generate a title slide and content slides, placing the images in the appropriate positions. The user then checks the generated slides, makes minor edits to the content, and finally saves them in .pptx format via the output saving means.
[0388] Example prompt sentence:
[0389] "Create a presentation that explains market growth. This slide explains market growth."
[0390] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0391] Step 1:
[0392] The user speaks into the microphone to input the content of the presentation. The server receives the voice data through a voice recognition means. This voice data is the input, and the server sends this voice data to voice recognition software (e.g., a voice recognition API) to obtain text data. The obtained text data is the output.
[0393] Specific behavior:
[0394] The user speaks into the microphone, "This slide explains market growth."
[0395] The server captures voice data in real time and sends it to the voice recognition API for analysis.
[0396] The speech recognition API converts the speech data into text data and returns the text data to the server.
[0397] Step 2:
[0398] A user uses a terminal to upload images or graphs they want to use to the server. This image data is the input, and the server receives this data using image processing means. The image processing means uses image processing software such as OpenCV to convert the data into a processable format and save it within the system. This converted image data is the output.
[0399] Specific behavior:
[0400] The user logs into the device's web interface, selects the image file "Market Growth.png," and clicks the upload button.
[0401] The server processes the received image data using OpenCV and stores it in its internal storage.
[0402] Step 3:
[0403] The server receives the text data obtained from the speech recognition means and the image data obtained from the image processing means, and automatically generates presentation materials using the presentation generation means. At this stage, the input is text data and image data, and the output is the presentation materials. The presentation generation means uses software such as Python-PPTx.
[0404] Specific behavior:
[0405] The server obtains the text data "This slide explains market growth" and the image data "market growth.png".
[0406] Use the Python-PPTx library to create a new presentation file and place text and images on slides.
[0407] The text will be displayed in the title area and the image will be automatically placed in the content area.
[0408] Step 4:
[0409] The user uses the terminal to check the generated slides and make minor corrections through the editing interface means. The input of this step is the generated presentation material, and the output is the corrected presentation material.
[0410] Specific behavior:
[0411] The user previews the generated slides in a web interface.
[0412] Users can click on the text boxes to modify their content or drag the image to change its position.
[0413] Once the user has completed the edits, they click the "Save" button and the edits are sent to the server.
[0414] Step 5:
[0415] The user signals the server that they are finished editing. The server uses the output storage facility to save the final presentation in the specified format (e.g., .pptx) and generate a download link. The input to this step is the modified presentation, and the output is the saved presentation and a download link.
[0416] Specific behavior:
[0417] The user clicks the "Done" button to notify the server that editing is complete.
[0418] The server uses Python-PPTx to export the final presentation in .pptx format.
[0419] The server saves the generated file in the user's download folder, generates a download link, and notifies the user.
[0420] (Application example 1)
[0421] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0422] Current virtual stores lack systems that can quickly and effectively provide product explanations and comparison presentations. This makes it difficult for users to easily obtain detailed information about products, which can hinder purchasing behavior. Furthermore, traditional methods require a lot of time and effort to create presentation materials, making them unsuitable for use in virtual stores, where immediacy is required.
[0423] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0424] In this invention, the server includes a voice recognition means, an image input means, a generation means, an editing means, an output means, a visual output means for automatically generating product descriptions and comparative presentations in the virtual store, and an integration means for creating a presentation based on the voice input and image data, thereby enabling a user to quickly automatically generate product descriptions and comparative presentations in the virtual store based on the voice input and image data.
[0425] A "voice recognition means" is a device or system that has the function of converting voice data into text data.
[0426] "Image input means" refers to a device or system that has the function of receiving and processing images and graphs uploaded by users.
[0427] The "generation means" is a device or system that has an algorithm that combines text data obtained from the voice recognition means and image data obtained from the image input means to automatically create presentation materials.
[0428] The "editing means" is a device or system that has the function of providing an interface that allows the user to make minor edits to the generated presentation materials.
[0429] "Output means" refers to a device or system that has the function of saving edited presentation materials in a predetermined format, making them available for download, or transmitting them to a specified location.
[0430] The "visual output means for automatically generating product explanations and comparative presentations within a virtual store" is a device or system that has the function of visually providing product explanations and comparative presentations to users within a virtual store.
[0431] The "integration means for creating a presentation based on voice input and image data" is a device or system having an algorithm for automatically generating presentation materials by integrating voice input and image data.
[0432] As a specific embodiment of the present invention, a system for automatically generating product explanations and comparison presentations in a virtual store, which is configured by the following steps, will be described.
[0433] 1. Voice recognition means:
[0434] The server receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. For example, when a user uses a microphone to explain the features and advantages of a product, the voice data is sent to the server and converted into text data.
[0435] 2. Image input method:
[0436] Users use their terminals to upload images and graphs related to products to the server. The image input means reads the uploaded image data and stores it in a format that can be processed within the system. Users can upload multiple images and graphs at once.
[0437] 3. Generation means:
[0438] The server combines the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. This generation means uses a generative AI model to optimize the arrangement of text and images and the composition of slides.
[0439] 4. Editing methods:
[0440] The generated presentation materials can then be fine-tuned by the user. The user can check the contents of the slides via their terminal and modify the text or rearrange images as necessary. This editing tool provides an intuitive interface and is designed to allow users to make edits easily.
[0441] 5. Output Method:
[0442] After the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means that can save the materials in a downloadable format or send them to a specified email address or cloud storage.
[0443] Hardware / Software used
[0444] Hardware:
[0445] Microphone (e.g. USB microphone)
[0446] Computer or smartphone
[0447] software:
[0448] Python 3.x
[0449] speech_recognition module
[0450] python-pptx module
[0451] Google Speech Recognition API
[0452] Specific examples
[0453] For example, imagine a user is trying to explain a new electronics product in a virtual store. The user might say into the microphone, "This product has the latest technology and twice the battery life of previous models."
[0454] Next, the user uploads product images and graphs of specifications via their device. The server receives this data and uses a generative AI model to automatically generate a presentation in the form of the following prompt:
[0455] Example prompt sentence:
[0456] "I'm going to create a presentation in a virtual store to explain a new electronics product. Here's what I'm saying: This product is equipped with the latest technology and has twice the battery life of previous models. I've uploaded some voice-recognized text and related product images. Please generate the presentation materials based on this."
[0457] In this way, the present invention enables users to create high-quality presentation materials in a short amount of time.
[0458] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0459] Step 1:
[0460] The user uses a microphone to input a description of the product by voice. The voice input is sent to the server as voice data via the microphone. The server converts this voice data into text data using a voice recognition means. Specifically, the server analyzes the voice data using the Google voice recognition API and generates corresponding text data.
[0461] Input: Audio data of product description
[0462] Data processing: Converting voice data into text data
[0463] Output: Text data
[0464] Step 2:
[0465] Users use their terminals to upload images and graphs related to products. The server receives these image data using image input means and stores them in the system in a processable format. Users can upload multiple images and graphs at once.
[0466] Input: Image file (e.g. product image, specification graph)
[0467] Data processing: Converting image data into a format that can be processed within the system
[0468] Output: Image data in a processable format
[0469] Step 3:
[0470] The server uses the generation means to automatically generate presentation materials by combining the text data obtained from the voice recognition means and the image data obtained from the image input means, and uses the generative AI model to determine the optimal arrangement of text and images and the composition of slides to create the presentation materials.
[0471] Input: Text data, image data
[0472] Data processing: Automatically generate presentation materials by combining text and images
[0473] Output: Presentation materials
[0474] Step 4:
[0475] The user can check the content of the generated presentation materials through the terminal and make any necessary corrections using the editing tool, which provides an intuitive interface and is designed to allow the user to easily modify text and rearrange images.
[0476] Input: Generated presentation materials
[0477] Data processing: Minor edits to presentation content (text edits, image rearrangement)
[0478] Output: Revised presentation
[0479] Step 5:
[0480] After the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means that can save the materials in a downloadable format or send them to a specified email address or cloud storage.
[0481] Input: Revised presentation
[0482] Data processing: Save and output presentation materials in a specified format
[0483] Output: Final presentation file (.pptx format)
[0484] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0485] As a specific embodiment of the present invention, a system including a voice recognition means, an image input means, a generation means, an editing means, an output means, and an emotion engine will be described.
[0486] First, the system receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. The user uses a microphone to input the content of their presentation by voice, and the server converts this voice into text data. At the same time, the emotion engine recognizes the user's emotional state from their voice tone and speaking style.
[0487] Next, the user uses the terminal to upload the images and graphs they want to use to the server. The image input means reads the uploaded image data and stores it in a format that can be processed within the system. At this time, the user can upload multiple images and graphs at once.
[0488] The server then uses a generation means to combine the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. The generation means uses algorithms to optimize the placement of text and images and the composition of slides. Furthermore, an emotion engine adjusts the content and tone of the presentation materials based on the user's emotional state.
[0489] For example, if the user is in a positive emotional state, the generator will suggest recommended expressions and images. As a specific example, if the user requests "I want to create a presentation explaining market growth" and the emotion engine detects the user's excitement, the server will adjust the insertion of emphasized expressions and bright images.
[0490] Conversely, if the user shows a negative emotional state, the server will provide encouraging messages and advice. For example, if the user shows signs of lacking confidence, the server will display a message such as, "This is your area of expertise. Continue your presentation with confidence."
[0491] The generated presentation materials can then be edited by the user through the editing tool. The user can check the generated slides on their device and edit text or rearrange images as needed. This editing tool provides an intuitive interface and is designed to allow users to make edits easily.
[0492] Finally, after the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means, which has the function of saving the materials in a downloadable format or sending them to a specified email address or cloud storage.
[0493] This system enables users to create high-quality, emotionally sensitive presentation materials in a short amount of time, significantly reducing the effort required for creating materials and helping users maximize the effectiveness of their presentations.
[0494] The processing flow will be explained below.
[0495] Step 1:
[0496] The user starts voice input. When the user speaks into the microphone, the device captures the voice data through the microphone.
[0497] Step 2:
[0498] The terminal transmits the captured voice data to the server, which converts the voice data into text data via a voice recognition means.
[0499] Step 3:
[0500] The server uses an emotion engine to analyze the user's emotional state from the voice data, and this analysis determines whether the user has a positive emotion, a negative emotion, or a neutral emotion.
[0501] Step 4:
[0502] The server displays the converted text data to the user for confirmation, and the user checks the converted text data and corrects it if necessary.
[0503] Step 5:
[0504] The user uses the device to select the images and graphs they want to use and upload them to the server, where the device sends the image files to the server.
[0505] Step 6:
[0506] The server reads the uploaded image data using the image input means and stores it in a format that can be processed within the system.
[0507] Step 7:
[0508] The server uses the generation means to automatically generate presentation materials based on the text data and image data. During this generation process, the tone and content of the materials are adjusted based on the analysis results of the emotion engine.
[0509] Step 8:
[0510] If the user is in a positive emotional state, the server adds suggested phrases and upbeat images, such as inserting powerful words and graphs to highlight a slide that explains market growth.
[0511] Step 9:
[0512] If the user is showing a negative emotional state, the server will provide encouraging messages and advice, such as "You've spent a lot of time preparing for your presentation. Go ahead and give it with confidence."
[0513] Step 10:
[0514] The server delivers the generated presentation materials to the user, who then checks the generated slides through their terminal.
[0515] Step 11:
[0516] The user uses the editing means to make necessary minor corrections to the generated presentation material, such as correcting text or rearranging images.
[0517] Step 12:
[0518] Once the editing is complete, the user notifies the server, which then uses the output means to save the final presentation in a predetermined format (e.g., .pptx format).
[0519] Step 13:
[0520] The server provides the saved materials to the user's device in a downloadable format or sends them to a specified email address or cloud storage.
[0521] Example 2
[0522] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0523] Conventional presentation creation systems have the ability to convert audio data into text data and then combine it with image data to automatically generate presentation materials. However, they lack the ability to adjust the content and tone based on the user's emotions, making it difficult to efficiently create presentation materials that take emotions into consideration. This has resulted in the problem of not maximizing the effectiveness of the presentation. Additionally, editing the generated materials is often not intuitive, which increases the user's workload.
[0524] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a voice recognition means, an image input means, a generation means, an emotion analysis means, an editing means, and an output means. This makes it possible to convert voice data to text, import image data, adjust the content of materials based on the emotional state, intuitively edit materials, and save them in a predetermined format.
[0525] The "voice recognition means" is a means having a function of inputting voice data and converting the voice data into text data.
[0526] "Image input means" refers to a means for importing image data and graph data uploaded by users into a format that can be processed within the system.
[0527] The "generation means" is a means for automatically generating presentation materials by combining text data obtained from the voice recognition means and image data obtained from the image input means.
[0528] "Editing means" refers to a means by which users can modify and edit the generated presentation materials through an intuitive interface.
[0529] "Emotion analysis means" is a means for analyzing the user's emotional state from the tone of voice and speaking style, and reflecting that information in the presentation materials.
[0530] The "output means" is a means for saving the final edited presentation materials in a predetermined format and transmitting them to a specified location or in a format that allows the user to download them.
[0531] A "server" is a device or system that receives audio data and image data from a client, processes and analyzes them, and supports the creation and editing of presentation materials.
[0532] The present invention relates to a presentation material creation system that includes a voice recognition unit, an image input unit, a generation unit, an emotion analysis unit, an editing unit, and an output unit.
[0533] First, the system receives voice input from the user via a voice recognition means. The user uses a microphone to input the presentation content by voice. Voice input begins through a dedicated application on the device or a browser screen. This voice data is sent to the server. The server then converts the received voice data into text data in real time using a voice recognition means such as the Google Speech-to-Text API. At this time, an emotion analysis means analyzes the tone and rhythm of the voice to recognize the user's emotional state. This emotional state is then used in the subsequent stage of generating presentation materials.
[0534] Next, the user uses the terminal to upload images and graphs to be used in the presentation to the server. The user selects multiple image files at once from a file selection dialog using the terminal's browser and presses the upload button. The server receives these image data, and the image input means stores them in a format that can be processed within the system.
[0535] The server uses a generation means to combine text data obtained by the voice recognition means with image data obtained by the image input means to automatically generate presentation materials. Libraries such as Apache POI are used in this generation process. The generation means arranges the text data and image data on slides in an optimal layout and adjusts the content and tone of the materials based on the results of the emotion analysis means. If positive emotions are recognized, bright colors and emphasized expressions are used frequently. On the other hand, if negative emotions are recognized, encouraging messages and advice are added.
[0536] For example, if a user enters the following prompt into the system:
[0537] I want to create a presentation that explains the growth of the market. Emotional state: Excitement
[0538] In this case, the server uses sentiment analysis to recognize the user's excitement and generates slides with a positive tone, featuring emphatic expressions and upbeat images.
[0539] The generated presentation materials can be edited by the user through the editing tool. The user can check the generated slides using the terminal and edit the text or rearrange the images as needed. This editing tool provides an intuitive interface and is designed to allow the user to make edits easily.
[0540] Finally, after the user has finished editing, the server uses the output means to save the final presentation in a predetermined format (e.g., .pptx format), which is then made available to the user in a downloadable format or sent to a specified email address or cloud storage.
[0541] This system allows users to create high-quality, emotionally sensitive presentation materials in a short amount of time, significantly reducing the effort required to create materials.
[0542] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0543] Step 1: Accepting voice input
[0544] The user uses the device's microphone to input the presentation content by voice, starting voice input through a dedicated application on the device or through the browser screen.
[0545] Input: User's voice data
[0546] Output: Audio data sent to the server
[0547] Specific behavior:
[0548] The user speaks the content of the presentation into the microphone.
[0549] An application on the device captures the user's voice in real time.
[0550] Send the captured audio data to the server.
[0551] Step 2: Convert audio data to text
[0552] The server converts the received voice data into text data in real time using a voice recognition means.
[0553] Input: Audio data
[0554] Output: Text data
[0555] Specific behavior:
[0556] The server passes the audio data to the Google Speech-to-Text API.
[0557] The API analyzes the audio data and generates corresponding text data.
[0558] The server receives the generated text data.
[0559] Step 3: Recognizing your emotional state
[0560] The server uses an emotion analysis means to analyze the user's emotional state from the text data and voice tone.
[0561] Input: Text data, audio tone
[0562] Output: User's emotional state
[0563] Specific behavior:
[0564] The server passes the text data and voice tone information to the emotion analysis engine.
[0565] The emotion analysis engine analyzes voice characteristics such as intonation and speed to identify the user's emotional state (e.g., positive, negative, excited, etc.).
[0566] The server obtains the user's emotional state as the analysis result.
[0567] Step 4: Upload images and graphs
[0568] The user uses the terminal to upload images and graphs to be used in the presentation to the server.
[0569] Input: Image data, graph data
[0570] Output: Image and graph data sent to the server
[0571] Specific behavior:
[0572] Displays a file selection dialog in the device's browser.
[0573] The user selects the file to upload and presses the upload button.
[0574] The selected file is sent from the terminal to the server.
[0575] Step 5: Join the data and generate a presentation
[0576] The server uses the generating means to combine the text data obtained from the voice recognition means and the image data obtained from the image input means to automatically generate presentation materials.
[0577] Input: Text data, image data, emotional state
[0578] Output: Presentation materials
[0579] Specific behavior:
[0580] The server acquires the text data and the image data.
[0581] A generating means places the text data and image data on the slide.
[0582] Adjust the content and tone of your materials based on your emotional state (use bright colors if you're positive).
[0583] Step 6: Edit your presentation
[0584] The user checks the generated slides via the terminal and makes any necessary corrections using the editing means.
[0585] Input: Presentation materials
[0586] Output: Revised presentation materials
[0587] Specific behavior:
[0588] The presentation materials generated on the terminal are displayed.
[0589] The user modifies text and images through a GUI.
[0590] The changes are updated on the server in real time.
[0591] Step 7: Finalize your presentation
[0592] The server uses the output means to save the final presentation materials in a predetermined format (for example, .pptx format).
[0593] Input: Revised presentation materials
[0594] Output: Presentation material saved in the specified format
[0595] Specific behavior:
[0596] The final presentation materials are generated on the server.
[0597] Export the generated materials in a specified format (such as pptx).
[0598] Generate and provide a URL from which users can download the materials, or send it to a specified email address.
[0599] This process allows users to efficiently create high-quality, emotionally sensitive presentation materials in a short amount of time.
[0600] (Application example 2)
[0601] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0602] Existing customer service systems have difficulty generating appropriate responses to customer questions and requests immediately. Furthermore, they respond without taking into account the emotional state of staff, resulting in inconsistent quality of customer service. Furthermore, it is difficult for staff to gain confidence when dealing with customers, which can lead to a decline in motivation.
[0603] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, an image input means, an emotion analysis means, a generation means, an editing means, and an output means. This allows the staff member to instantly convert voice input into text data through the smart glasses, generate a response based on the emotional state, and display it on the smart glasses' display. This improves the quality of customer service, increases staff motivation, and enables efficient work performance.
[0604] "Speech recognition means" is a means for inputting speech and converting it into text data.
[0605] "Image input means" refers to means for taking in image data and converting it into a format that can be processed within the system.
[0606] The "emotion analysis means" is a means for analyzing the user's emotional state from voice data, image data, and the like.
[0607] The "generation means" is a means for automatically generating presentation materials and responses by combining text data obtained from the speech recognition means with image data and emotional states obtained from the image input means.
[0608] "Editing tools" are tools that allow users to fine-tune and adjust automatically generated materials and responses.
[0609] "Output means" refers to a means for saving the final generated materials and responses in a predetermined format and transmitting them in a downloadable format or to cloud storage as needed.
[0610] As an embodiment of the present invention, a system including a voice recognition means, an image input means, an emotion analysis means, a generation means, an editing means, and an output means will be described.
[0611] First, the voice recognition means has the function of capturing the user's voice through a microphone installed in the smart glasses and converting that voice into text data. For voice recognition, for example, the speech_recognition library is used. This library converts the user's voice into text through Google's voice recognition service.
[0612] Next, the emotion analysis means analyzes the user's emotional state from the voice data. This analysis is performed using the emotion_recognition library. This library analyzes the tone and pitch of the voice data to determine whether the user is friendly or stressed.
[0613] The generation means combines the text data obtained from the speech recognition means with the emotional state obtained from the emotion analysis means to generate an appropriate response or guide message. This generation uses a generative AI model (e.g., GPT-3). The generated response is adjusted based on the user's emotional state.
[0614] For example, if a customer asks, "What's the difference between these products?", the voice is first converted into text data. Next, an emotion analysis tool detects the friendly emotional state of the staff member. Based on this information, the generative AI model generates a response such as, "I'll explain the difference between these products in an easy-to-understand way."
[0615] The editing tool allows users to fine-tune the generated responses and presentation materials. An intuitive user interface (UI) is used for editing, allowing users to easily modify text and rearrange images.
[0616] Finally, the output means saves the generated materials and responses in a predetermined format, displays them on the smart glasses display, or saves them in a downloadable format. This function can also be linked to cloud storage services (e.g., Google Drive or Dropbox).
[0617] As a concrete example, the following prompt sentences may be fed into a generative AI model to generate a response:
[0618] Voice: A customer asks, "What's different about this product?" The staff member's emotional state is friendly. What response would you generate?
[0619] This system will improve the quality of customer service, increase staff motivation, and enable efficient work execution.
[0620] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0621] Step 1:
[0622] The user wears the smart glasses and speaks their questions and requests into the microphone. This voice data is temporarily stored in the smart glasses' internal memory.
[0623] Step 2:
[0624] The voice data is sent to the server, and the server converts the voice data to text data using the speech_recognition library. The voice data is analyzed and output as string data. This step converts voice input to text data.
[0625] Step 3:
[0626] Based on the converted text data, the server performs emotion analysis of the audio data using the emotion_recognition library. It analyzes the tone, pitch, and speed of the audio to detect the user's emotional state. It then outputs the emotional state (e.g., friendly, stressed).
[0627] Step 4:
[0628] The server uses a generative AI model (e.g., GPT-3) to generate an appropriate response based on the input text data and emotional state. In this step, the appropriate response is input as a prompt to the generative AI model, and the generated text response is output.
[0629] Step 5:
[0630] The generated response is displayed for the user to review on the smart glasses display, and the user has an interface to edit the response if necessary. If edits are made, the revised response data is resubmitted to the server and the final version is saved.
[0631] Step 6:
[0632] The server saves the final response data in a specified format and sends it to cloud storage or a specified email address as needed. In this step, data is saved and output.
[0633] Through the above processing steps, the user can instantly generate, display, edit, and save appropriate responses to customer questions and requests.
[0634] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0635] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0636] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0637] [Third embodiment]
[0638] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0639] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0640] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0641] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0642] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0643] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0644] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0645] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0646] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0647] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0648] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0649] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0650] As a specific embodiment of the present invention, a system including a voice recognition means, an image input means, a generation means, an editing means, and an output means will be described.
[0651] First, this system receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. The user uses a microphone to input the content of their presentation by voice, and the server converts this voice into text data.
[0652] Next, the user uses the terminal to upload the images and graphs they want to use to the server. The image input means reads the uploaded image data and stores it in a format that can be processed within the system. At this time, the user can upload multiple images and graphs at once.
[0653] The server then uses a generation means to combine the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. This generation means uses an algorithm to optimize the arrangement of text and images and the composition of slides.
[0654] The generated presentation materials can then be edited by the user through the editing tool. The user can check the contents of the slides via their terminal and modify the text or rearrange images as necessary. This editing tool provides an intuitive interface and is designed to allow users to make edits easily.
[0655] Finally, after the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means, which has the function of saving the materials in a downloadable format or sending them to a specified email address or cloud storage.
[0656] As a concrete example, consider the following scenario. A user wishes to "create a presentation that explains market growth," and says "This slide will explain market growth" via speech recognition means. The user then uploads an image file called "market growth.png." The server receives this data and uses generation means to automatically generate a title slide and content slides, placing the images in the appropriate positions. The user then checks the generated slides, makes minor edits to the content, and finally saves them in .pptx format via output means.
[0657] In this way, the present invention allows users to create high-quality presentation materials in a short amount of time.
[0658] The processing flow will be explained below.
[0659] Step 1:
[0660] The user starts voice input. When the user speaks into the microphone, the device captures the voice data through the microphone.
[0661] Step 2:
[0662] The terminal transmits the captured voice data to the server, which converts the voice data into text data via a voice recognition means.
[0663] Step 3:
[0664] The server displays the converted text data to the user for confirmation, and the user checks the converted text data and corrects it if necessary.
[0665] Step 4:
[0666] The user uses the device to select the images and graphs they want to use and upload them to the server, where the device sends the image files to the server.
[0667] Step 5:
[0668] The server reads the uploaded image data using the image input means and stores it in a format that can be processed within the system.
[0669] Step 6:
[0670] The server uses a generating means to automatically generate presentation materials based on the text data and image data, and the generating means optimizes the layout of slides, the positioning of text, and the positioning of images.
[0671] Step 7:
[0672] The server delivers the generated presentation materials to the user, who then checks the generated slides through their terminal.
[0673] Step 8:
[0674] The user uses the editing means to make necessary minor corrections to the generated presentation material, such as correcting text or rearranging images.
[0675] Step 9:
[0676] Once the editing is complete, the user notifies the server, which then uses the output means to save the final presentation in a predetermined format (e.g., .pptx format).
[0677] Step 10:
[0678] The server provides the saved materials to the user's device in a downloadable format or sends them to a specified email address or cloud storage.
[0679] Example 1
[0680] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0681] Conventional presentation creation systems require users to manually input text, arrange images, and consider the slide layout, which is time-consuming and labor-intensive. Furthermore, the lack of integration of voice recognition and automatic generation functions makes it difficult to efficiently create high-quality presentation materials.
[0682] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0683] In this invention, the server includes a speech recognition means, an image processing means, a text generation means, a presentation generation means, an editing interface means, and an output saving means, which enables the conversion of speech input into text data, the processing of image data, and the automatic generation of presentation materials by combining text and images, which can then be easily edited and saved.
[0684] "Speech recognition means" is a system that has the function of analyzing voice data and converting it into text data.
[0685] The "image processing means" is a system that receives image data uploaded by users, converts it into a processable format, and saves it.
[0686] The "text generation means" is a system that has the function of using text data generated from voice data analyzed by the voice recognition means.
[0687] A "presentation generation means" is a system that combines text data and image data to automatically generate slides and presentation materials.
[0688] The "editing interface means" is a system that has the function of allowing users to easily modify and edit the generated presentation materials.
[0689] The "output saving means" is a system that has the function of saving the final presentation materials in a specified format and generating a link that allows users to download them.
[0690] As a specific embodiment of the present invention, a system including a speech recognition means, an image processing means, a text generation means, a presentation generation means, an editing interface means, and an output storage means will be described.
[0691] This system receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. The user uses a microphone to input the content of their presentation by voice, and the server uses voice recognition software (e.g., a voice recognition API) to convert this voice into text data.
[0692] Next, the user uses the terminal to upload the images and graphs they want to use to the server. The image processing means reads the uploaded image data, converts it into a format that can be processed within the system, and stores it. At this time, the user can upload multiple images and graphs at once. Specifically, the server uses image processing software (e.g., OpenCV) to perform the processing.
[0693] The server then uses a presentation generation means to combine the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. This presentation generation means includes algorithms for determining the optimal layout of text and images and the composition of slides. Specifically, the server uses presentation generation software (e.g., Python-PPTx).
[0694] The generated presentation materials can then be fine-tuned by the user through an editing interface. The user can check the contents of the slides via their terminal and modify text or rearrange images as necessary. This editing interface is designed to provide an intuitive interface that allows users to easily make modifications.
[0695] Finally, after the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output saving means, which has the function of saving the materials in a downloadable format or sending them to a specified email address or cloud storage.
[0696] As a concrete example, consider the following scenario. A user wishes to "create a presentation that explains market growth," and says via speech recognition means, "This slide will explain market growth." The user then uploads an image file called "market growth.png." The server receives this data and uses the presentation generation means to automatically generate a title slide and content slides, placing the images in the appropriate positions. The user then checks the generated slides, makes minor edits to the content, and finally saves them in .pptx format via the output saving means.
[0697] Example prompt sentence:
[0698] "Create a presentation that explains market growth. This slide explains market growth."
[0699] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0700] Step 1:
[0701] The user speaks into the microphone to input the content of the presentation. The server receives the voice data through a voice recognition means. This voice data is the input, and the server sends this voice data to voice recognition software (e.g., a voice recognition API) to obtain text data. The obtained text data is the output.
[0702] Specific behavior:
[0703] The user speaks into the microphone, "This slide explains market growth."
[0704] The server captures voice data in real time and sends it to the voice recognition API for analysis.
[0705] The speech recognition API converts the speech data into text data and returns the text data to the server.
[0706] Step 2:
[0707] A user uses a terminal to upload images or graphs they want to use to the server. This image data is the input, and the server receives this data using image processing means. The image processing means uses image processing software such as OpenCV to convert the data into a processable format and save it within the system. This converted image data is the output.
[0708] Specific behavior:
[0709] The user logs into the device's web interface, selects the image file "Market Growth.png," and clicks the upload button.
[0710] The server processes the received image data using OpenCV and stores it in its internal storage.
[0711] Step 3:
[0712] The server receives the text data obtained from the speech recognition means and the image data obtained from the image processing means, and automatically generates presentation materials using the presentation generation means. At this stage, the input is text data and image data, and the output is the presentation materials. The presentation generation means uses software such as Python-PPTx.
[0713] Specific behavior:
[0714] The server obtains the text data "This slide explains market growth" and the image data "market growth.png".
[0715] Use the Python-PPTx library to create a new presentation file and place text and images on slides.
[0716] The text will be displayed in the title area and the image will be automatically placed in the content area.
[0717] Step 4:
[0718] The user uses the terminal to check the generated slides and make minor corrections through the editing interface means. The input of this step is the generated presentation material, and the output is the corrected presentation material.
[0719] Specific behavior:
[0720] The user previews the generated slides in a web interface.
[0721] Users can click on the text boxes to modify their content or drag the image to change its position.
[0722] Once the user has completed the edits, they click the "Save" button and the edits are sent to the server.
[0723] Step 5:
[0724] The user signals the server that they are finished editing. The server uses the output storage facility to save the final presentation in the specified format (e.g., .pptx) and generate a download link. The input to this step is the modified presentation, and the output is the saved presentation and a download link.
[0725] Specific behavior:
[0726] The user clicks the "Done" button to notify the server that editing is complete.
[0727] The server uses Python-PPTx to export the final presentation in .pptx format.
[0728] The server saves the generated file in the user's download folder, generates a download link, and notifies the user.
[0729] (Application example 1)
[0730] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0731] Current virtual stores lack systems that can quickly and effectively provide product explanations and comparison presentations. This makes it difficult for users to easily obtain detailed information about products, which can hinder purchasing behavior. Furthermore, traditional methods require a lot of time and effort to create presentation materials, making them unsuitable for use in virtual stores, where immediacy is required.
[0732] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0733] In this invention, the server includes a voice recognition means, an image input means, a generation means, an editing means, an output means, a visual output means for automatically generating product descriptions and comparative presentations in the virtual store, and an integration means for creating a presentation based on the voice input and image data, thereby enabling a user to quickly automatically generate product descriptions and comparative presentations in the virtual store based on the voice input and image data.
[0734] A "voice recognition means" is a device or system that has the function of converting voice data into text data.
[0735] "Image input means" refers to a device or system that has the function of receiving and processing images and graphs uploaded by users.
[0736] The "generation means" is a device or system having an algorithm that combines text data obtained from the voice recognition means and image data obtained from the image input means to automatically create presentation materials.
[0737] The "editing means" is a device or system that has the function of providing an interface that allows the user to make minor edits to the generated presentation materials.
[0738] "Output means" refers to a device or system that has the function of saving edited presentation materials in a predetermined format, making them available for download, or sending them to a specified location.
[0739] The "visual output means for automatically generating product explanations and comparative presentations within a virtual store" is a device or system that has the function of visually providing product explanations and comparative presentations to users within a virtual store.
[0740] The "integration means for creating a presentation based on voice input and image data" is a device or system having an algorithm for automatically generating presentation materials by integrating voice input and image data.
[0741] As a specific embodiment of the present invention, a system for automatically generating product explanations and comparison presentations in a virtual store, which is configured by the following steps, will be described.
[0742] 1. Voice recognition means:
[0743] The server receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. For example, when a user uses a microphone to explain the features and advantages of a product, the voice data is sent to the server and converted into text data.
[0744] 2. Image input method:
[0745] Users use their terminals to upload images and graphs related to products to the server. The image input means reads the uploaded image data and stores it in a format that can be processed within the system. Users can upload multiple images and graphs at once.
[0746] 3. Generation means:
[0747] The server combines the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. This generation means uses a generative AI model to optimize the arrangement of text and images and the composition of slides.
[0748] 4. Editing methods:
[0749] The generated presentation materials can then be fine-tuned by the user. The user can check the contents of the slides via their terminal and modify the text or rearrange images as necessary. This editing tool provides an intuitive interface and is designed to allow users to make edits easily.
[0750] 5. Output Method:
[0751] After the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means that can save the materials in a downloadable format or send them to a specified email address or cloud storage.
[0752] Hardware / Software used
[0753] Hardware:
[0754] Microphone (e.g. USB microphone)
[0755] Computer or smartphone
[0756] software:
[0757] Python 3.x
[0758] speech_recognition module
[0759] python-pptx module
[0760] Google Speech Recognition API
[0761] Specific examples
[0762] For example, imagine a user is trying to explain a new electronics product in a virtual store. The user might say into the microphone, "This product has the latest technology and twice the battery life of previous models."
[0763] Next, the user uploads product images and graphs of specifications via their device. The server receives this data and uses a generative AI model to automatically generate a presentation in the form of the following prompt:
[0764] Example prompt sentence:
[0765] "I'm going to create a presentation in a virtual store to explain a new electronics product. Here's what I'm saying: This product is equipped with the latest technology and has twice the battery life of previous models. I've uploaded some voice-recognized text and related product images. Please generate the presentation materials based on this."
[0766] In this way, the present invention enables users to create high-quality presentation materials in a short amount of time.
[0767] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0768] Step 1:
[0769] The user uses a microphone to input a description of the product by voice. The voice input is sent to the server as voice data via the microphone. The server converts this voice data into text data using a voice recognition means. Specifically, the server analyzes the voice data using the Google voice recognition API and generates corresponding text data.
[0770] Input: Audio data of product description
[0771] Data processing: Converting voice data into text data
[0772] Output: Text data
[0773] Step 2:
[0774] Users use their terminals to upload images and graphs related to products. The server receives these image data using image input means and stores them in the system in a processable format. Users can upload multiple images and graphs at once.
[0775] Input: Image file (e.g. product image, specification graph)
[0776] Data processing: Converting image data into a format that can be processed within the system
[0777] Output: Image data in a processable format
[0778] Step 3:
[0779] The server uses the generation means to automatically generate presentation materials by combining the text data obtained from the voice recognition means and the image data obtained from the image input means, and uses the generative AI model to determine the optimal arrangement of text and images and the composition of slides to create the presentation materials.
[0780] Input: Text data, image data
[0781] Data processing: Automatically generate presentation materials by combining text and images
[0782] Output: Presentation materials
[0783] Step 4:
[0784] The user can check the content of the generated presentation materials through the terminal and make any necessary corrections using the editing tool, which provides an intuitive interface and is designed to allow the user to easily modify text and rearrange images.
[0785] Input: Generated presentation materials
[0786] Data processing: Minor edits to presentation content (text edits, image rearrangement)
[0787] Output: Revised presentation
[0788] Step 5:
[0789] After the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means that can save the materials in a downloadable format or send them to a specified email address or cloud storage.
[0790] Input: Revised presentation
[0791] Data processing: Save and output presentation materials in a specified format
[0792] Output: Final presentation file (.pptx format)
[0793] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0794] As a specific embodiment of the present invention, a system including a voice recognition means, an image input means, a generation means, an editing means, an output means, and an emotion engine will be described.
[0795] First, the system receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. The user uses a microphone to input the content of their presentation by voice, and the server converts this voice into text data. At the same time, the emotion engine recognizes the user's emotional state from their voice tone and speaking style.
[0796] Next, the user uses the terminal to upload the images and graphs they want to use to the server. The image input means reads the uploaded image data and stores it in a format that can be processed within the system. At this time, the user can upload multiple images and graphs at once.
[0797] The server then uses a generation means to combine the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. The generation means uses algorithms to optimize the placement of text and images and the composition of slides. Furthermore, an emotion engine adjusts the content and tone of the presentation materials based on the user's emotional state.
[0798] For example, if the user is in a positive emotional state, the generator will suggest recommended expressions and images. As a specific example, if the user requests "I want to create a presentation explaining market growth" and the emotion engine detects the user's excitement, the server will adjust the insertion of emphasized expressions and bright images.
[0799] Conversely, if the user shows a negative emotional state, the server will provide encouraging messages and advice. For example, if the user shows signs of lacking confidence, the server will display a message such as, "This is your area of expertise. Continue your presentation with confidence."
[0800] The generated presentation materials can then be edited by the user through the editing tool. The user can check the generated slides on their device and edit text or rearrange images as needed. This editing tool provides an intuitive interface and is designed to allow users to make edits easily.
[0801] Finally, after the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means, which has the function of saving the materials in a downloadable format or sending them to a specified email address or cloud storage.
[0802] This system enables users to create high-quality, emotionally sensitive presentation materials in a short amount of time, significantly reducing the effort required for creating materials and helping users maximize the effectiveness of their presentations.
[0803] The processing flow will be explained below.
[0804] Step 1:
[0805] The user starts voice input. When the user speaks into the microphone, the device captures the voice data through the microphone.
[0806] Step 2:
[0807] The terminal transmits the captured voice data to the server, which converts the voice data into text data via a voice recognition means.
[0808] Step 3:
[0809] The server uses an emotion engine to analyze the user's emotional state from the voice data, and this analysis determines whether the user has a positive emotion, a negative emotion, or a neutral emotion.
[0810] Step 4:
[0811] The server displays the converted text data to the user for confirmation, and the user checks the converted text data and corrects it if necessary.
[0812] Step 5:
[0813] The user uses the device to select the images and graphs they want to use and upload them to the server, where the device sends the image files to the server.
[0814] Step 6:
[0815] The server reads the uploaded image data using the image input means and stores it in a format that can be processed within the system.
[0816] Step 7:
[0817] The server uses the generation means to automatically generate presentation materials based on the text data and image data. During this generation process, the tone and content of the materials are adjusted based on the analysis results of the emotion engine.
[0818] Step 8:
[0819] If the user is in a positive emotional state, the server adds suggested phrases and upbeat images, such as inserting powerful words and graphs to highlight a slide that explains market growth.
[0820] Step 9:
[0821] If the user is showing a negative emotional state, the server will provide encouraging messages and advice, such as "You've spent a lot of time preparing for your presentation. Go ahead and give it with confidence."
[0822] Step 10:
[0823] The server delivers the generated presentation materials to the user, who then checks the generated slides through their terminal.
[0824] Step 11:
[0825] The user uses the editing means to make necessary minor corrections to the generated presentation material, such as correcting text or rearranging images.
[0826] Step 12:
[0827] Once the editing is complete, the user notifies the server, which then uses the output means to save the final presentation in a predetermined format (e.g., .pptx format).
[0828] Step 13:
[0829] The server provides the saved materials to the user's device in a downloadable format or sends them to a specified email address or cloud storage.
[0830] Example 2
[0831] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0832] Conventional presentation creation systems have the ability to convert audio data into text data and then combine it with image data to automatically generate presentation materials. However, they lack the ability to adjust the content and tone based on the user's emotions, making it difficult to efficiently create presentation materials that take emotions into consideration. This has resulted in the problem of not maximizing the effectiveness of the presentation. Additionally, editing the generated materials is often not intuitive, which increases the user's workload.
[0833] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a voice recognition means, an image input means, a generation means, an emotion analysis means, an editing means, and an output means. This makes it possible to convert voice data to text, import image data, adjust the content of materials based on the emotional state, intuitively edit materials, and save them in a predetermined format.
[0834] The "voice recognition means" is a means having a function of inputting voice data and converting the voice data into text data.
[0835] "Image input means" refers to a means for importing image data and graph data uploaded by users into a format that can be processed within the system.
[0836] The "generation means" is a means for automatically generating presentation materials by combining text data obtained from the voice recognition means and image data obtained from the image input means.
[0837] "Editing means" refers to a means by which users can modify and edit the generated presentation materials through an intuitive interface.
[0838] "Emotion analysis means" is a means for analyzing the user's emotional state from the tone of voice and speaking style, and reflecting that information in the presentation materials.
[0839] The "output means" is a means for saving the final edited presentation materials in a predetermined format and transmitting them to a specified location or in a format that allows the user to download them.
[0840] A "server" is a device or system that receives audio data and image data from a client, processes and analyzes them, and supports the creation and editing of presentation materials.
[0841] The present invention relates to a presentation material creation system that includes a voice recognition unit, an image input unit, a generation unit, an emotion analysis unit, an editing unit, and an output unit.
[0842] First, the system receives voice input from the user via a voice recognition means. The user uses a microphone to input the presentation content by voice. Voice input begins through a dedicated application on the device or a browser screen. This voice data is sent to the server. The server then converts the received voice data into text data in real time using a voice recognition means such as the Google Speech-to-Text API. At this time, an emotion analysis means analyzes the tone and rhythm of the voice to recognize the user's emotional state. This emotional state is then used in the subsequent stage of generating presentation materials.
[0843] Next, the user uses the terminal to upload images and graphs to be used in the presentation to the server. The user selects multiple image files at once from a file selection dialog using the terminal's browser and presses the upload button. The server receives these image data, and the image input means stores them in a format that can be processed within the system.
[0844] The server uses a generation means to combine text data obtained by the voice recognition means with image data obtained by the image input means to automatically generate presentation materials. Libraries such as Apache POI are used in this generation process. The generation means arranges the text data and image data on slides in an optimal layout and adjusts the content and tone of the materials based on the results of the emotion analysis means. If positive emotions are recognized, bright colors and emphasized expressions are used frequently. On the other hand, if negative emotions are recognized, encouraging messages and advice are added.
[0845] For example, if a user enters the following prompt into the system:
[0846] I want to create a presentation that explains the growth of the market. Emotional state: Excitement
[0847] In this case, the server uses sentiment analysis to recognize the user's excitement and generates slides with a positive tone, featuring emphatic expressions and upbeat images.
[0848] The generated presentation materials can be edited by the user through the editing tool. The user can check the generated slides using the terminal and edit the text or rearrange the images as needed. This editing tool provides an intuitive interface and is designed to allow the user to make edits easily.
[0849] Finally, after the user has finished editing, the server uses the output means to save the final presentation in a predetermined format (e.g., .pptx format), which is then made available to the user in a downloadable format or sent to a specified email address or cloud storage.
[0850] This system allows users to create high-quality, emotionally sensitive presentation materials in a short amount of time, significantly reducing the effort required to create materials.
[0851] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0852] Step 1: Accepting voice input
[0853] The user uses the device's microphone to input the presentation content by voice, starting voice input through a dedicated application or browser screen on the device.
[0854] Input: User's voice data
[0855] Output: Audio data sent to the server
[0856] Specific behavior:
[0857] The user speaks the content of the presentation into the microphone.
[0858] An application on the device captures the user's voice in real time.
[0859] Send the captured audio data to the server.
[0860] Step 2: Convert audio data to text
[0861] The server converts the received voice data into text data in real time using a voice recognition means.
[0862] Input: Audio data
[0863] Output: Text data
[0864] Specific behavior:
[0865] The server passes the audio data to the Google Speech-to-Text API.
[0866] The API analyzes the audio data and generates corresponding text data.
[0867] The server receives the generated text data.
[0868] Step 3: Recognizing your emotional state
[0869] The server uses an emotion analysis means to analyze the user's emotional state from the text data and voice tone.
[0870] Input: Text data, audio tone
[0871] Output: User's emotional state
[0872] Specific behavior:
[0873] The server passes the text data and voice tone information to the emotion analysis engine.
[0874] The emotion analysis engine analyzes voice characteristics such as intonation and speed to identify the user's emotional state (e.g., positive, negative, excited, etc.).
[0875] The server obtains the user's emotional state as the analysis result.
[0876] Step 4: Upload images and graphs
[0877] The user uses the terminal to upload images and graphs to be used in the presentation to the server.
[0878] Input: Image data, graph data
[0879] Output: Image and graph data sent to the server
[0880] Specific behavior:
[0881] Displays a file selection dialog in the device's browser.
[0882] The user selects the file to upload and presses the upload button.
[0883] The selected file is sent from the terminal to the server.
[0884] Step 5: Join the data and generate a presentation
[0885] The server uses the generating means to combine the text data obtained from the voice recognition means and the image data obtained from the image input means to automatically generate presentation materials.
[0886] Input: Text data, image data, emotional state
[0887] Output: Presentation materials
[0888] Specific behavior:
[0889] The server acquires the text data and the image data.
[0890] A generating means places the text data and image data on the slide.
[0891] Adjust the content and tone of your materials based on your emotional state (use bright colors if you're positive).
[0892] Step 6: Edit your presentation
[0893] The user checks the generated slides via the terminal and makes any necessary corrections using the editing means.
[0894] Input: Presentation materials
[0895] Output: Revised presentation materials
[0896] Specific behavior:
[0897] The presentation materials generated on the terminal are displayed.
[0898] The user modifies text and images through a GUI.
[0899] The changes are updated on the server in real time.
[0900] Step 7: Finalize your presentation
[0901] The server uses the output means to save the final presentation materials in a predetermined format (for example, .pptx format).
[0902] Input: Revised presentation materials
[0903] Output: Presentation material saved in the specified format
[0904] Specific behavior:
[0905] The final presentation materials are generated on the server.
[0906] Export the generated materials in a specified format (such as pptx).
[0907] Generate and provide a URL from which users can download the materials, or send it to a specified email address.
[0908] This process allows users to efficiently create high-quality, emotionally sensitive presentation materials in a short amount of time.
[0909] (Application example 2)
[0910] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0911] Existing customer service systems have difficulty generating appropriate responses to customer questions and requests immediately. Furthermore, they respond without taking into account the emotional state of staff, resulting in inconsistent quality of customer service. Furthermore, it is difficult for staff to gain confidence when dealing with customers, which can lead to a decline in motivation.
[0912] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, an image input means, an emotion analysis means, a generation means, an editing means, and an output means. This allows the staff member to instantly convert voice input into text data through the smart glasses, generate a response based on the emotional state, and display it on the smart glasses' display. This improves the quality of customer service, increases staff motivation, and enables efficient work performance.
[0913] "Speech recognition means" is a means for inputting speech and converting it into text data.
[0914] "Image input means" refers to means for taking in image data and converting it into a format that can be processed within the system.
[0915] The "emotion analysis means" is a means for analyzing the user's emotional state from voice data, image data, and the like.
[0916] The "generation means" is a means for automatically generating presentation materials and responses by combining text data obtained from the speech recognition means with image data and emotional states obtained from the image input means.
[0917] "Editing tools" are tools that allow users to fine-tune and adjust automatically generated materials and responses.
[0918] "Output means" refers to a means for saving the final generated materials and responses in a predetermined format and transmitting them in a downloadable format or to cloud storage as needed.
[0919] As an embodiment of the present invention, a system including a voice recognition means, an image input means, an emotion analysis means, a generation means, an editing means, and an output means will be described.
[0920] First, the voice recognition means has the function of capturing the user's voice through a microphone installed in the smart glasses and converting that voice into text data. For voice recognition, for example, the speech_recognition library is used. This library converts the user's voice into text through Google's voice recognition service.
[0921] Next, the emotion analysis means analyzes the user's emotional state from the voice data. This analysis is performed using the emotion_recognition library. This library analyzes the tone and pitch of the voice data to determine whether the user is friendly or stressed.
[0922] The generation means combines the text data obtained from the speech recognition means with the emotional state obtained from the emotion analysis means to generate an appropriate response or guide message. This generation uses a generative AI model (e.g., GPT-3). The generated response is adjusted based on the user's emotional state.
[0923] For example, if a customer asks, "What's the difference between these products?", the voice is first converted into text data. Next, an emotion analysis tool detects the friendly emotional state of the staff member. Based on this information, the generative AI model generates a response such as, "I'll explain the difference between these products in an easy-to-understand way."
[0924] The editing tool allows users to fine-tune the generated responses and presentation materials. An intuitive user interface (UI) is used for editing, allowing users to easily modify text and rearrange images.
[0925] Finally, the output means saves the generated materials and responses in a predetermined format, displays them on the smart glasses display, or saves them in a downloadable format. This function can also be linked to cloud storage services (e.g., Google Drive or Dropbox).
[0926] As a concrete example, the following prompt sentences may be fed into a generative AI model to generate a response:
[0927] Voice: A customer asks, "What's different about this product?" The staff member's emotional state is friendly. What response would you generate?
[0928] This system will improve the quality of customer service, increase staff motivation, and enable efficient work execution.
[0929] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0930] Step 1:
[0931] The user wears the smart glasses and speaks their questions and requests into the microphone. This voice data is temporarily stored in the smart glasses' internal memory.
[0932] Step 2:
[0933] The voice data is sent to the server, and the server converts the voice data to text data using the speech_recognition library. The voice data is analyzed and output as string data. This step converts voice input to text data.
[0934] Step 3:
[0935] Based on the converted text data, the server performs emotion analysis of the audio data using the emotion_recognition library. It analyzes the tone, pitch, and speed of the audio to detect the user's emotional state. It then outputs the emotional state (e.g., friendly, stressed).
[0936] Step 4:
[0937] The server uses a generative AI model (e.g., GPT-3) to generate an appropriate response based on the input text data and emotional state. In this step, the appropriate response is input as a prompt to the generative AI model, and the generated text response is output.
[0938] Step 5:
[0939] The generated response is displayed for the user to review on the smart glasses display, and the user has an interface to edit the response if necessary. If edits are made, the revised response data is resubmitted to the server and the final version is saved.
[0940] Step 6:
[0941] The server saves the final response data in a specified format and sends it to cloud storage or a specified email address as needed. In this step, data is saved and output.
[0942] Through the above processing steps, the user can instantly generate, display, edit, and save appropriate responses to customer questions and requests.
[0943] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0944] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0945] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0946] [Fourth embodiment]
[0947] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0948] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0949] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0950] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0951] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0952] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0953] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0954] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0955] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0956] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0957] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0958] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0959] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0960] As a specific embodiment of the present invention, a system including a voice recognition means, an image input means, a generation means, an editing means, and an output means will be described.
[0961] First, this system receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. The user uses a microphone to input the content of their presentation by voice, and the server converts this voice into text data.
[0962] Next, the user uses the terminal to upload the images and graphs they want to use to the server. The image input means reads the uploaded image data and stores it in a format that can be processed within the system. At this time, the user can upload multiple images and graphs at once.
[0963] The server then uses a generation means to combine the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. This generation means uses an algorithm to optimize the arrangement of text and images and the composition of slides.
[0964] The generated presentation materials can then be edited by the user through the editing tool. The user can check the contents of the slides via their terminal and modify the text or rearrange images as necessary. This editing tool provides an intuitive interface and is designed to allow users to make edits easily.
[0965] Finally, after the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means, which has the function of saving the materials in a downloadable format or sending them to a specified email address or cloud storage.
[0966] As a concrete example, consider the following scenario. A user wishes to "create a presentation that explains market growth," and says "This slide will explain market growth" via speech recognition means. The user then uploads an image file called "market growth.png." The server receives this data and uses generation means to automatically generate a title slide and content slides, placing the images in the appropriate positions. The user then checks the generated slides, makes minor edits to the content, and finally saves them in .pptx format via output means.
[0967] In this way, the present invention allows users to create high-quality presentation materials in a short amount of time.
[0968] The processing flow will be explained below.
[0969] Step 1:
[0970] The user starts voice input. When the user speaks into the microphone, the device captures the voice data through the microphone.
[0971] Step 2:
[0972] The terminal transmits the captured voice data to the server, which converts the voice data into text data via a voice recognition means.
[0973] Step 3:
[0974] The server displays the converted text data to the user for confirmation, and the user checks the converted text data and corrects it if necessary.
[0975] Step 4:
[0976] The user uses the device to select the images and graphs they want to use and upload them to the server, where the device sends the image files to the server.
[0977] Step 5:
[0978] The server reads the uploaded image data using the image input means and stores it in a format that can be processed within the system.
[0979] Step 6:
[0980] The server uses a generating means to automatically generate presentation materials based on the text data and image data, and the generating means optimizes the layout of slides, the positioning of text, and the positioning of images.
[0981] Step 7:
[0982] The server delivers the generated presentation materials to the user, who then checks the generated slides through their terminal.
[0983] Step 8:
[0984] The user uses the editing means to make necessary minor corrections to the generated presentation material, such as correcting text or rearranging images.
[0985] Step 9:
[0986] Once the editing is complete, the user notifies the server, which then uses the output means to save the final presentation in a predetermined format (e.g., .pptx format).
[0987] Step 10:
[0988] The server provides the saved materials to the user's device in a downloadable format or sends them to a specified email address or cloud storage.
[0989] Example 1
[0990] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0991] Conventional presentation creation systems require users to manually input text, arrange images, and consider the slide layout, which is time-consuming and labor-intensive. Furthermore, the lack of integration of voice recognition and automatic generation functions makes it difficult to efficiently create high-quality presentation materials.
[0992] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0993] In this invention, the server includes a speech recognition means, an image processing means, a text generation means, a presentation generation means, an editing interface means, and an output saving means, which enables the conversion of speech input into text data, the processing of image data, and the automatic generation of presentation materials by combining text and images, which can then be easily edited and saved.
[0994] "Speech recognition means" is a system that has the function of analyzing voice data and converting it into text data.
[0995] The "image processing means" is a system that receives image data uploaded by users, converts it into a processable format, and saves it.
[0996] The "text generation means" is a system that has the function of using text data generated from voice data analyzed by the voice recognition means.
[0997] A "presentation generation means" is a system that combines text data and image data to automatically generate slides and presentation materials.
[0998] The "editing interface means" is a system that has the function of allowing users to easily modify and edit the generated presentation materials.
[0999] The "output saving means" is a system that has the function of saving the final presentation materials in a specified format and generating a link that allows users to download them.
[1000] As a specific embodiment of the present invention, a system including a speech recognition means, an image processing means, a text generation means, a presentation generation means, an editing interface means, and an output storage means will be described.
[1001] This system receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. The user uses a microphone to input the content of their presentation by voice, and the server uses voice recognition software (e.g., a voice recognition API) to convert this voice into text data.
[1002] Next, the user uses the terminal to upload the images and graphs they want to use to the server. The image processing means reads the uploaded image data, converts it into a format that can be processed within the system, and stores it. At this time, the user can upload multiple images and graphs at once. Specifically, the server uses image processing software (e.g., OpenCV) to perform the processing.
[1003] The server then uses a presentation generation means to combine the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. This presentation generation means includes algorithms for determining the optimal layout of text and images and the composition of slides. Specifically, the server uses presentation generation software (e.g., Python-PPTx).
[1004] The generated presentation materials can then be fine-tuned by the user through an editing interface. The user can check the contents of the slides via their terminal and modify text or rearrange images as necessary. This editing interface is designed to provide an intuitive interface that allows users to easily make modifications.
[1005] Finally, after the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output saving means, which has the function of saving the materials in a downloadable format or sending them to a specified email address or cloud storage.
[1006] As a concrete example, consider the following scenario. A user wishes to "create a presentation that explains market growth," and says via speech recognition means, "This slide will explain market growth." The user then uploads an image file called "market growth.png." The server receives this data and uses the presentation generation means to automatically generate a title slide and content slides, placing the images in the appropriate positions. The user then checks the generated slides, makes minor edits to the content, and finally saves them in .pptx format via the output saving means.
[1007] Example prompt sentence:
[1008] "Create a presentation that explains market growth. This slide explains market growth."
[1009] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1010] Step 1:
[1011] The user speaks into the microphone to input the content of the presentation. The server receives the voice data through a voice recognition means. This voice data is the input, and the server sends this voice data to voice recognition software (e.g., a voice recognition API) to obtain text data. The obtained text data is the output.
[1012] Specific behavior:
[1013] The user speaks into the microphone, "This slide explains market growth."
[1014] The server captures voice data in real time and sends it to the voice recognition API for analysis.
[1015] The speech recognition API converts the speech data into text data and returns the text data to the server.
[1016] Step 2:
[1017] A user uses a terminal to upload images or graphs they want to use to the server. This image data is the input, and the server receives this data using image processing means. The image processing means uses image processing software such as OpenCV to convert the data into a processable format and save it within the system. This converted image data is the output.
[1018] Specific behavior:
[1019] The user logs into the device's web interface, selects the image file "Market Growth.png," and clicks the upload button.
[1020] The server processes the received image data using OpenCV and stores it in its internal storage.
[1021] Step 3:
[1022] The server receives the text data obtained from the speech recognition means and the image data obtained from the image processing means, and automatically generates presentation materials using the presentation generation means. At this stage, the input is text data and image data, and the output is the presentation materials. The presentation generation means uses software such as Python-PPTx.
[1023] Specific behavior:
[1024] The server obtains the text data "This slide explains market growth" and the image data "market growth.png".
[1025] Use the Python-PPTx library to create a new presentation file and place text and images on slides.
[1026] The text will be displayed in the title area and the image will be automatically placed in the content area.
[1027] Step 4:
[1028] The user uses the terminal to check the generated slides and make minor corrections through the editing interface means. The input of this step is the generated presentation material, and the output is the corrected presentation material.
[1029] Specific behavior:
[1030] The user previews the generated slides in a web interface.
[1031] Users can click on the text boxes to modify their content or drag the image to change its position.
[1032] Once the user has completed the edits, they click the "Save" button and the edits are sent to the server.
[1033] Step 5:
[1034] The user signals the server that they are finished editing. The server uses the output storage facility to save the final presentation in the specified format (e.g., .pptx) and generate a download link. The input to this step is the modified presentation, and the output is the saved presentation and a download link.
[1035] Specific behavior:
[1036] The user clicks the "Done" button to notify the server that editing is complete.
[1037] The server uses Python-PPTx to export the final presentation in .pptx format.
[1038] The server saves the generated file in the user's download folder, generates a download link, and notifies the user.
[1039] (Application example 1)
[1040] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1041] Current virtual stores lack systems that can quickly and effectively provide product explanations and comparison presentations. This makes it difficult for users to easily obtain detailed information about products, which can hinder purchasing behavior. Furthermore, traditional methods require a lot of time and effort to create presentation materials, making them unsuitable for use in virtual stores, where immediacy is required.
[1042] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1043] In this invention, the server includes a voice recognition means, an image input means, a generation means, an editing means, an output means, a visual output means for automatically generating product descriptions and comparative presentations in the virtual store, and an integration means for creating a presentation based on the voice input and image data, thereby enabling a user to quickly automatically generate product descriptions and comparative presentations in the virtual store based on the voice input and image data.
[1044] A "voice recognition means" is a device or system that has the function of converting voice data into text data.
[1045] "Image input means" refers to a device or system that has the function of receiving and processing images and graphs uploaded by users.
[1046] The "generation means" is a device or system having an algorithm that combines text data obtained from the voice recognition means and image data obtained from the image input means to automatically create presentation materials.
[1047] The "editing means" is a device or system that has the function of providing an interface that allows the user to make minor edits to the generated presentation materials.
[1048] "Output means" refers to a device or system that has the function of saving edited presentation materials in a predetermined format, making them available for download, or sending them to a specified location.
[1049] The "visual output means for automatically generating product explanations and comparative presentations within a virtual store" is a device or system that has the function of visually providing product explanations and comparative presentations to users within a virtual store.
[1050] The "integration means for creating a presentation based on voice input and image data" is a device or system having an algorithm for automatically generating presentation materials by integrating voice input and image data.
[1051] As a specific embodiment of the present invention, a system for automatically generating product explanations and comparison presentations in a virtual store, which is configured by the following steps, will be described.
[1052] 1. Voice recognition means:
[1053] The server receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. For example, when a user uses a microphone to explain the features and advantages of a product, the voice data is sent to the server and converted into text data.
[1054] 2. Image input method:
[1055] Users use their terminals to upload images and graphs related to products to the server. The image input means reads the uploaded image data and stores it in a format that can be processed within the system. Users can upload multiple images and graphs at once.
[1056] 3. Generation means:
[1057] The server combines the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. This generation means uses a generative AI model to optimize the arrangement of text and images and the composition of slides.
[1058] 4. Editing methods:
[1059] The generated presentation materials can then be fine-tuned by the user. The user can check the contents of the slides via their terminal and modify the text or rearrange images as necessary. This editing tool provides an intuitive interface and is designed to allow users to make edits easily.
[1060] 5. Output Method:
[1061] After the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means that can save the materials in a downloadable format or send them to a specified email address or cloud storage.
[1062] Hardware / Software used
[1063] Hardware:
[1064] Microphone (e.g. USB microphone)
[1065] Computer or smartphone
[1066] software:
[1067] Python 3.x
[1068] speech_recognition module
[1069] python-pptx module
[1070] Google Speech Recognition API
[1071] Specific examples
[1072] For example, imagine a user is trying to explain a new electronics product in a virtual store. The user might say into the microphone, "This product has the latest technology and twice the battery life of previous models."
[1073] Next, the user uploads product images and graphs of specifications via their device. The server receives this data and uses a generative AI model to automatically generate a presentation in the form of the following prompt:
[1074] Example prompt sentence:
[1075] "I'm going to create a presentation in a virtual store to explain a new electronics product. Here's what I'm saying: This product is equipped with the latest technology and has twice the battery life of previous models. I've uploaded some voice-recognized text and related product images. Please generate the presentation materials based on this."
[1076] In this way, the present invention enables users to create high-quality presentation materials in a short amount of time.
[1077] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1078] Step 1:
[1079] The user uses a microphone to input a description of the product by voice. The voice input is sent to the server as voice data via the microphone. The server converts this voice data into text data using a voice recognition means. Specifically, the server analyzes the voice data using the Google voice recognition API and generates corresponding text data.
[1080] Input: Audio data of product description
[1081] Data processing: Converting voice data into text data
[1082] Output: Text data
[1083] Step 2:
[1084] Users use their terminals to upload images and graphs related to products. The server receives these image data using image input means and stores them in the system in a processable format. Users can upload multiple images and graphs at once.
[1085] Input: Image file (e.g. product image, specification graph)
[1086] Data processing: Converting image data into a format that can be processed within the system
[1087] Output: Image data in a processable format
[1088] Step 3:
[1089] The server uses the generation means to automatically generate presentation materials by combining the text data obtained from the voice recognition means and the image data obtained from the image input means, and uses the generative AI model to determine the optimal arrangement of text and images and the composition of slides to create the presentation materials.
[1090] Input: Text data, image data
[1091] Data processing: Automatically generate presentation materials by combining text and images
[1092] Output: Presentation materials
[1093] Step 4:
[1094] The user can check the content of the generated presentation materials through the terminal and make any necessary corrections using the editing tool, which provides an intuitive interface and is designed to allow the user to easily modify text and rearrange images.
[1095] Input: Generated presentation materials
[1096] Data processing: Minor edits to presentation content (text edits, image rearrangement)
[1097] Output: Revised presentation
[1098] Step 5:
[1099] After the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means that can save the materials in a downloadable format or send them to a specified email address or cloud storage.
[1100] Input: Revised presentation
[1101] Data processing: Save and output presentation materials in a specified format
[1102] Output: Final presentation file (.pptx format)
[1103] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1104] As a specific embodiment of the present invention, a system including a voice recognition means, an image input means, a generation means, an editing means, an output means, and an emotion engine will be described.
[1105] First, the system receives voice input from the user via a voice recognition means. This voice recognition means has the function of converting voice data into text data. The user uses a microphone to input the content of their presentation by voice, and the server converts this voice into text data. At the same time, the emotion engine recognizes the user's emotional state from their voice tone and speaking style.
[1106] Next, the user uses the terminal to upload the images and graphs they want to use to the server. The image input means reads the uploaded image data and stores it in a format that can be processed within the system. At this time, the user can upload multiple images and graphs at once.
[1107] The server then uses a generation means to combine the text data obtained from the speech recognition means and the image data obtained from the image input means to automatically generate presentation materials. The generation means uses algorithms to optimize the placement of text and images and the composition of slides. Furthermore, an emotion engine adjusts the content and tone of the presentation materials based on the user's emotional state.
[1108] For example, if the user is in a positive emotional state, the generator will suggest recommended expressions and images. As a specific example, if the user requests "I want to create a presentation explaining market growth" and the emotion engine detects the user's excitement, the server will adjust the insertion of emphasized expressions and bright images.
[1109] Conversely, if the user shows a negative emotional state, the server will provide encouraging messages and advice. For example, if the user shows signs of lacking confidence, the server will display a message such as, "This is your area of expertise. Continue your presentation with confidence."
[1110] The generated presentation materials can then be edited by the user through the editing tool. The user can check the generated slides on their device and edit text or rearrange images as needed. This editing tool provides an intuitive interface and is designed to allow users to make edits easily.
[1111] Finally, after the user has finished editing, the server saves the final presentation materials in a predetermined format (e.g., .pptx format) using an output means, which has the function of saving the materials in a downloadable format or sending them to a specified email address or cloud storage.
[1112] This system enables users to create high-quality, emotionally sensitive presentation materials in a short amount of time, significantly reducing the effort required for creating materials and helping users maximize the effectiveness of their presentations.
[1113] The processing flow will be explained below.
[1114] Step 1:
[1115] The user starts voice input. When the user speaks into the microphone, the device captures the voice data through the microphone.
[1116] Step 2:
[1117] The terminal transmits the captured voice data to the server, which converts the voice data into text data via a voice recognition means.
[1118] Step 3:
[1119] The server uses an emotion engine to analyze the user's emotional state from the voice data, and this analysis determines whether the user has a positive emotion, a negative emotion, or a neutral emotion.
[1120] Step 4:
[1121] The server displays the converted text data to the user for confirmation, and the user checks the converted text data and corrects it if necessary.
[1122] Step 5:
[1123] The user uses the device to select the images and graphs they want to use and upload them to the server, where the device sends the image files to the server.
[1124] Step 6:
[1125] The server reads the uploaded image data using the image input means and stores it in a format that can be processed within the system.
[1126] Step 7:
[1127] The server uses the generation means to automatically generate presentation materials based on the text data and image data. During this generation process, the tone and content of the materials are adjusted based on the analysis results of the emotion engine.
[1128] Step 8:
[1129] If the user is in a positive emotional state, the server adds suggested phrases and upbeat images, such as inserting powerful words and graphs to highlight a slide that explains market growth.
[1130] Step 9:
[1131] If the user is showing a negative emotional state, the server will provide encouraging messages and advice, such as "You've spent a lot of time preparing for your presentation. Go ahead and give it with confidence."
[1132] Step 10:
[1133] The server delivers the generated presentation materials to the user, who then checks the generated slides through their terminal.
[1134] Step 11:
[1135] The user uses the editing means to make necessary minor corrections to the generated presentation material, such as correcting text or rearranging images.
[1136] Step 12:
[1137] Once the editing is complete, the user notifies the server, which then uses the output means to save the final presentation in a predetermined format (e.g., .pptx format).
[1138] Step 13:
[1139] The server provides the saved materials to the user's device in a downloadable format or sends them to a specified email address or cloud storage.
[1140] Example 2
[1141] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1142] Conventional presentation creation systems have the ability to convert audio data into text data and then combine it with image data to automatically generate presentation materials. However, they lack the ability to adjust the content and tone based on the user's emotions, making it difficult to efficiently create presentation materials that take emotions into consideration. This has resulted in the problem of not maximizing the effectiveness of the presentation. Additionally, editing the generated materials is often not intuitive, which increases the user's workload.
[1143] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a voice recognition means, an image input means, a generation means, an emotion analysis means, an editing means, and an output means. This makes it possible to convert voice data to text, import image data, adjust the content of materials based on the emotional state, intuitively edit materials, and save them in a predetermined format.
[1144] The "voice recognition means" is a means having a function of inputting voice data and converting the voice data into text data.
[1145] "Image input means" refers to a means for importing image data and graph data uploaded by users into a format that can be processed within the system.
[1146] The "generation means" is a means for automatically generating presentation materials by combining text data obtained from the voice recognition means and image data obtained from the image input means.
[1147] "Editing means" refers to a means by which users can modify and edit the generated presentation materials through an intuitive interface.
[1148] "Emotion analysis means" is a means for analyzing the user's emotional state from the tone of voice and speaking style, and reflecting that information in the presentation materials.
[1149] The "output means" is a means for saving the final edited presentation materials in a predetermined format and transmitting them to a specified location or in a format that allows the user to download them.
[1150] A "server" is a device or system that receives audio data and image data from a client, processes and analyzes them, and supports the creation and editing of presentation materials.
[1151] The present invention relates to a presentation material creation system that includes a voice recognition unit, an image input unit, a generation unit, an emotion analysis unit, an editing unit, and an output unit.
[1152] First, the system receives voice input from the user via a voice recognition means. The user uses a microphone to input the presentation content by voice. Voice input begins through a dedicated application on the device or a browser screen. This voice data is sent to the server. The server then converts the received voice data into text data in real time using a voice recognition means such as the Google Speech-to-Text API. At this time, an emotion analysis means analyzes the tone and rhythm of the voice to recognize the user's emotional state. This emotional state is then used in the subsequent stage of generating presentation materials.
[1153] Next, the user uses the terminal to upload images and graphs to be used in the presentation to the server. The user selects multiple image files at once from a file selection dialog using the terminal's browser and presses the upload button. The server receives these image data, and the image input means stores them in a format that can be processed within the system.
[1154] The server uses a generation means to combine text data obtained by the voice recognition means with image data obtained by the image input means to automatically generate presentation materials. Libraries such as Apache POI are used in this generation process. The generation means arranges the text data and image data on slides in an optimal layout and adjusts the content and tone of the materials based on the results of the emotion analysis means. If positive emotions are recognized, bright colors and emphasized expressions are used frequently. On the other hand, if negative emotions are recognized, encouraging messages and advice are added.
[1155] For example, if a user enters the following prompt into the system:
[1156] I want to create a presentation that explains the growth of the market. Emotional state: Excitement
[1157] In this case, the server uses sentiment analysis to recognize the user's excitement and generates slides with a positive tone, featuring emphatic expressions and upbeat images.
[1158] The generated presentation materials can be edited by the user through the editing tool. The user can check the generated slides using the terminal and edit the text or rearrange the images as needed. This editing tool provides an intuitive interface and is designed to allow the user to make edits easily.
[1159] Finally, after the user has finished editing, the server uses the output means to save the final presentation in a predetermined format (e.g., .pptx format), which is then made available to the user in a downloadable format or sent to a specified email address or cloud storage.
[1160] This system allows users to create high-quality, emotionally sensitive presentation materials in a short amount of time, significantly reducing the effort required to create materials.
[1161] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1162] Step 1: Accepting voice input
[1163] The user uses the device's microphone to input the presentation content by voice, starting voice input through a dedicated application or browser screen on the device.
[1164] Input: User's voice data
[1165] Output: Audio data sent to the server
[1166] Specific behavior:
[1167] The user speaks the content of the presentation into the microphone.
[1168] An application on the device captures the user's voice in real time.
[1169] Send the captured audio data to the server.
[1170] Step 2: Convert audio data to text
[1171] The server converts the received voice data into text data in real time using a voice recognition means.
[1172] Input: Audio data
[1173] Output: Text data
[1174] Specific behavior:
[1175] The server passes the audio data to the Google Speech-to-Text API.
[1176] The API analyzes the audio data and generates corresponding text data.
[1177] The server receives the generated text data.
[1178] Step 3: Recognizing your emotional state
[1179] The server uses an emotion analysis means to analyze the user's emotional state from the text data and voice tone.
[1180] Input: Text data, audio tone
[1181] Output: User's emotional state
[1182] Specific behavior:
[1183] The server passes the text data and voice tone information to the emotion analysis engine.
[1184] The emotion analysis engine analyzes voice characteristics such as intonation and speed to identify the user's emotional state (e.g., positive, negative, excited, etc.).
[1185] The server obtains the user's emotional state as the analysis result.
[1186] Step 4: Upload images and graphs
[1187] The user uses the terminal to upload images and graphs to be used in the presentation to the server.
[1188] Input: Image data, graph data
[1189] Output: Image and graph data sent to the server
[1190] Specific behavior:
[1191] Displays a file selection dialog in the device's browser.
[1192] The user selects the file to upload and presses the upload button.
[1193] The selected file is sent from the terminal to the server.
[1194] Step 5: Join the data and generate a presentation
[1195] The server uses the generating means to combine the text data obtained from the voice recognition means and the image data obtained from the image input means to automatically generate presentation materials.
[1196] Input: Text data, image data, emotional state
[1197] Output: Presentation materials
[1198] Specific behavior:
[1199] The server acquires the text data and the image data.
[1200] A generating means places the text data and image data on the slide.
[1201] Adjust the content and tone of your materials based on your emotional state (use bright colors if you're positive).
[1202] Step 6: Edit your presentation
[1203] The user checks the generated slides via the terminal and makes any necessary corrections using the editing means.
[1204] Input: Presentation materials
[1205] Output: Revised presentation materials
[1206] Specific behavior:
[1207] The presentation materials generated on the terminal are displayed.
[1208] The user modifies text and images through a GUI.
[1209] The changes are updated on the server in real time.
[1210] Step 7: Finalize your presentation
[1211] The server uses the output means to save the final presentation materials in a predetermined format (for example, .pptx format).
[1212] Input: Revised presentation materials
[1213] Output: Presentation material saved in the specified format
[1214] Specific behavior:
[1215] The final presentation materials are generated on the server.
[1216] Export the generated materials in a specified format (such as pptx).
[1217] Generate and provide a URL from which users can download the materials, or send it to a specified email address.
[1218] This process allows users to efficiently create high-quality, emotionally sensitive presentation materials in a short amount of time.
[1219] (Application example 2)
[1220] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1221] Existing customer service systems have difficulty generating appropriate responses to customer questions and requests immediately. Furthermore, they respond without taking into account the emotional state of staff, resulting in inconsistent quality of customer service. Furthermore, it is difficult for staff to gain confidence when dealing with customers, which can lead to a decline in motivation.
[1222] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, an image input means, an emotion analysis means, a generation means, an editing means, and an output means. This allows the staff member to instantly convert voice input into text data through the smart glasses, generate a response based on the emotional state, and display it on the smart glasses' display. This improves the quality of customer service, increases staff motivation, and enables efficient work performance.
[1223] "Speech recognition means" is a means for inputting speech and converting it into text data.
[1224] "Image input means" refers to means for taking in image data and converting it into a format that can be processed within the system.
[1225] The "emotion analysis means" is a means for analyzing the user's emotional state from voice data, image data, and the like.
[1226] The "generation means" is a means for automatically generating presentation materials and responses by combining text data obtained from the speech recognition means with image data and emotional states obtained from the image input means.
[1227] "Editing tools" are tools that allow users to fine-tune and adjust automatically generated materials and responses.
[1228] "Output means" refers to a means for saving the final generated materials and responses in a predetermined format and transmitting them in a downloadable format or to cloud storage as needed.
[1229] As an embodiment of the present invention, a system including a voice recognition means, an image input means, an emotion analysis means, a generation means, an editing means, and an output means will be described.
[1230] First, the voice recognition means has the function of capturing the user's voice through a microphone installed in the smart glasses and converting that voice into text data. For voice recognition, for example, the speech_recognition library is used. This library converts the user's voice into text through Google's voice recognition service.
[1231] Next, the emotion analysis means analyzes the user's emotional state from the voice data. This analysis is performed using the emotion_recognition library. This library analyzes the tone and pitch of the voice data to determine whether the user is friendly or stressed.
[1232] The generation means combines the text data obtained from the speech recognition means with the emotional state obtained from the emotion analysis means to generate an appropriate response or guide message. This generation uses a generative AI model (e.g., GPT-3). The generated response is adjusted based on the user's emotional state.
[1233] For example, if a customer asks, "What's the difference between these products?", the voice is first converted into text data. Next, an emotion analysis tool detects the friendly emotional state of the staff member. Based on this information, the generative AI model generates a response such as, "I'll explain the difference between these products in an easy-to-understand way."
[1234] The editing tool provides users with the ability to fine-tune generated responses and presentation materials. An intuitive user interface (UI) is used for editing, allowing users to easily modify text and rearrange images.
[1235] Finally, the output means saves the generated materials and responses in a predetermined format, displays them on the smart glasses display, or saves them in a downloadable format. This function can also be linked to cloud storage services (e.g., Google Drive or Dropbox).
[1236] As a concrete example, the following prompt sentences may be fed into a generative AI model to generate a response:
[1237] Voice: A customer asks, "What's different about this product?" The staff member's emotional state is friendly. What response would you generate?
[1238] This system will improve the quality of customer service, increase staff motivation, and enable efficient work execution.
[1239] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1240] Step 1:
[1241] The user wears the smart glasses and speaks their questions and requests into the microphone. This voice data is temporarily stored in the smart glasses' internal memory.
[1242] Step 2:
[1243] The voice data is sent to the server, and the server converts the voice data to text data using the speech_recognition library. The voice data is analyzed and output as string data. This step converts voice input to text data.
[1244] Step 3:
[1245] Based on the converted text data, the server performs emotion analysis of the audio data using the emotion_recognition library. It analyzes the tone, pitch, and speed of the audio to detect the user's emotional state. It then outputs the emotional state (e.g., friendly, stressed).
[1246] Step 4:
[1247] The server uses a generative AI model (e.g., GPT-3) to generate an appropriate response based on the input text data and emotional state. In this step, the appropriate response is input as a prompt to the generative AI model, and the generated text response is output.
[1248] Step 5:
[1249] The generated response is displayed for the user to review on the smart glasses display, and the user has an interface to edit the response if necessary. If edits are made, the revised response data is resubmitted to the server and the final version is saved.
[1250] Step 6:
[1251] The server saves the final response data in a specified format and sends it to cloud storage or a specified email address as needed. In this step, data is saved and output.
[1252] Through the above processing steps, the user can instantly generate, display, edit, and save appropriate responses to customer questions and requests.
[1253] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1254] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1255] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1256] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1257] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1258] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1259] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1260] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1261] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1262] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1263] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1264] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1265] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1266] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1267] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1268] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1269] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1270] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1271] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1272] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1273] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1274] The following is further disclosed regarding the above embodiment.
[1275] (Claim 1)
[1276] a voice recognition means;
[1277] Image input means;
[1278] generating means;
[1279] Editing means;
[1280] A system including an output means.
[1281] (Claim 2)
[1282] 2. The system according to claim 1, wherein the input voice data is converted into text data by a voice recognition means.
[1283] (Claim 3)
[1284] 2. The system according to claim 1, wherein the image data captured by the image input means is used to automatically create presentation materials.
[1285] (Claim 4)
[1286] 2. The system according to claim 1, wherein the generating means automatically determines the structure of the presentation materials based on the text data and image data.
[1287] (Claim 5)
[1288] 2. The system according to claim 1, wherein the user can make minor edits to the presentation materials generated by the editing means.
[1289] (Claim 6)
[1290] 2. The system according to claim 1, wherein the output means saves or outputs the generated presentation materials in a predetermined format.
[1291] "Example 1"
[1292] (Claim 1)
[1293] a voice recognition means;
[1294] image processing means;
[1295] a text generation means;
[1296] a presentation generation means;
[1297] an editing interface means;
[1298] The system includes an output storage means.
[1299] (Claim 2)
[1300] 2. The system according to claim 1, wherein the input voice data is converted into text data by a voice recognition means.
[1301] (Claim 3)
[1302] 2. The system according to claim 1, wherein the image data captured by the image processing means is used to automatically create presentation materials.
[1303] (Claim 4)
[1304] 2. The system according to claim 1, wherein the presentation generating means automatically generates presentation materials by combining text data and image data.
[1305] (Claim 5)
[1306] 10. The system of claim 1, wherein the editing interface means allows a user to modify the presentation material.
[1307] (Claim 6)
[1308] 2. The system of claim 1, wherein the output saving means saves the final presentation materials in a specified format and generates a download link.
[1309] "Application Example 1"
[1310] (Claim 1)
[1311] a voice recognition means;
[1312] Image input means;
[1313] generating means;
[1314] Editing means;
[1315] An output means;
[1316] A visual output means for automatically generating product explanations and comparison presentations within the virtual store;
[1317] A system including an integrated means for creating presentations based on audio input and image data.
[1318] (Claim 2)
[1319] 2. The system according to claim 1, wherein the voice recognition means converts input voice data into text data and provides explanations of products in the virtual store.
[1320] (Claim 3)
[1321] 2. The system according to claim 1, wherein the image data captured by the image input means is used to automatically create presentation materials instantly within the virtual store.
[1322] "Example 2: Combining Emotion Engines"
[1323] (Claim 1)
[1324] a voice recognition means;
[1325] Image input means;
[1326] generating means;
[1327] Editing means;
[1328] A sentiment analysis means;
[1329] A system including an output means.
[1330] (Claim 2)
[1331] 2. The system according to claim 1, wherein the input voice data is converted into text data by a voice recognition means.
[1332] (Claim 3)
[1333] 2. The system according to claim 1, wherein the image data captured by the image input means is used to automatically create presentation materials.
[1334] (Claim 4)
[1335] 2. The system according to claim 1, wherein the emotion analysis means analyzes the emotional state of the user, and adjusts the content and tone of the presentation materials based on the analysis results.
[1336] (Claim 5)
[1337] 10. The system of claim 1, wherein the user can modify the generated presentation materials using editing means.
[1338] (Claim 6)
[1339] 2. The system according to claim 1, wherein the output means saves the final presentation materials in a predetermined format and provides the saved materials to the user.
[1340] "Application example 2 when combining emotion engines"
[1341] (Claim 1)
[1342] a voice recognition means;
[1343] Image input means;
[1344] A sentiment analysis means;
[1345] generating means;
[1346] Editing means;
[1347] A system including an output means.
[1348] (Claim 2)
[1349] 2. The system according to claim 1, wherein the voice recognition means converts input voice data into text data, and the emotion analysis means analyzes an emotional state from the voice data.
[1350] (Claim 3)
[1351] 2. The system according to claim 1, wherein presentation materials are automatically created using image data, text data, and emotional states captured by the image input means. [Explanation of symbols]
[1352] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a voice recognition means; Image input means; generating means; Editing means; A system including an output means.
2. 2. The system according to claim 1, wherein the voice recognition means converts input voice data into text data.
3. 2. The system according to claim 1, wherein the image data captured by the image input means is used to automatically create presentation materials.
4. 2. The system according to claim 1, wherein the generating means automatically determines the structure of the presentation material based on the text data and image data.
5. 2. The system according to claim 1, wherein the user can make minor edits to the presentation material generated by the editing means.
6. 2. The system according to claim 1, wherein the output means stores or outputs the generated presentation materials in a predetermined format.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A