System

The system addresses limitations in existing technologies by enabling real-time generation and display of interactive 3D holograms using voice input, facilitating versatile applications in entertainment, education, and business presentations.

JP2026018093APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119154
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing systems are limited to static presentations and lack flexibility for dynamic, interactive 3D stereoscopic image generation using voice input, primarily in specific fields and applications, and do not support multiple uses such as entertainment, education, and business presentations.

Method used

A system that includes voice input, communication, speech recognition, 3D image generation, hologram conversion, and display components, allowing real-time projection of interactive 3D holograms based on user voice input, compatible with various devices for diverse applications.

Benefits of technology

Enables users to experience advanced visual content in real-time across multiple fields, including entertainment, education, and business presentations, through seamless audio-to-hologram conversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018093000001_ABST
    Figure 2026018093000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: an input device for a user to input a voice; a communication device for the terminal 1 to transmit a voice datum to a server; a voice recognition device for the server to convert the voice datum into a text; a generation device for the server to generate a 3D stereoscopic image based on the text information; a conversion device for the server to convert the 3D stereoscopic image into a hologram datum; and a display device for the terminal 1 to receive and project the hologram datum in real time.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Recent technological advances have made it possible to combine voice recognition and hologram technologies, but previous systems have been limited to static presentations and limited interactive functions. Challenges remain in developing a system that generates and displays dynamic, interactive 3D stereoscopic images in real time using only the user's voice in the real world. Furthermore, existing technologies are limited to specific fields and applications and lack the flexibility to accommodate multiple uses. This invention aims to solve these challenges and provide a highly versatile system that can be used in a variety of fields, including entertainment, education, and business presentations. [Means for solving the problem]

[0005] The present invention provides a system that includes an input means for inputting a user's voice, a communication means for the terminal to transmit the voice data to a server, a voice recognition means for the server to convert the voice data into text, a generation means for the server to generate a 3D stereoscopic image based on the text information, a conversion means for the server to convert the 3D stereoscopic image into hologram data, and a display means for the terminal to receive the hologram data and project it in real time. Furthermore, because the generation means generates interactive 3D images based on text information and the display means is compatible with multiple support devices, this system can be used for a wide range of applications, including entertainment, education, and business presentations. This allows users to enjoy an advanced visual experience in real time in the real world.

[0006] "User" refers to a person who uses this system to input voice and experience hologram content.

[0007] "Speech recognition means" is a general term for software and hardware for converting voice data into text data.

[0008] A "terminal" refers to a device that has functions such as voice input and hologram display and communicates with a server.

[0009] "Communication means" refers to technology including interfaces and protocols for transmitting and receiving data from a terminal to a server or from a server to a terminal.

[0010] "Server" refers to the computing resources that perform voice recognition, text analysis, 3D image generation, and hologram data conversion.

[0011] "Generation means" refers to the algorithms and software used to generate 3D stereoscopic images based on text information.

[0012] The "conversion means" is a technology for converting the generated 3D stereoscopic image data into hologram data that can be visually confirmed by the user.

[0013] "Display means" is a general term for devices and software for displaying hologram data received by a terminal in real time.

[0014] "Support device" refers to hardware such as a projector or holographic display that supports the real-time projection of holograms. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] The system of the present invention generates a 3D stereoscopic image in real time based on the user's voice input and displays it as a hologram. Below, we will create a program for this system and explain its processing in detail.

[0037] overview

[0038] The system consists of the following main components:

[0039] 1. A voice input device (terminal) that inputs the user's voice

[0040] 2. A communication device (terminal) that transmits voice data to a server

[0041] 3. Speech recognition engine (server) that converts voice data into text

[0042] 4. Generative AI (server) that generates 3D images based on text information

[0043] 5. Hologram generation engine (server) that converts the generated 3D stereoscopic image into hologram data

[0044] 6. Support device (terminal) that receives and projects hologram data

[0045] Program processing flow

[0046] Voice input

[0047] 1. The user speaks into a voice input device, giving a command such as "Describe the dog."

[0048] 2. The device captures the user's voice and collects voice data in real time.

[0049] 3. Communication is established to send the audio data to the server.

[0050] Voice Recognition

[0051] 4. The server passes the received voice data to the voice recognition engine.

[0052] 5. The speech recognition engine analyzes the voice data and converts it into text information. For example, the text data generated is "Please describe the dog."

[0053] Text analysis and 3D image generation

[0054] 6. The server analyzes the generated text information and passes it to the generation AI to extract relevant content.

[0055] 7. Generative AI generates a 3D image based on the input text information. For example, it creates a 3D model of a dog based on the keyword "dog."

[0056] 8. The generated 3D stereoscopic image data is sent to the hologram generation engine.

[0057] Hologram Generation

[0058] 9. The server's hologram generation engine converts the 3D image data into hologram data.

[0059] 10. The converted hologram data is sent to the terminal.

[0060] Hologram display

[0061] 11. The terminal receives the hologram data and transmits it to the support device.

[0062] 12. The assistive device uses the holographic data to project a 3D image in real time.

[0063] Specific examples

[0064] Educational Demo Scenario

[0065] User: "Tell me about the planets in our solar system."

[0066] The device captures the audio and sends it to the server.

[0067] The server's speech recognition engine converts the received voice data into text such as "Tell me about the planets in the solar system."

[0068] The server's generation AI generates a series of 3D stereoscopic images (planet models) based on the data of each planet.

[0069] The server's hologram generation engine converts these 3D models into hologram data.

[0070] The terminal receives the hologram data and transmits it to the display device.

[0071] The assistive device displays a 3D model of the planets in the solar system in real time.

[0072] This system allows users to learn visually through holograms in real time using only voice input, and is expected to be used in a variety of situations, not just education, but also entertainment and business presentations.

[0073] The processing flow will be explained below.

[0074] Step 1:

[0075] A user speaks into a voice input device, for example, "Describe my dog."

[0076] Step 2:

[0077] The terminal captures the user's voice and converts it into digital voice data. The voice input device collects the voice and pre-processes it to convert it into a data format.

[0078] Step 3:

[0079] The device transmits the captured audio data to the server in real time. The audio data is sent to the server as a stream via the communication interface.

[0080] Step 4:

[0081] The server passes the received voice data to the speech recognition engine, which then decodes the voice data and provides it as input to the speech recognition engine.

[0082] Step 5:

[0083] The server's speech recognition engine converts the voice data into text data. Specifically, the speech recognition engine analyzes the voice and generates text data such as "Please describe the dog."

[0084] Step 6:

[0085] The server analyzes the text data obtained from the speech recognition engine and passes it to the generation AI to extract relevant content.

[0086] Step 7:

[0087] The server's generation AI generates a 3D image based on text data. For example, based on the data for "dog," it generates a 3D model of a dog.

[0088] Step 8:

[0089] The server sends the generated 3D model data to the hologram generation engine. The data from the generation AI is passed to the conversion engine, which prepares the hologram data for creation.

[0090] Step 9:

[0091] The server's hologram generation engine converts the 3D stereoscopic image data into hologram data, specifically, converting the 3D model into a format suitable for holographic display.

[0092] Step 10:

[0093] The server sends the converted hologram data to the terminal via a communication interface.

[0094] Step 11:

[0095] The terminal receives the hologram data, temporarily stores it in a buffer, retrieves the data through the receiving interface, and starts processing it immediately.

[0096] Step 12:

[0097] The terminal initiates the process for transmitting hologram data to an assistive device, such as a projector or hologram display.

[0098] Step 13:

[0099] The assistive device uses the received hologram data to project a 3D stereoscopic image in real time. The hologram projection device interprets the data and displays a realistic stereoscopic image in space.

[0100] Each step allows the user to seamlessly experience the audio-to-hologram conversion process.

[0101] Example 1

[0102] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0103] Many modern learning and presentation systems provide information through text or 2D images, but these do not adequately support visual comprehension. Furthermore, conventional voice input systems struggle to visually display the input content in real time, requiring numerous steps and devices. This complicates the user experience and makes intuitive operation difficult.

[0104] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0105] In this invention, the server includes a voice recognition unit that converts voice data into text, an analysis unit that analyzes the text information, a generation unit that generates a 3D stereoscopic image based on the text information, and a conversion unit that converts the 3D stereoscopic image into hologram data. This makes it possible to generate a 3D stereoscopic image in real time based on a user's voice input and display the image as a hologram.

[0106] The "input means" is a device that allows the user to input voice.

[0107] "Communication means" is a mechanism for transmitting voice data from a terminal to a server.

[0108] "Speech recognition means" refers to software or algorithms that convert voice data into text on the server.

[0109] "Analysis means" refers to the process or technology by which the server analyzes text information and extracts the necessary information.

[0110] The "generation means" is a technology that allows the server to generate a 3D stereoscopic image based on text information.

[0111] The "conversion means" is the technology that the server uses to convert the 3D stereoscopic image into hologram data.

[0112] "Display means" refers to a device or technology that allows a terminal to receive hologram data and project it in real time.

[0113] The "support device" is an external device for actually projecting and displaying hologram data.

[0114] This invention relates to a system that generates a 3D stereoscopic image in real time based on a user's voice input and displays it as a hologram. Specific embodiments of this system will be described below.

[0115] The system consists of the following main components:

[0116] 1. Audio input device (terminal)

[0117] 2. Communication Devices (Terminals)

[0118] 3. Speech recognition engine (server)

[0119] 4. Generative AI (server)

[0120] 5. Hologram generation engine (server)

[0121] 6. Projection Support Devices (Terminals)

[0122] Voice input

[0123] First, a user can operate the system by speaking into a voice input device. For example, the user can speak commands such as "Tell me about the planets in the solar system." The voice input device captures the voice and generates audio data. Specific voice input devices can be built-in microphones or dedicated voice input devices.

[0124] Sending audio data

[0125] The collected audio data is then transmitted to a server via the communication device, using standard wireless communication technologies such as Wi-Fi or Bluetooth.

[0126] Voice Recognition

[0127] The server then passes the received voice data to a voice recognition engine for analysis. An example of a voice recognition engine is the Google Cloud Speech-to-Text API. This engine analyzes phonemes and phonology to convert the voice data into text, and generates the corresponding text. For example, a command such as "Tell me about the planets in the solar system" is generated as text data.

[0128] Text analysis and 3D image generation

[0129] Once the text data is generated, an analysis tool on the server analyzes the text data and extracts the necessary information. Natural language processing (NLP) technology is used as the analysis tool. Based on the results of this analysis, the server's generative AI (such as OpenAI's GPT-3) creates data for generating 3D stereoscopic images. For example, a series of 3D stereoscopic images (planet models) can be generated based on data on the planets in the solar system.

[0130] Hologram Generation

[0131] The generated 3D stereoscopic image data is passed to the server's hologram generation engine, which converts it into hologram data. The hologram generation engine can use, for example, Looking Glass Factory's hologram generation technology. This engine converts the 3D model data into images from multiple viewpoints and combines them to create hologram data.

[0132] Hologram display

[0133] The hologram data is sent from the server to the terminal and displayed. The terminal receives this data and passes it to a projection assistance device, specifically a hologram projector, which uses light interference patterns to display 3D images in space.

[0134] Specific examples

[0135] For example, consider the following scenario for an educational system:

[0136] User: "Tell me about the planets in our solar system."

[0137] The device captures the audio and sends it to the server. At this stage, the built-in microphone picks up the audio signal and sends it to the server as digital data.

[0138] The server's speech recognition engine converts the text to "Tell me about the planets in the solar system," using feature extraction and statistical models of the speech signal.

[0139] The server's NLP engine parses this text and understands it as a specific information request.

[0140] The server sends the generated AI a prompt saying, "Provide information about planets in the solar system."

[0141] Generative AI generates a series of 3D stereoscopic images based on data for each planet.

[0142] The server's hologram generation engine converts these 3D models into hologram data.

[0143] The terminal receives the hologram data and transmits it to the display device, using Wi-Fi or a wired connection.

[0144] The assistive device displays a real-time stereoscopic model of the planets in the solar system. The projector uses light patterns to create 3D images in space that can be seen from multiple perspectives.

[0145] In this way, users can visually study detailed 3D holograms in real time through voice input.

[0146] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0147] Step 1:

[0148] The user inputs commands to the system by speaking into a voice input device. For example, the user might say, "Please describe the dog." The input is the user's voice data.

[0149] Step 2:

[0150] The terminal captures the user's voice and collects the voice data. The microphone inside the device converts the voice signal into digital data and processes the data in real time. The input is the user's voice data, and the output is the digitized voice data.

[0151] Step 3:

[0152] The voice data collected by the device is sent to a server via a communication device. Specifically, the data is transferred using wireless communication protocols such as Wi-Fi or Bluetooth. The input is digitized voice data, and the output is the voice data sent to the server.

[0153] Step 4:

[0154] The server passes the received voice data to a voice recognition engine (voice recognition means) for analysis. The voice recognition engine (for example, general voice recognition software) analyzes the voice data and converts it into text information. The input is voice data and the output is text information.

[0155] Step 5:

[0156] The server's analysis means analyzes the generated text information and extracts the necessary information. Natural language processing (NLP) technology is used for this analysis, and specific keywords and command content are extracted from the analysis results. The input is text information, and the output is the extracted keywords and command data.

[0157] Step 6:

[0158] The server passes a prompt to the generation AI based on the extracted command data. For example, a prompt such as "Please describe the dog" is input to the generation AI. The input is the command data, and the output is the prompt.

[0159] Step 7:

[0160] Based on the prompt received, the generation AI generates relevant 3D stereoscopic image data. For example, the keyword "dog" generates 3D model data of a dog. The input is the prompt, and the output is 3D stereoscopic image data.

[0161] Step 8:

[0162] The server's hologram generation engine (conversion means) converts the generated 3D stereoscopic image data into hologram data. Specifically, it uses 3D model data to generate images from multiple viewpoints and then combines them to create hologram data. The input is 3D stereoscopic image data, and the output is hologram data.

[0163] Step 9:

[0164] The server transmits the converted hologram data to the terminal via Wi-Fi or a wired connection. The input is the hologram data, and the output is the hologram data transmitted to the terminal.

[0165] Step 10:

[0166] The terminal receives the hologram data and passes it to the support device. For example, it transmits the data to a hologram projector. The input is the hologram data, and the output is the hologram data passed to the support device.

[0167] Step 11:

[0168] The assistive device uses the hologram data to project a 3D image in real time. For example, a hologram projector displays a 3D image in space based on the data it receives. The input is hologram data, and the output is a 3D hologram displayed in space.

[0169] (Application example 1)

[0170] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0171] Currently, physical stores have limited means of providing detailed product information, making it difficult for customers to fully understand product features. Furthermore, there is a lack of efficient means for providing visual product information to multiple customers simultaneously. Therefore, there is a demand for technology that can easily display product information in real time using voice input to improve the user experience and stimulate purchasing motivation.

[0172] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0173] In this invention, the server includes an input means for a user to input voice, a communication means for a terminal to transmit voice data to the server, a voice recognition means for the server to convert the voice data into text, a generation means for the server to generate a 3D stereoscopic image using a generative artificial intelligence model based on the text information, a conversion means for the server to convert the 3D stereoscopic image into hologram data, and a display means for the terminal to receive the hologram data and project it in real time. This makes it possible to display interactive 3D holograms in physical stores that allow customers to intuitively understand detailed product information through voice commands.

[0174] "Input means for inputting voice" refers to a device or technology that allows a user to give instructions to a system via voice.

[0175] "Communication means for transmitting audio data to a server" refers to a method or device for sending audio data to a server via the Internet or other communication network.

[0176] "Speech recognition means for converting voice data into text" refers to the algorithms or software used to analyze voice data and convert it into text data.

[0177] "Generative AI model" refers to an AI model or system used to generate 3D stereoscopic images based on text information.

[0178] "Means for converting a 3D stereoscopic image into hologram data" refers to an algorithm or system for converting a 3D stereoscopic image into a data format that can be projected as a hologram.

[0179] "Display means for receiving holographic data and projecting it in real time" refers to a device or technology for visually projecting the received holographic data in real time.

[0180] A "prompt" is a piece of text used to instruct a generative AI model to generate a particular output.

[0181] "Display device" is a general term for devices and hardware for visually displaying generated hologram data.

[0182] A "physical store" refers to a store located in a physical location where customers can visit in person to view and purchase products.

[0183] The present invention relates to a system that displays detailed product information in a physical store as a 3D hologram in real time. This system allows a user to input voice, and then projects a 3D stereoscopic image generated based on the voice as a hologram. Specific embodiments of this system are described below.

[0184] System configuration

[0185] The system consists of the following main components:

[0186] 1. Voice input means: A microphone for users to input voice.

[0187] 2. Communication method: The method by which the voice data is transmitted to the server via the Internet.

[0188] 3. Speech recognition means: A speech recognition engine (e.g., Google Speech-to-Text) that the server uses to convert voice data into text.

[0189] 4. Generation: The process in which the server uses a generative AI model (e.g., OpenAI's GPT-3) to generate a 3D stereoscopic image from text information.

[0190] 5. Conversion method: The algorithm used by the server to convert the 3D stereoscopic image into hologram data (e.g., Unity's hologram generation plug-in).

[0191] 6. Display means: A device that projects holographic data in real time using a smartphone or a holographic display (e.g., Looking Glass Factory's holographic display).

[0192] Processing flow

[0193] 1. Voice input: The user speaks into the smartphone, for example, "Please tell me the features of this new smartphone."

[0194] 2. Sending audio data: The smartphone sends the audio data to the server via an HTTP request over the internet.

[0195] 3. Speech recognition: The server passes the received voice data to a speech recognition engine and converts it into text. For example, the generated text is "Please tell me the features of this new smartphone."

[0196] 4. Text analysis and 3D image generation: The server uses generative AI models to generate relevant 3D images from text information. For example, a 3D model is generated based on the detailed features and specifications of a smartphone.

[0197] 5. Hologram generation: The server's hologram generation engine analyzes the 3D stereoscopic image and converts it into data that can be displayed as a hologram.

[0198] 6. Hologram display: A smartphone receives hologram data and projects it in real time on a built-in hologram display, etc. This device can display interactive 3D images in real time.

[0199] Specific example explanation

[0200] Example 1:

[0201] When a user speaks to a smartphone sales counter in a physical store and says, "What are the features of this new smartphone?", the system instantly projects a 3D model of the smartphone and its features as a hologram, allowing the user to understand the details of the product without having to touch it.

[0202] Example prompt sentence:

[0203] User dictates: "What are the features of this new phone?"

[0204] Server prompt: "Please describe the features of the new smartphone XYZ. Generate a 3D model of it."

[0205] This system will significantly improve the customer experience in physical stores, increasing product understanding and purchasing intent.

[0206] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0207] Step 1:

[0208] A user uses the smartphone's voice input means to ask a product-related question. Specifically, the user says, "What are the features of this new smartphone?" The input is the user's voice, and the output is the voice data captured by the smartphone's microphone.

[0209] Step 2:

[0210] The device sends the captured audio data to the server via a communication means. The audio data is sent to the server as an HTTP request over the Internet. The input is the audio data captured by the smartphone's microphone, and the output is the audio data sent to the server.

[0211] Step 3:

[0212] The server receives the voice data and passes it to a voice recognition engine to convert it into text. Specifically, the server analyzes the received voice data and generates text data such as "Please tell me the features of this new smartphone." The input is voice data and the output is text data.

[0213] Step 4:

[0214] The server passes the generated text data to a generative AI model, which then generates a 3D stereoscopic image based on the text information. Specifically, the generative AI model generates a 3D model of a smartphone based on the prompt, "Features of the new smartphone." The input is text data, and the output is 3D stereoscopic image data.

[0215] Step 5:

[0216] The server passes the generated 3D stereoscopic image data to an algorithm for converting it into hologram data. Here, the server's hologram generation engine converts the 3D stereoscopic image into a data format that can be displayed as a hologram. The input is 3D stereoscopic image data, and the output is hologram data.

[0217] Step 6:

[0218] The device receives the hologram data from the server and projects it in real time using the built-in display means. Specifically, a smartphone or holographic display visually displays the hologram data. The input is the hologram data, and the output is the displayed hologram.

[0219] This will allow users to view interactive 3D holograms in physical stores that provide intuitive understanding of product details.

[0220] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0221] The system of the present invention generates 3D stereoscopic images in real time based on the user's voice input and displays them as holograms. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide a more interactive and personalized experience. Below, we will create a program for this system and explain its processing in detail.

[0222] overview

[0223] The system consists of the following main components:

[0224] 1. A voice input device (terminal) that inputs the user's voice

[0225] 2. A communication device (terminal) that transmits voice data to a server

[0226] 3. Speech recognition engine (server) that converts voice data into text

[0227] 4. Generative AI (server) that generates 3D images based on text information

[0228] 5. Hologram generation engine (server) that converts the generated 3D stereoscopic image into hologram data

[0229] 6. Support device (terminal) that receives and projects hologram data

[0230] 7. Emotion engine (server) that recognizes user emotions

[0231] Program processing flow

[0232] Voice input

[0233] 1. The user speaks into a voice input device, for example, "Please describe my dog."

[0234] 2. The terminal captures the user's voice and converts it into digital voice data. The voice input device collects the voice and performs pre-processing to convert it into a data format.

[0235] 3. Communication is established to send the audio data to the server.

[0236] Voice Recognition

[0237] 4. The server passes the received voice data to the speech recognition engine, which decodes the voice data and provides it as input to the speech recognition engine.

[0238] 5. The speech recognition engine converts the voice data into text data. For example, the text data generated is "Please describe the dog."

[0239] emotion recognition

[0240] 6. The server passes the voice data to the emotion engine, which analyzes the user's emotional state. The emotion engine analyzes the tone and tempo of the voice to identify the user's emotion (happiness, sadness, surprise, etc.).

[0241] 7. Provide the emotion engine's analysis results (user's emotional state) to the generation AI.

[0242] Text analysis and 3D image generation

[0243] 8. The server analyzes the generated text information and sentiment analysis results and passes them to the generation AI to extract relevant content.

[0244] 9. The generation AI generates a 3D image based on the input text and emotion information. For example, it generates a 3D model of a dog based on the keyword "dog," and adds movements and expressions that correspond to the user's emotions.

[0245] 10. The generated 3D stereoscopic image data is sent to the hologram generation engine.

[0246] Hologram Generation

[0247] 11. The server's hologram generation engine converts the 3D stereoscopic image data into hologram data, specifically converting the 3D model into a format suitable for holographic display.

[0248] 12. The converted hologram data is sent to the terminal.

[0249] Hologram display

[0250] 13. The terminal receives the hologram data and temporarily stores it in a buffer. It then retrieves the data through the receiving interface and starts processing it immediately.

[0251] 14. The terminal starts the process of transmitting hologram data to the assistive device, such as a projector or hologram display.

[0252] 15. The assistive device uses the received hologram data to project a 3D stereoscopic image in real time. The hologram projection device interprets the data and displays a realistic stereoscopic image in space.

[0253] Specific examples

[0254] Educational Demo Scenario

[0255] User: "Tell me about the planets in our solar system."

[0256] The device captures the user's voice and sends it to the server.

[0257] The server's speech recognition engine converts the received voice data into text such as "Tell me about the planets in the solar system."

[0258] The server's emotion engine analyzes the user's emotions from the voice data, and recognizes the emotion of joy.

[0259] The server's generation AI generates a series of 3D stereoscopic images (planet models) based on the data of each planet, adding bright colors and movement according to the user's pleasure.

[0260] The server's hologram generation engine converts these 3D models into hologram data.

[0261] The terminal receives the hologram data and transmits it to the display device.

[0262] The assistive device displays a 3D model of the planets in the solar system in real time.

[0263] The system allows users to enjoy a richer, more personalized holographic experience in real time, utilizing both voice input and emotion analysis.

[0264] The processing flow will be explained below.

[0265] Step 1:

[0266] A user speaks into a voice input device, for example, "Describe my dog."

[0267] Step 2:

[0268] The terminal captures the user's voice and converts it into digital voice data. The voice input device collects the voice and pre-processes it to convert it into a data format.

[0269] Step 3:

[0270] The device transmits the captured audio data to the server in real time. The audio data is sent to the server as a stream via the communication interface.

[0271] Step 4:

[0272] The server passes the received voice data to the speech recognition engine, which then decodes the voice data and provides it as input to the speech recognition engine.

[0273] Step 5:

[0274] The server's speech recognition engine converts the voice data into text data. Specifically, the speech recognition engine analyzes the voice and generates text data such as "Please describe the dog."

[0275] Step 6:

[0276] The server passes the voice data to the emotion engine, which analyzes the user's emotional state. The emotion engine analyzes the tone and tempo of the voice to identify the user's emotion (happiness, sadness, surprise, etc.).

[0277] Step 7:

[0278] The server's emotion engine provides the analysis results (user's emotional state) to the generation AI, which then performs the following processes based on the received emotional information.

[0279] Step 8:

[0280] The server passes the generated text information and emotion analysis results to the generation AI, which extracts relevant content. The generation AI then generates a 3D image based on the input text information and emotion information.

[0281] Step 9:

[0282] The server's AI generates 3D model data for a dog based on a keyword, such as "dog," and adds movements and facial expressions that correspond to the user's emotions.

[0283] Step 10:

[0284] The server sends the generated 3D model data to the hologram generation engine. The data from the generation AI is passed to the conversion engine, which prepares the hologram data for creation.

[0285] Step 11:

[0286] The server's hologram generation engine converts the 3D stereoscopic image data into hologram data, specifically, converting the 3D model into a format suitable for holographic display.

[0287] Step 12:

[0288] The server sends the converted hologram data to the terminal via a communication interface.

[0289] Step 13:

[0290] The terminal receives the hologram data, temporarily stores it in a buffer, retrieves the data through the receiving interface, and starts processing it immediately.

[0291] Step 14:

[0292] The terminal initiates the process for transmitting hologram data to an assistive device, such as a projector or hologram display.

[0293] Step 15:

[0294] The assistive device uses the received hologram data to project a 3D stereoscopic image in real time. The hologram projection device interprets the data and displays a realistic stereoscopic image in space.

[0295] Specific examples

[0296] Educational Demo Scenario

[0297] Step 1:

[0298] A user speaks into a voice input device, "Tell me about the planets in the solar system."

[0299] Step 2:

[0300] The device captures the user's voice and converts it into digital voice data.

[0301] Step 3:

[0302] The device transmits the captured audio data to the server in real time.

[0303] Step 4:

[0304] The server passes the received voice data to the voice recognition engine.

[0305] Step 5:

[0306] The server's voice recognition engine converts the voice data into text data such as "Tell me about the planets in the solar system."

[0307] Step 6:

[0308] The server passes the voice data to the emotion engine, which analyzes the user's emotional state (interest and pleasure).

[0309] Step 7:

[0310] The server's emotion engine provides the analysis results to the generation AI.

[0311] Step 8:

[0312] The server passes the generated text information and sentiment analysis results to the generation AI, which extracts relevant content.

[0313] Step 9:

[0314] The server's generation AI generates a series of 3D stereoscopic images (planet models) based on data for each planet, adding bright colors and movement according to the user's emotions.

[0315] Step 10:

[0316] The server sends the generated 3D model data to the hologram generation engine.

[0317] Step 11:

[0318] The server's hologram generation engine converts the generated 3D model data into hologram data.

[0319] Step 12:

[0320] The server transmits the converted hologram data to the terminal.

[0321] Step 13:

[0322] The terminal receives the hologram data and temporarily stores it in a buffer.

[0323] Step 14:

[0324] The terminal initiates the process for transmitting the hologram data to the assistive device.

[0325] Step 15:

[0326] The assistive device displays a 3D model of the planets in the solar system in real time.

[0327] Example 2

[0328] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0329] Conventional voice input systems simply recognize a user's voice, convert it into text, and display a static 3D stereoscopic image based on that text. This limited the user experience, making it difficult to provide an interactive and personalized experience. Furthermore, the lack of technology to analyze user emotions and generate content based on those emotions prevented users from achieving high levels of satisfaction.

[0330] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0331] In this invention, the server includes means for converting voice data into text, means for generating a 3D stereoscopic image based on the text information, means for recognizing a user's emotion from the voice data, and means for adjusting the movements and facial expressions of the generated 3D stereoscopic image based on the user's emotion information, thereby making it possible to provide an interactive and personalized hologram experience in real time according to the user's emotion.

[0332] "Input means" refers to a device that allows a user to input voice.

[0333] "Communication means" refers to a network interface that enables a terminal to transmit audio data to a server.

[0334] "Speech recognition means" refers to the algorithm or engine that converts the voice data received by the server into text data.

[0335] "Generation means" refers to the algorithms or programs that the server uses to generate 3D stereoscopic images based on text information.

[0336] "Conversion means" refers to the algorithm or engine that the server uses to convert 3D stereoscopic images into hologram data.

[0337] "Display means" refers to a device for projecting hologram data received by the terminal in real time.

[0338] "Recognition means" refers to the algorithms or engines that the server uses to analyze the user's emotions from voice data.

[0339] "Adjustment means" refers to an algorithm or program that the generation means uses to adjust the movements and facial expressions of the 3D stereoscopic image based on the user's emotional information.

[0340] "Support equipment" refers to an external device for projecting a hologram.

[0341] This invention provides a system that generates 3D stereoscopic images in real time based on the user's voice input and displays them as holograms, and also recognizes the user's emotions and provides a personalized experience based on them.

[0342] Hardware and software used

[0343] Hardware:

[0344] Audio input devices: microphones, smartphones, etc.

[0345] Communication device: smartphone, tablet, or computer.

[0346] Server: High-performance cloud or on-premise servers.

[0347] Assistive devices: Hologram projector, VR / AR headset (e.g. Microsoft HoloLens).

[0348] software:

[0349] Speech recognition engine: Google Speech-to-Text API, Azure Speech Service, etc.

[0350] Emotion recognition engine: IBM Watson Tone Analyzer.

[0351] Generative AI: OpenAI GPT-3 or similar advanced language model.

[0352] Hologram generation engine: Unity or Unreal Engine.

[0353] Program processing explanation

[0354] 1. Voice input

[0355] A user speaks into a voice input device, for example, "Please describe my dog."

[0356] 2. Communications

[0357] In order for the terminal to transmit the audio data to the server, the digital audio data is divided into packets and transmitted to the server via secure communication.

[0358] 3. Voice Recognition

[0359] The server passes the received voice data to a voice recognition engine, which converts the voice into text data.

[0360] 4. Emotion recognition

[0361] The server passes the voice data to an emotion recognition engine, which analyzes the user's emotional state. The emotion engine analyzes the tone and tempo of the voice to identify the user's emotion (happiness, sadness, surprise, etc.).

[0362] 5. Text Analysis and 3D Image Generation

[0363] The server inputs the generated text information and emotion analysis results into a generative AI model to generate appropriate 3D stereoscopic image data. For example, based on the text "dog" and emotion information, a 3D model of a dog is generated and its movements and facial expressions are adjusted according to the emotion.

[0364] 6. Hologram Generation

[0365] The server's hologram generation engine converts the 3D stereoscopic image data into hologram data. Specifically, it uses Unity or Unreal Engine to encode the 3D model into a format suitable for holograms.

[0366] 7. Data Transmission

[0367] The server transmits the converted hologram data to the terminal.

[0368] 8. Hologram display

[0369] The terminal receives the hologram data and transmits it to the support device, which then displays the hologram in real time.

[0370] Specific examples

[0371] Educational Demo Scenario

[0372] User: "Tell me about the planets in our solar system."

[0373] The device captures the user's voice, converts it into digital audio data, and sends it to the server.

[0374] The server's speech recognition engine converts the received voice data into text such as "Tell me about the planets in the solar system."

[0375] The server's emotion recognition engine analyzes the user's emotions from the voice and recognizes the emotion of joy.

[0376] The server's generative AI model generates a series of 3D stereoscopic images (planet models) based on data for each planet, adding bright colors and movement depending on the user's emotions.

[0377] The server's hologram generation engine converts these 3D models into hologram data.

[0378] The terminal receives the hologram data and transmits it to the support device.

[0379] The support device displays a 3D model of the planets in the solar system in real time.

[0380] Example prompt: "Tell me about the planets in the solar system."

[0381] The system allows users to enjoy a richer, more personalized holographic experience in real time, utilizing both voice input and emotion analysis.

[0382] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0383] Step 1:

[0384] The user speaks into the voice input device, saying, "Please describe my dog," and the voice input device captures the voice and converts it into digital voice data.

[0385] Input: User spoken words.

[0386] Output: Digital audio data.

[0387] What happens: An audio input device (e.g., a smartphone microphone) converts an analog audio signal into a digital signal, which is sampled at 44.1 kHz and recorded in PCM format.

[0388] Step 2:

[0389] The terminal sends voice data to the server using a communication method. The voice data is divided into packets and sent to the server using secure communication (SSL / TLS).

[0390] Input: Digital audio data.

[0391] Output: The audio data sent to the server.

[0392] What it does: The device splits the audio data into packets of 1024 bytes each and sends the series of packets to the server over an SSL / TLS connection.

[0393] Step 3:

[0394] The server receives the voice data and passes it to the voice recognition engine, which converts the received voice data into text data.

[0395] Input: Audio data sent to the server.

[0396] Output: Text data.

[0397] Specific operation: The server temporarily stores the voice data in storage and calls the Google Speech-to-Text API to convert the voice data into text data. For example, the text generated is "Please describe the dog."

[0398] Step 4:

[0399] The server passes the voice data to an emotion recognition engine, which analyzes the user's emotional state. The emotion recognition engine analyzes the tone and tempo of the voice to recognize the user's emotions.

[0400] Input: Audio data.

[0401] Output: Emotion analysis results.

[0402] How it works: The server calls IBM Watson Tone Analyzer, which analyzes the tone, pitch, speed, etc. of the voice data to identify emotions. For example, the voice may be judged to be "joyful," and the joy score may be high.

[0403] Step 5:

[0404] The server inputs text and emotion information into the generative AI model and obtains the generated 3D stereoscopic image data. The generative AI generates a 3D model based on the text and adds movements and facial expressions according to the emotion information.

[0405] Input: Text data, sentiment analysis results.

[0406] Output: 3D stereoscopic image data.

[0407] Specific operation: The server calls OpenAI GPT-3, inputs the text prompt "Please describe the dog" and the emotion "joy" as input, and generates 3D stereoscopic image data. For example, a joyful motion of the dog jumping up and down is added.

[0408] Step 6:

[0409] The server converts the 3D stereoscopic image into hologram data, which is then converted into a format suitable for holographic projection.

[0410] Input: 3D stereoscopic image data.

[0411] Output: Hologram data.

[0412] How it works: The server uses Unity or Unreal Engine to convert 3D stereoscopic image data into a volumetric hologram format, for example encoding a 3D model of a dog into a holographic format.

[0413] Step 7:

[0414] The server transmits hologram data to the terminal, which receives the data and transmits it to the support device.

[0415] Input: Hologram data.

[0416] Output: Hologram data sent to the device.

[0417] Specific operation: The server packetizes the hologram data and sends it to the device via an SSL / TLS connection. The device temporarily stores the received data in a buffer.

[0418] Step 8:

[0419] The device transmits the hologram data to the support device, which displays the hologram in real time. The support device interprets the hologram data and projects it into real space.

[0420] Input: Hologram data.

[0421] Output: The projected hologram.

[0422] How it works: The device uses Wi-Fi 6 (IEEE 802.11ax) to transmit hologram data to an assistive device (e.g., Microsoft HoloLens). The assistive device interprets the data and projects a real-time 3D hologram of a bouncing dog into space.

[0423] (Application example 2)

[0424] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0425] In today's brick-and-mortar stores, customers are limited to a single explanation method and fixed information when receiving product explanations and store guidance. It is also difficult to provide a personalized experience tailored to the customer's emotions and state, making it difficult to provide effective customer service. With such limited information provided, it is difficult to increase customer satisfaction and purchasing motivation.

[0426] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0427] In this invention, the server includes an emotion recognition means for analyzing voice data to recognize a user's emotion, a generation means for personalizing a 3D stereoscopic image based on the user's emotion, and a generation means for generating an interactive 3D stereoscopic image, thereby enabling a richer and more personalized hologram experience based on the user's voice input and emotion analysis.

[0428] The "voice input means" is a device that allows the user to input voice.

[0429] "Communication means" is a function that enables a terminal to transmit voice data to a server.

[0430] "Speech recognition means" is a function by which the server converts voice data into text.

[0431] The "generation means" is a function that allows the server to generate a 3D stereoscopic image based on text information.

[0432] The "conversion means" is a function by which the server converts a 3D stereoscopic image into hologram data.

[0433] "Emotion recognition means" is a function in which the server analyzes voice data and recognizes the user's emotions.

[0434] The "means for personalizing" is a function by which the generating means personalizes the 3D stereoscopic image based on the user's emotions.

[0435] "Display means" refers to the function of the terminal to receive hologram data and project it in real time.

[0436] As an embodiment of this invention, we will describe an interactive product explanation system for use in a brick-and-mortar store. This system provides a personalized experience by generating 3D stereoscopic images in real time based on the user's voice input and emotion analysis, and displaying the images as holograms.

[0437] System configuration

[0438] 1. Voice input means: The user uses a device to input voice. This can include a microphone on a smartphone or smart glasses.

[0439] 2. Communication means: The device has the ability to send captured audio data to the server.

[0440] 3. Speech recognition: The server uses a speech recognition engine to convert the voice data into text, for example, Google Cloud Speech-to-Text API.

[0441] 4. Emotion recognition: The server analyzes the voice data and uses an emotion engine to recognize the user's emotions, for example, IBM Watson Tone Analyzer.

[0442] 5. Generation method: The server uses a generative AI model, such as OpenAI GPT-4, to generate a 3D stereoscopic image based on text information and emotions.

[0443] 6. Conversion means: The server uses a hologram generation engine to convert the generated 3D stereoscopic image into hologram data.

[0444] 7. Display means: Includes a support device for the terminal to receive holographic data and project it in real time, such as a holographic display from Looking Glass Factory.

[0445] Program processing

[0446] The system does the following:

[0447] First, the user uses the voice input means to speak about the product or information they want to hear explained about. For example, they might say, "Please explain the features of this product." At that time, the voice input means captures the user's voice and converts it into digital voice data. This data is then sent to the server via the communication means.

[0448] The voice data that arrives at the server is converted into text data using a voice recognition means. In parallel, an emotion recognition means analyzes the voice data and recognizes the user's emotions. The user's emotional state (e.g., curiosity or surprise) is used to further refine the interaction.

[0449] Next, the text information and emotion information are passed to a generation means, which uses a generative AI model to generate a 3D stereoscopic image. The generated 3D stereoscopic image is personalized by applying movements and facial expressions according to the user's emotions. The 3D stereoscopic image is then converted into hologram data by a conversion means.

[0450] Finally, the hologram data is transmitted to the support device through the display means and projected in real time in the physical store, allowing users to visualize the products and information as interactive holograms right in front of their eyes.

[0451] Specific examples

[0452] For example, suppose a user uses a voice input method in a store to say, "Please explain the features of this product." The voice is immediately sent to the server, where speech recognition and emotion analysis are performed. As a result, the text data "Please explain the features of this product" and the user's emotional state (curiosity) are acquired. Based on the text and emotion information, the generative AI model generates an easy-to-understand and interesting 3D stereoscopic image corresponding to the "curiosity." For example, this could include an animation or exploded view showing the internal structure of the product.

[0453] The generated 3D stereoscopic image is converted into hologram data and projected as a hologram in real time through a display device. Users can visually observe the hologram and manipulate it interactively.

[0454] Example prompt sentence:

[0455] "Please explain. Analyze the user's emotions and generate a 3D model based on those emotions. Now, the user said, 'Explain the features of this product.' The user's emotion is 'Curiosity.'"

[0456] This allows users to enjoy a more personalized experience in physical stores.

[0457] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0458] Program processing

[0459] Step 1:

[0460] The user uses a voice input means to talk about the product or information they want to hear explained. For example, they might say, "Please explain the features of this product." In this case, the input is the user's voice data. This voice data is captured by the microphone of a smartphone or smart glasses and converted into digital voice data. The output is digitized voice data.

[0461] Step 2:

[0462] The terminal transmits the captured voice data to the server via a communication means. At this time, the input is the captured digital voice data. The server receives this data and prepares to pass it to the voice recognition engine. The output is the result of transmitting the voice data to the server.

[0463] Step 3:

[0464] The server converts the voice data into text data using a speech recognition tool. The input is digital voice data. A speech recognition engine (e.g., Google Cloud Speech-to-Text) analyzes the data and generates text data such as "Describe the features of this product." The output is text data.

[0465] Step 4:

[0466] At the same time, the server passes the voice data to an emotion recognition means to analyze the user's emotions. The input is digital voice data. The emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the tone and tempo of the voice to identify the user's emotional state (e.g., data expressing curiosity or surprise). The output is the user's emotional data.

[0467] Step 5:

[0468] The server passes the generated text information and emotion analysis results to the generation means. At this time, the inputs are text data and emotion data. The generative AI model (e.g., OpenAI GPT-4) generates a 3D stereoscopic image based on the text and emotion information, corresponding to the user's emotion. The output is 3D stereoscopic image data.

[0469] Step 6:

[0470] The server passes the generated 3D stereoscopic image data to a conversion means, which converts it into hologram data. At this time, the input is the 3D stereoscopic image data. The hologram generation engine converts the 3D model data into a format suitable for hologram display. The output is hologram data.

[0471] Step 7:

[0472] The server sends the converted hologram data to the terminal. At this time, the input is the hologram data. The terminal receives the data through the receiving interface and temporarily stores it in a buffer. The output is the data storage result on the terminal.

[0473] Step 8:

[0474] The terminal transmits hologram data to the assistive device, which projects it in real time. The input is the hologram data. The assistive device (e.g., Looking Glass Factory's holographic display) interprets the received data and projects a realistic 3D image in space. The output is the user's visual holographic experience.

[0475] Specific operation example

[0476] For example, a user might say, "Explain the features of this product" using a voice input method while in a store. The voice is immediately sent to the server, where speech recognition and emotion analysis are performed. As a result, the text data "Explain the features of this product" and the user's emotional state (curiosity) are acquired. Based on the text and emotion information, a generative AI model generates an easy-to-understand and interesting 3D stereoscopic image that corresponds to the "curiosity." For example, this could include animations or exploded views that show the internal structure of the product.

[0477] Example prompt sentence:

[0478] "Please explain. Analyze the user's emotions and generate a 3D model based on those emotions. Now, the user said, 'Explain the features of this product.' The user's emotion is 'Curiosity.'"

[0479] In this way, a personalized and interactive hologram experience can be provided based on the user's voice input and emotion analysis.

[0480] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0481] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0482] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0483] [Second embodiment]

[0484] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0485] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0486] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0487] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0488] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0489] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0490] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0491] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0492] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0493] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0494] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0495] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0496] The system of the present invention generates a 3D stereoscopic image in real time based on the user's voice input and displays it as a hologram. Below, we will create a program for this system and explain its processing in detail.

[0497] overview

[0498] The system consists of the following main components:

[0499] 1. A voice input device (terminal) that inputs the user's voice

[0500] 2. A communication device (terminal) that transmits voice data to a server

[0501] 3. Speech recognition engine (server) that converts voice data into text

[0502] 4. Generative AI (server) that generates 3D images based on text information

[0503] 5. Hologram generation engine (server) that converts the generated 3D stereoscopic image into hologram data

[0504] 6. Support device (terminal) that receives and projects hologram data

[0505] Program processing flow

[0506] Voice input

[0507] 1. The user speaks into a voice input device, giving a command such as "Describe the dog."

[0508] 2. The device captures the user's voice and collects voice data in real time.

[0509] 3. Communication is established to send the audio data to the server.

[0510] Voice Recognition

[0511] 4. The server passes the received voice data to the voice recognition engine.

[0512] 5. The speech recognition engine analyzes the voice data and converts it into text information. For example, the text data generated is "Please describe the dog."

[0513] Text analysis and 3D image generation

[0514] 6. The server analyzes the generated text information and passes it to the generation AI to extract relevant content.

[0515] 7. Generative AI generates a 3D image based on the input text information. For example, it creates a 3D model of a dog based on the keyword "dog."

[0516] 8. The generated 3D stereoscopic image data is sent to the hologram generation engine.

[0517] Hologram Generation

[0518] 9. The server's hologram generation engine converts the 3D image data into hologram data.

[0519] 10. The converted hologram data is sent to the terminal.

[0520] Hologram display

[0521] 11. The terminal receives the hologram data and transmits it to the support device.

[0522] 12. The assistive device uses the holographic data to project a 3D image in real time.

[0523] Specific examples

[0524] Educational Demo Scenario

[0525] User: "Tell me about the planets in our solar system."

[0526] The device captures the audio and sends it to the server.

[0527] The server's speech recognition engine converts the received voice data into text such as "Tell me about the planets in the solar system."

[0528] The server's generation AI generates a series of 3D stereoscopic images (planet models) based on the data of each planet.

[0529] The server's hologram generation engine converts these 3D models into hologram data.

[0530] The terminal receives the hologram data and transmits it to the display device.

[0531] The assistive device displays a 3D model of the planets in the solar system in real time.

[0532] This system allows users to learn visually through holograms in real time using only voice input, and is expected to be used in a variety of situations, not just education, but also entertainment and business presentations.

[0533] The processing flow will be explained below.

[0534] Step 1:

[0535] A user speaks into a voice input device, for example, "Describe my dog."

[0536] Step 2:

[0537] The terminal captures the user's voice and converts it into digital voice data. The voice input device collects the voice and pre-processes it to convert it into a data format.

[0538] Step 3:

[0539] The device transmits the captured audio data to the server in real time. The audio data is sent to the server as a stream via the communication interface.

[0540] Step 4:

[0541] The server passes the received voice data to the speech recognition engine, which then decodes the voice data and provides it as input to the speech recognition engine.

[0542] Step 5:

[0543] The server's speech recognition engine converts the voice data into text data. Specifically, the speech recognition engine analyzes the voice and generates text data such as "Please describe the dog."

[0544] Step 6:

[0545] The server analyzes the text data obtained from the speech recognition engine and passes it to the generation AI to extract relevant content.

[0546] Step 7:

[0547] The server's generation AI generates a 3D image based on text data. For example, based on the data for "dog," it generates a 3D model of a dog.

[0548] Step 8:

[0549] The server sends the generated 3D model data to the hologram generation engine. The data from the generation AI is passed to the conversion engine, which prepares the hologram data for creation.

[0550] Step 9:

[0551] The server's hologram generation engine converts the 3D stereoscopic image data into hologram data, specifically, converting the 3D model into a format suitable for holographic display.

[0552] Step 10:

[0553] The server sends the converted hologram data to the terminal via a communication interface.

[0554] Step 11:

[0555] The terminal receives the hologram data, temporarily stores it in a buffer, retrieves the data through the receiving interface, and starts processing it immediately.

[0556] Step 12:

[0557] The terminal initiates the process for transmitting hologram data to an assistive device, such as a projector or hologram display.

[0558] Step 13:

[0559] The assistive device uses the received hologram data to project a 3D stereoscopic image in real time. The hologram projection device interprets the data and displays a realistic stereoscopic image in space.

[0560] Each step allows the user to seamlessly experience the audio-to-hologram conversion process.

[0561] Example 1

[0562] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0563] Many modern learning and presentation systems provide information through text or 2D images, but these do not adequately support visual comprehension. Furthermore, conventional voice input systems struggle to visually display the input content in real time, requiring numerous steps and devices. This complicates the user experience and makes intuitive operation difficult.

[0564] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0565] In this invention, the server includes a voice recognition unit that converts voice data into text, an analysis unit that analyzes the text information, a generation unit that generates a 3D stereoscopic image based on the text information, and a conversion unit that converts the 3D stereoscopic image into hologram data. This makes it possible to generate a 3D stereoscopic image in real time based on a user's voice input and display the image as a hologram.

[0566] The "input means" is a device that allows the user to input voice.

[0567] "Communication means" is a mechanism for transmitting voice data from a terminal to a server.

[0568] "Speech recognition means" refers to software or algorithms that convert voice data into text on the server.

[0569] "Analysis means" refers to the process or technology by which the server analyzes text information and extracts the necessary information.

[0570] The "generation means" is a technology that allows the server to generate a 3D stereoscopic image based on text information.

[0571] The "conversion means" is the technology that the server uses to convert the 3D stereoscopic image into hologram data.

[0572] "Display means" refers to a device or technology that allows a terminal to receive hologram data and project it in real time.

[0573] The "support device" is an external device for actually projecting and displaying hologram data.

[0574] This invention relates to a system that generates a 3D stereoscopic image in real time based on a user's voice input and displays it as a hologram. Specific embodiments of this system will be described below.

[0575] The system consists of the following main components:

[0576] 1. Audio input device (terminal)

[0577] 2. Communication Devices (Terminals)

[0578] 3. Speech recognition engine (server)

[0579] 4. Generative AI (server)

[0580] 5. Hologram generation engine (server)

[0581] 6. Projection Support Devices (Terminals)

[0582] Voice input

[0583] First, a user can operate the system by speaking into a voice input device. For example, the user can speak commands such as "Tell me about the planets in the solar system." The voice input device captures the voice and generates audio data. Specific voice input devices can be built-in microphones or dedicated voice input devices.

[0584] Sending audio data

[0585] The collected audio data is then transmitted to a server via the communication device, using standard wireless communication technologies such as Wi-Fi or Bluetooth.

[0586] Voice Recognition

[0587] The server then passes the received voice data to a voice recognition engine for analysis. An example of a voice recognition engine is the Google Cloud Speech-to-Text API. This engine analyzes phonemes and phonology to convert the voice data into text, and generates the corresponding text. For example, a command such as "Tell me about the planets in the solar system" is generated as text data.

[0588] Text analysis and 3D image generation

[0589] Once the text data is generated, an analysis tool on the server analyzes the text data and extracts the necessary information. Natural language processing (NLP) technology is used as the analysis tool. Based on the results of this analysis, the server's generative AI (such as OpenAI's GPT-3) creates data for generating 3D stereoscopic images. For example, a series of 3D stereoscopic images (planet models) can be generated based on data on the planets in the solar system.

[0590] Hologram Generation

[0591] The generated 3D stereoscopic image data is passed to the server's hologram generation engine, which converts it into hologram data. The hologram generation engine can use, for example, Looking Glass Factory's hologram generation technology. This engine converts the 3D model data into images from multiple viewpoints and combines them to create hologram data.

[0592] Hologram display

[0593] The hologram data is sent from the server to the terminal and displayed. The terminal receives this data and passes it to a projection assistance device, specifically a hologram projector, which uses light interference patterns to display 3D images in space.

[0594] Specific examples

[0595] For example, consider the following scenario for an educational system:

[0596] User: "Tell me about the planets in our solar system."

[0597] The device captures the audio and sends it to the server. At this stage, the built-in microphone picks up the audio signal and sends it to the server as digital data.

[0598] The server's speech recognition engine converts the text to "Tell me about the planets in the solar system," using feature extraction and statistical models of the speech signal.

[0599] The server's NLP engine parses this text and understands it as a specific information request.

[0600] The server sends the generated AI a prompt saying, "Provide information about planets in the solar system."

[0601] Generative AI generates a series of 3D stereoscopic images based on data for each planet.

[0602] The server's hologram generation engine converts these 3D models into hologram data.

[0603] The terminal receives the hologram data and transmits it to the display device, using Wi-Fi or a wired connection.

[0604] The assistive device displays a real-time stereoscopic model of the planets in the solar system. The projector uses light patterns to create 3D images in space that can be seen from multiple perspectives.

[0605] In this way, users can visually study detailed 3D holograms in real time through voice input.

[0606] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0607] Step 1:

[0608] The user inputs commands to the system by speaking into a voice input device. For example, the user might say, "Please describe the dog." The input is the user's voice data.

[0609] Step 2:

[0610] The terminal captures the user's voice and collects the voice data. The microphone inside the device converts the voice signal into digital data and processes the data in real time. The input is the user's voice data, and the output is the digitized voice data.

[0611] Step 3:

[0612] The voice data collected by the device is sent to a server via a communication device. Specifically, the data is transferred using wireless communication protocols such as Wi-Fi or Bluetooth. The input is digitized voice data, and the output is the voice data sent to the server.

[0613] Step 4:

[0614] The server passes the received voice data to a voice recognition engine (voice recognition means) for analysis. The voice recognition engine (for example, general voice recognition software) analyzes the voice data and converts it into text information. The input is voice data and the output is text information.

[0615] Step 5:

[0616] The server's analysis means analyzes the generated text information and extracts the necessary information. Natural language processing (NLP) technology is used for this analysis, and specific keywords and command content are extracted from the analysis results. The input is text information, and the output is the extracted keywords and command data.

[0617] Step 6:

[0618] The server passes a prompt to the generation AI based on the extracted command data. For example, a prompt such as "Please describe the dog" is input to the generation AI. The input is the command data, and the output is the prompt.

[0619] Step 7:

[0620] Based on the prompt received, the generation AI generates relevant 3D stereoscopic image data. For example, the keyword "dog" generates 3D model data of a dog. The input is the prompt, and the output is 3D stereoscopic image data.

[0621] Step 8:

[0622] The server's hologram generation engine (conversion means) converts the generated 3D stereoscopic image data into hologram data. Specifically, it uses 3D model data to generate images from multiple viewpoints and then combines them to create hologram data. The input is 3D stereoscopic image data, and the output is hologram data.

[0623] Step 9:

[0624] The server transmits the converted hologram data to the terminal via Wi-Fi or a wired connection. The input is the hologram data, and the output is the hologram data transmitted to the terminal.

[0625] Step 10:

[0626] The terminal receives the hologram data and passes it to the support device. For example, it transmits the data to a hologram projector. The input is the hologram data, and the output is the hologram data passed to the support device.

[0627] Step 11:

[0628] The assistive device uses the hologram data to project a 3D image in real time. For example, a hologram projector displays a 3D image in space based on the data it receives. The input is hologram data, and the output is a 3D hologram displayed in space.

[0629] (Application example 1)

[0630] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0631] Currently, physical stores have limited means of providing detailed product information, making it difficult for customers to fully understand product features. Furthermore, there is a lack of efficient means for providing visual product information to multiple customers simultaneously. Therefore, there is a demand for technology that can easily display product information in real time using voice input to improve the user experience and stimulate purchasing motivation.

[0632] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0633] In this invention, the server includes an input means for a user to input voice, a communication means for a terminal to transmit voice data to the server, a voice recognition means for the server to convert the voice data into text, a generation means for the server to generate a 3D stereoscopic image using a generative artificial intelligence model based on the text information, a conversion means for the server to convert the 3D stereoscopic image into hologram data, and a display means for the terminal to receive the hologram data and project it in real time. This makes it possible to display interactive 3D holograms in physical stores that allow customers to intuitively understand detailed product information through voice commands.

[0634] "Input means for inputting voice" refers to a device or technology that allows a user to give instructions to a system via voice.

[0635] "Communication means for transmitting audio data to a server" refers to a method or device for sending audio data to a server via the Internet or other communication network.

[0636] "Speech recognition means for converting voice data into text" refers to the algorithms or software used to analyze voice data and convert it into text data.

[0637] "Generative AI model" refers to an AI model or system used to generate 3D stereoscopic images based on text information.

[0638] "Means for converting a 3D stereoscopic image into hologram data" refers to an algorithm or system for converting a 3D stereoscopic image into a data format that can be projected as a hologram.

[0639] "Display means for receiving holographic data and projecting it in real time" refers to a device or technology for visually projecting the received holographic data in real time.

[0640] A "prompt" is a piece of text used to instruct a generative AI model to generate a particular output.

[0641] "Display device" is a general term for devices and hardware for visually displaying generated hologram data.

[0642] A "physical store" refers to a store located in a physical location where customers can visit in person to view and purchase products.

[0643] The present invention relates to a system that displays detailed product information in a physical store as a 3D hologram in real time. This system allows a user to input voice, and then projects a 3D stereoscopic image generated based on the voice as a hologram. Specific embodiments of this system are described below.

[0644] System configuration

[0645] The system consists of the following main components:

[0646] 1. Voice input means: A microphone for users to input voice.

[0647] 2. Communication method: The method by which the voice data is transmitted to the server via the Internet.

[0648] 3. Speech recognition means: A speech recognition engine (e.g., Google Speech-to-Text) that the server uses to convert voice data into text.

[0649] 4. Generation: The process in which the server uses a generative AI model (e.g., OpenAI's GPT-3) to generate a 3D stereoscopic image from text information.

[0650] 5. Conversion method: The algorithm used by the server to convert the 3D stereoscopic image into hologram data (e.g., Unity's hologram generation plug-in).

[0651] 6. Display means: A device that projects holographic data in real time using a smartphone or a holographic display (e.g., Looking Glass Factory's holographic display).

[0652] Processing flow

[0653] 1. Voice input: The user speaks into the smartphone, for example, "Please tell me the features of this new smartphone."

[0654] 2. Sending audio data: The smartphone sends the audio data to the server via an HTTP request over the internet.

[0655] 3. Speech recognition: The server passes the received voice data to a speech recognition engine and converts it into text. For example, the generated text is "Please tell me the features of this new smartphone."

[0656] 4. Text analysis and 3D image generation: The server uses generative AI models to generate relevant 3D images from text information. For example, a 3D model is generated based on the detailed features and specifications of a smartphone.

[0657] 5. Hologram generation: The server's hologram generation engine analyzes the 3D stereoscopic image and converts it into data that can be displayed as a hologram.

[0658] 6. Hologram display: A smartphone receives hologram data and projects it in real time on a built-in hologram display, etc. This device can display interactive 3D images in real time.

[0659] Specific example explanation

[0660] Example 1:

[0661] When a user speaks to a smartphone sales counter in a physical store and says, "What are the features of this new smartphone?", the system instantly projects a 3D model of the smartphone and its features as a hologram, allowing the user to understand the details of the product without having to touch it.

[0662] Example prompt sentence:

[0663] User dictates: "What are the features of this new phone?"

[0664] Server prompt: "Please describe the features of the new smartphone XYZ. Generate a 3D model of it."

[0665] This system will significantly improve the customer experience in physical stores, increasing product understanding and purchasing intent.

[0666] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0667] Step 1:

[0668] A user uses the smartphone's voice input means to ask a product-related question. Specifically, the user says, "What are the features of this new smartphone?" The input is the user's voice, and the output is the voice data captured by the smartphone's microphone.

[0669] Step 2:

[0670] The device sends the captured audio data to the server via a communication means. The audio data is sent to the server as an HTTP request over the Internet. The input is the audio data captured by the smartphone's microphone, and the output is the audio data sent to the server.

[0671] Step 3:

[0672] The server receives the voice data and passes it to a voice recognition engine to convert it into text. Specifically, the server analyzes the received voice data and generates text data such as "Please tell me the features of this new smartphone." The input is voice data and the output is text data.

[0673] Step 4:

[0674] The server passes the generated text data to a generative AI model, which then generates a 3D stereoscopic image based on the text information. Specifically, the generative AI model generates a 3D model of a smartphone based on the prompt, "Features of the new smartphone." The input is text data, and the output is 3D stereoscopic image data.

[0675] Step 5:

[0676] The server passes the generated 3D stereoscopic image data to an algorithm for converting it into hologram data. Here, the server's hologram generation engine converts the 3D stereoscopic image into a data format that can be displayed as a hologram. The input is 3D stereoscopic image data, and the output is hologram data.

[0677] Step 6:

[0678] The device receives the hologram data from the server and projects it in real time using the built-in display means. Specifically, a smartphone or holographic display visually displays the hologram data. The input is the hologram data, and the output is the displayed hologram.

[0679] This will allow users to view interactive 3D holograms in physical stores that provide intuitive understanding of product details.

[0680] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0681] The system of the present invention generates 3D stereoscopic images in real time based on the user's voice input and displays them as holograms. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide a more interactive and personalized experience. Below, we will create a program for this system and explain its processing in detail.

[0682] overview

[0683] The system consists of the following main components:

[0684] 1. A voice input device (terminal) that inputs the user's voice

[0685] 2. A communication device (terminal) that transmits voice data to a server

[0686] 3. Speech recognition engine (server) that converts voice data into text

[0687] 4. Generative AI (server) that generates 3D images based on text information

[0688] 5. Hologram generation engine (server) that converts the generated 3D stereoscopic image into hologram data

[0689] 6. Support device (terminal) that receives and projects hologram data

[0690] 7. Emotion engine (server) that recognizes user emotions

[0691] Program processing flow

[0692] Voice input

[0693] 1. The user speaks into a voice input device, for example, "Please describe my dog."

[0694] 2. The terminal captures the user's voice and converts it into digital voice data. The voice input device collects the voice and performs pre-processing to convert it into a data format.

[0695] 3. Communication is established to send the audio data to the server.

[0696] Voice Recognition

[0697] 4. The server passes the received voice data to the speech recognition engine, which decodes the voice data and provides it as input to the speech recognition engine.

[0698] 5. The speech recognition engine converts the voice data into text data. For example, the text data generated is "Please describe the dog."

[0699] emotion recognition

[0700] 6. The server passes the voice data to the emotion engine, which analyzes the user's emotional state. The emotion engine analyzes the tone and tempo of the voice to identify the user's emotion (happiness, sadness, surprise, etc.).

[0701] 7. Provide the emotion engine's analysis results (user's emotional state) to the generation AI.

[0702] Text analysis and 3D image generation

[0703] 8. The server analyzes the generated text information and sentiment analysis results and passes them to the generation AI to extract relevant content.

[0704] 9. The generation AI generates a 3D image based on the input text and emotion information. For example, it generates a 3D model of a dog based on the keyword "dog," and adds movements and expressions that correspond to the user's emotions.

[0705] 10. The generated 3D stereoscopic image data is sent to the hologram generation engine.

[0706] Hologram Generation

[0707] 11. The server's hologram generation engine converts the 3D stereoscopic image data into hologram data, specifically converting the 3D model into a format suitable for holographic display.

[0708] 12. The converted hologram data is sent to the terminal.

[0709] Hologram display

[0710] 13. The terminal receives the hologram data and temporarily stores it in a buffer. It then retrieves the data through the receiving interface and starts processing it immediately.

[0711] 14. The terminal starts the process of transmitting hologram data to the assistive device, such as a projector or hologram display.

[0712] 15. The assistive device uses the received hologram data to project a 3D stereoscopic image in real time. The hologram projection device interprets the data and displays a realistic stereoscopic image in space.

[0713] Specific examples

[0714] Educational Demo Scenario

[0715] User: "Tell me about the planets in our solar system."

[0716] The device captures the user's voice and sends it to the server.

[0717] The server's speech recognition engine converts the received voice data into text such as "Tell me about the planets in the solar system."

[0718] The server's emotion engine analyzes the user's emotions from the voice data, and recognizes the emotion of joy.

[0719] The server's generation AI generates a series of 3D stereoscopic images (planet models) based on the data of each planet, adding bright colors and movement according to the user's pleasure.

[0720] The server's hologram generation engine converts these 3D models into hologram data.

[0721] The terminal receives the hologram data and transmits it to the display device.

[0722] The assistive device displays a 3D model of the planets in the solar system in real time.

[0723] The system allows users to enjoy a richer, more personalized holographic experience in real time, utilizing both voice input and emotion analysis.

[0724] The processing flow will be explained below.

[0725] Step 1:

[0726] A user speaks into a voice input device, for example, "Describe my dog."

[0727] Step 2:

[0728] The terminal captures the user's voice and converts it into digital voice data. The voice input device collects the voice and pre-processes it to convert it into a data format.

[0729] Step 3:

[0730] The device transmits the captured audio data to the server in real time. The audio data is sent to the server as a stream via the communication interface.

[0731] Step 4:

[0732] The server passes the received voice data to the speech recognition engine, which then decodes the voice data and provides it as input to the speech recognition engine.

[0733] Step 5:

[0734] The server's speech recognition engine converts the voice data into text data. Specifically, the speech recognition engine analyzes the voice and generates text data such as "Please describe the dog."

[0735] Step 6:

[0736] The server passes the voice data to the emotion engine, which analyzes the user's emotional state. The emotion engine analyzes the tone and tempo of the voice to identify the user's emotion (happiness, sadness, surprise, etc.).

[0737] Step 7:

[0738] The server's emotion engine provides the analysis results (user's emotional state) to the generation AI, which then performs the following processes based on the received emotional information.

[0739] Step 8:

[0740] The server passes the generated text information and emotion analysis results to the generation AI, which extracts relevant content. The generation AI then generates a 3D image based on the input text information and emotion information.

[0741] Step 9:

[0742] The server's AI generates 3D model data for a dog based on a keyword, such as "dog," and adds movements and facial expressions that correspond to the user's emotions.

[0743] Step 10:

[0744] The server sends the generated 3D model data to the hologram generation engine. The data from the generation AI is passed to the conversion engine, which prepares the hologram data for creation.

[0745] Step 11:

[0746] The server's hologram generation engine converts the 3D stereoscopic image data into hologram data, specifically, converting the 3D model into a format suitable for holographic display.

[0747] Step 12:

[0748] The server sends the converted hologram data to the terminal via a communication interface.

[0749] Step 13:

[0750] The terminal receives the hologram data, temporarily stores it in a buffer, retrieves the data through the receiving interface, and starts processing it immediately.

[0751] Step 14:

[0752] The terminal initiates the process for transmitting hologram data to an assistive device, such as a projector or hologram display.

[0753] Step 15:

[0754] The assistive device uses the received hologram data to project a 3D stereoscopic image in real time. The hologram projection device interprets the data and displays a realistic stereoscopic image in space.

[0755] Specific examples

[0756] Educational Demo Scenario

[0757] Step 1:

[0758] A user speaks into a voice input device, "Tell me about the planets in the solar system."

[0759] Step 2:

[0760] The device captures the user's voice and converts it into digital voice data.

[0761] Step 3:

[0762] The device transmits the captured audio data to the server in real time.

[0763] Step 4:

[0764] The server passes the received voice data to the voice recognition engine.

[0765] Step 5:

[0766] The server's voice recognition engine converts the voice data into text data such as "Tell me about the planets in the solar system."

[0767] Step 6:

[0768] The server passes the voice data to the emotion engine, which analyzes the user's emotional state (interest and pleasure).

[0769] Step 7:

[0770] The server's emotion engine provides the analysis results to the generation AI.

[0771] Step 8:

[0772] The server passes the generated text information and sentiment analysis results to the generation AI, which extracts relevant content.

[0773] Step 9:

[0774] The server's generation AI generates a series of 3D stereoscopic images (planet models) based on data for each planet, adding bright colors and movement according to the user's emotions.

[0775] Step 10:

[0776] The server sends the generated 3D model data to the hologram generation engine.

[0777] Step 11:

[0778] The server's hologram generation engine converts the generated 3D model data into hologram data.

[0779] Step 12:

[0780] The server transmits the converted hologram data to the terminal.

[0781] Step 13:

[0782] The terminal receives the hologram data and temporarily stores it in a buffer.

[0783] Step 14:

[0784] The terminal initiates the process for transmitting the hologram data to the assistive device.

[0785] Step 15:

[0786] The assistive device displays a 3D model of the planets in the solar system in real time.

[0787] Example 2

[0788] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0789] Conventional voice input systems simply recognize a user's voice, convert it into text, and display a static 3D stereoscopic image based on that text. This limited the user experience, making it difficult to provide an interactive and personalized experience. Furthermore, the lack of technology to analyze user emotions and generate content based on those emotions prevented users from achieving high levels of satisfaction.

[0790] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0791] In this invention, the server includes means for converting voice data into text, means for generating a 3D stereoscopic image based on the text information, means for recognizing a user's emotion from the voice data, and means for adjusting the movements and facial expressions of the generated 3D stereoscopic image based on the user's emotion information, thereby making it possible to provide an interactive and personalized hologram experience in real time according to the user's emotion.

[0792] "Input means" refers to a device that allows a user to input voice.

[0793] "Communication means" refers to a network interface that enables a terminal to transmit audio data to a server.

[0794] "Speech recognition means" refers to the algorithm or engine that converts the voice data received by the server into text data.

[0795] "Generation means" refers to the algorithms or programs that the server uses to generate 3D stereoscopic images based on text information.

[0796] "Conversion means" refers to the algorithm or engine that the server uses to convert 3D stereoscopic images into hologram data.

[0797] "Display means" refers to a device for projecting hologram data received by the terminal in real time.

[0798] "Recognition means" refers to the algorithms or engines that the server uses to analyze the user's emotions from voice data.

[0799] "Adjustment means" refers to an algorithm or program that the generation means uses to adjust the movements and facial expressions of the 3D stereoscopic image based on the user's emotional information.

[0800] "Support equipment" refers to an external device for projecting a hologram.

[0801] This invention provides a system that generates 3D stereoscopic images in real time based on the user's voice input and displays them as holograms, and also recognizes the user's emotions and provides a personalized experience based on them.

[0802] Hardware and software used

[0803] Hardware:

[0804] Audio input devices: microphones, smartphones, etc.

[0805] Communication device: smartphone, tablet, or computer.

[0806] Server: High-performance cloud or on-premise servers.

[0807] Assistive devices: Hologram projector, VR / AR headset (e.g. Microsoft HoloLens).

[0808] software:

[0809] Speech recognition engine: Google Speech-to-Text API, Azure Speech Service, etc.

[0810] Emotion recognition engine: IBM Watson Tone Analyzer.

[0811] Generative AI: OpenAI GPT-3 or similar advanced language model.

[0812] Hologram generation engine: Unity or Unreal Engine.

[0813] Program processing explanation

[0814] 1. Voice input

[0815] A user speaks into a voice input device, for example, "Please describe my dog."

[0816] 2. Communications

[0817] In order for the terminal to transmit the audio data to the server, the digital audio data is divided into packets and transmitted to the server via secure communication.

[0818] 3. Voice Recognition

[0819] The server passes the received voice data to a voice recognition engine, which converts the voice into text data.

[0820] 4. Emotion recognition

[0821] The server passes the voice data to an emotion recognition engine, which analyzes the user's emotional state. The emotion engine analyzes the tone and tempo of the voice to identify the user's emotion (happiness, sadness, surprise, etc.).

[0822] 5. Text Analysis and 3D Image Generation

[0823] The server inputs the generated text information and emotion analysis results into a generative AI model to generate appropriate 3D stereoscopic image data. For example, based on the text "dog" and emotion information, a 3D model of a dog is generated and its movements and facial expressions are adjusted according to the emotion.

[0824] 6. Hologram Generation

[0825] The server's hologram generation engine converts the 3D stereoscopic image data into hologram data. Specifically, it uses Unity or Unreal Engine to encode the 3D model into a format suitable for holograms.

[0826] 7. Data Transmission

[0827] The server transmits the converted hologram data to the terminal.

[0828] 8. Hologram display

[0829] The terminal receives the hologram data and transmits it to the support device, which then displays the hologram in real time.

[0830] Specific examples

[0831] Educational Demo Scenario

[0832] User: "Tell me about the planets in our solar system."

[0833] The device captures the user's voice, converts it into digital audio data, and sends it to the server.

[0834] The server's speech recognition engine converts the received voice data into text such as "Tell me about the planets in the solar system."

[0835] The server's emotion recognition engine analyzes the user's emotions from the voice and recognizes the emotion of joy.

[0836] The server's generative AI model generates a series of 3D stereoscopic images (planet models) based on data for each planet, adding bright colors and movement depending on the user's emotions.

[0837] The server's hologram generation engine converts these 3D models into hologram data.

[0838] The terminal receives the hologram data and transmits it to the support device.

[0839] The support device displays a 3D model of the planets in the solar system in real time.

[0840] Example prompt: "Tell me about the planets in the solar system."

[0841] The system allows users to enjoy a richer, more personalized holographic experience in real time, utilizing both voice input and emotion analysis.

[0842] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0843] Step 1:

[0844] The user speaks into the voice input device, saying, "Please describe my dog," and the voice input device captures the voice and converts it into digital voice data.

[0845] Input: User spoken words.

[0846] Output: Digital audio data.

[0847] What happens: An audio input device (e.g., a smartphone microphone) converts an analog audio signal into a digital signal, which is sampled at 44.1 kHz and recorded in PCM format.

[0848] Step 2:

[0849] The terminal sends voice data to the server using a communication method. The voice data is divided into packets and sent to the server using secure communication (SSL / TLS).

[0850] Input: Digital audio data.

[0851] Output: The audio data sent to the server.

[0852] What it does: The device splits the audio data into packets of 1024 bytes each and sends the series of packets to the server over an SSL / TLS connection.

[0853] Step 3:

[0854] The server receives the voice data and passes it to the voice recognition engine, which converts the received voice data into text data.

[0855] Input: Audio data sent to the server.

[0856] Output: Text data.

[0857] Specific operation: The server temporarily stores the voice data in storage and calls the Google Speech-to-Text API to convert the voice data into text data. For example, the text generated is "Please describe the dog."

[0858] Step 4:

[0859] The server passes the voice data to an emotion recognition engine, which analyzes the user's emotional state. The emotion recognition engine analyzes the tone and tempo of the voice to recognize the user's emotions.

[0860] Input: Audio data.

[0861] Output: Emotion analysis results.

[0862] How it works: The server calls IBM Watson Tone Analyzer, which analyzes the tone, pitch, speed, etc. of the voice data to identify emotions. For example, the voice may be judged to be "joyful," and the joy score may be high.

[0863] Step 5:

[0864] The server inputs text and emotion information into the generative AI model and obtains the generated 3D stereoscopic image data. The generative AI generates a 3D model based on the text and adds movements and facial expressions according to the emotion information.

[0865] Input: Text data, sentiment analysis results.

[0866] Output: 3D stereoscopic image data.

[0867] Specific operation: The server calls OpenAI GPT-3, inputs the text prompt "Please describe the dog" and the emotion "joy" as input, and generates 3D stereoscopic image data. For example, a joyful motion of the dog jumping up and down is added.

[0868] Step 6:

[0869] The server converts the 3D stereoscopic image into hologram data, which is then converted into a format suitable for holographic projection.

[0870] Input: 3D stereoscopic image data.

[0871] Output: Hologram data.

[0872] How it works: The server uses Unity or Unreal Engine to convert 3D stereoscopic image data into a volumetric hologram format, for example encoding a 3D model of a dog into a holographic format.

[0873] Step 7:

[0874] The server transmits hologram data to the terminal, which receives the data and transmits it to the support device.

[0875] Input: Hologram data.

[0876] Output: Hologram data sent to the device.

[0877] Specific operation: The server packetizes the hologram data and sends it to the device via an SSL / TLS connection. The device temporarily stores the received data in a buffer.

[0878] Step 8:

[0879] The device transmits the hologram data to the support device, which displays the hologram in real time. The support device interprets the hologram data and projects it into real space.

[0880] Input: Hologram data.

[0881] Output: The projected hologram.

[0882] How it works: The device uses Wi-Fi 6 (IEEE 802.11ax) to transmit hologram data to an assistive device (e.g., Microsoft HoloLens). The assistive device interprets the data and projects a real-time 3D hologram of a bouncing dog into space.

[0883] (Application example 2)

[0884] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0885] In today's brick-and-mortar stores, customers are limited to a single explanation method and fixed information when receiving product explanations and store guidance. It is also difficult to provide a personalized experience tailored to the customer's emotions and state, making it difficult to provide effective customer service. With such limited information provided, it is difficult to increase customer satisfaction and purchasing motivation.

[0886] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0887] In this invention, the server includes an emotion recognition means for analyzing voice data to recognize a user's emotion, a generation means for personalizing a 3D stereoscopic image based on the user's emotion, and a generation means for generating an interactive 3D stereoscopic image, thereby enabling a richer and more personalized hologram experience based on the user's voice input and emotion analysis.

[0888] The "voice input means" is a device that allows the user to input voice.

[0889] "Communication means" is a function that enables a terminal to transmit voice data to a server.

[0890] "Speech recognition means" is a function by which the server converts voice data into text.

[0891] The "generation means" is a function that allows the server to generate a 3D stereoscopic image based on text information.

[0892] The "conversion means" is a function by which the server converts a 3D stereoscopic image into hologram data.

[0893] "Emotion recognition means" is a function in which the server analyzes voice data and recognizes the user's emotions.

[0894] The "means for personalizing" is a function by which the generating means personalizes the 3D stereoscopic image based on the user's emotions.

[0895] "Display means" refers to the function of the terminal to receive hologram data and project it in real time.

[0896] As an embodiment of this invention, we will describe an interactive product explanation system for use in a brick-and-mortar store. This system provides a personalized experience by generating 3D stereoscopic images in real time based on the user's voice input and emotion analysis, and displaying the images as holograms.

[0897] System configuration

[0898] 1. Voice input means: The user uses a device to input voice. This can include a microphone on a smartphone or smart glasses.

[0899] 2. Communication means: The device has the ability to send captured audio data to the server.

[0900] 3. Speech recognition: The server uses a speech recognition engine to convert the voice data into text, for example, Google Cloud Speech-to-Text API.

[0901] 4. Emotion recognition: The server analyzes the voice data and uses an emotion engine to recognize the user's emotions, for example, IBM Watson Tone Analyzer.

[0902] 5. Generation method: The server uses a generative AI model, such as OpenAI GPT-4, to generate a 3D stereoscopic image based on text information and emotions.

[0903] 6. Conversion means: The server uses a hologram generation engine to convert the generated 3D stereoscopic image into hologram data.

[0904] 7. Display means: Includes a support device for the terminal to receive holographic data and project it in real time, such as a holographic display from Looking Glass Factory.

[0905] Program processing

[0906] The system does the following:

[0907] First, the user uses the voice input means to speak about the product or information they want to hear explained about. For example, they might say, "Please explain the features of this product." At that time, the voice input means captures the user's voice and converts it into digital voice data. This data is then sent to the server via the communication means.

[0908] The voice data that arrives at the server is converted into text data using a voice recognition means. In parallel, an emotion recognition means analyzes the voice data and recognizes the user's emotions. The user's emotional state (e.g., curiosity or surprise) is used to further refine the interaction.

[0909] Next, the text information and emotion information are passed to a generation means, which uses a generative AI model to generate a 3D stereoscopic image. The generated 3D stereoscopic image is personalized by applying movements and facial expressions according to the user's emotions. The 3D stereoscopic image is then converted into hologram data by a conversion means.

[0910] Finally, the hologram data is transmitted to the support device through the display means and projected in real time in the physical store, allowing users to visualize the products and information as interactive holograms right in front of their eyes.

[0911] Specific examples

[0912] For example, suppose a user uses a voice input method in a store to say, "Please explain the features of this product." The voice is immediately sent to the server, where speech recognition and emotion analysis are performed. As a result, the text data "Please explain the features of this product" and the user's emotional state (curiosity) are acquired. Based on the text and emotion information, the generative AI model generates an easy-to-understand and interesting 3D stereoscopic image corresponding to the "curiosity." For example, this could include an animation or exploded view showing the internal structure of the product.

[0913] The generated 3D stereoscopic image is converted into hologram data and projected as a hologram in real time through a display device. Users can visually observe the hologram and manipulate it interactively.

[0914] Example prompt sentence:

[0915] "Please explain. Analyze the user's emotions and generate a 3D model based on those emotions. Now, the user said, 'Explain the features of this product.' The user's emotion is 'Curiosity.'"

[0916] This allows users to enjoy a more personalized experience in physical stores.

[0917] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0918] Program processing

[0919] Step 1:

[0920] The user uses a voice input means to talk about the product or information they want to hear explained. For example, they might say, "Please explain the features of this product." In this case, the input is the user's voice data. This voice data is captured by the microphone of a smartphone or smart glasses and converted into digital voice data. The output is digitized voice data.

[0921] Step 2:

[0922] The terminal transmits the captured voice data to the server via a communication means. At this time, the input is the captured digital voice data. The server receives this data and prepares to pass it to the voice recognition engine. The output is the result of transmitting the voice data to the server.

[0923] Step 3:

[0924] The server converts the voice data into text data using a speech recognition tool. The input is digital voice data. A speech recognition engine (e.g., Google Cloud Speech-to-Text) analyzes the data and generates text data such as "Describe the features of this product." The output is text data.

[0925] Step 4:

[0926] At the same time, the server passes the voice data to an emotion recognition means to analyze the user's emotions. The input is digital voice data. The emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the tone and tempo of the voice to identify the user's emotional state (e.g., data expressing curiosity or surprise). The output is the user's emotional data.

[0927] Step 5:

[0928] The server passes the generated text information and emotion analysis results to the generation means. At this time, the inputs are text data and emotion data. The generative AI model (e.g., OpenAI GPT-4) generates a 3D stereoscopic image based on the text and emotion information, corresponding to the user's emotion. The output is 3D stereoscopic image data.

[0929] Step 6:

[0930] The server passes the generated 3D stereoscopic image data to a conversion means, which converts it into hologram data. At this time, the input is the 3D stereoscopic image data. The hologram generation engine converts the 3D model data into a format suitable for hologram display. The output is hologram data.

[0931] Step 7:

[0932] The server sends the converted hologram data to the terminal. At this time, the input is the hologram data. The terminal receives the data through the receiving interface and temporarily stores it in a buffer. The output is the data storage result on the terminal.

[0933] Step 8:

[0934] The terminal transmits hologram data to the assistive device, which projects it in real time. The input is the hologram data. The assistive device (e.g., Looking Glass Factory's holographic display) interprets the received data and projects a realistic 3D image in space. The output is the user's visual holographic experience.

[0935] Specific operation example

[0936] For example, a user might say, "Explain the features of this product" using a voice input method while in a store. The voice is immediately sent to the server, where speech recognition and emotion analysis are performed. As a result, the text data "Explain the features of this product" and the user's emotional state (curiosity) are acquired. Based on the text and emotion information, a generative AI model generates an easy-to-understand and interesting 3D stereoscopic image that corresponds to the "curiosity." For example, this could include animations or exploded views that show the internal structure of the product.

[0937] Example prompt sentence:

[0938] "Please explain. Analyze the user's emotions and generate a 3D model based on those emotions. Now, the user said, 'Explain the features of this product.' The user's emotion is 'Curiosity.'"

[0939] In this way, a personalized and interactive hologram experience can be provided based on the user's voice input and emotion analysis.

[0940] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0941] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0942] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0943] [Third embodiment]

[0944] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0945] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0946] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0947] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0948] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0949] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0950] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0951] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0952] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0953] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0954] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0955] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0956] The system of the present invention generates a 3D stereoscopic image in real time based on the user's voice input and displays it as a hologram. Below, we will create a program for this system and explain its processing in detail.

[0957] overview

[0958] The system consists of the following main components:

[0959] 1. A voice input device (terminal) that inputs the user's voice

[0960] 2. A communication device (terminal) that transmits voice data to a server

[0961] 3. Speech recognition engine (server) that converts voice data into text

[0962] 4. Generative AI (server) that generates 3D images based on text information

[0963] 5. Hologram generation engine (server) that converts the generated 3D stereoscopic image into hologram data

[0964] 6. Support device (terminal) that receives and projects hologram data

[0965] Program processing flow

[0966] Voice input

[0967] 1. The user speaks into a voice input device, giving a command such as "Describe the dog."

[0968] 2. The device captures the user's voice and collects voice data in real time.

[0969] 3. Communication is established to send the audio data to the server.

[0970] Voice Recognition

[0971] 4. The server passes the received voice data to the voice recognition engine.

[0972] 5. The speech recognition engine analyzes the voice data and converts it into text information. For example, the text data generated is "Please describe the dog."

[0973] Text analysis and 3D image generation

[0974] 6. The server analyzes the generated text information and passes it to the generation AI to extract relevant content.

[0975] 7. Generative AI generates a 3D image based on the input text information. For example, it creates a 3D model of a dog based on the keyword "dog."

[0976] 8. The generated 3D stereoscopic image data is sent to the hologram generation engine.

[0977] Hologram Generation

[0978] 9. The server's hologram generation engine converts the 3D image data into hologram data.

[0979] 10. The converted hologram data is sent to the terminal.

[0980] Hologram display

[0981] 11. The terminal receives the hologram data and transmits it to the support device.

[0982] 12. The assistive device uses the holographic data to project a 3D image in real time.

[0983] Specific examples

[0984] Educational Demo Scenario

[0985] User: "Tell me about the planets in our solar system."

[0986] The device captures the audio and sends it to the server.

[0987] The server's speech recognition engine converts the received voice data into text such as "Tell me about the planets in the solar system."

[0988] The server's generation AI generates a series of 3D stereoscopic images (planet models) based on the data of each planet.

[0989] The server's hologram generation engine converts these 3D models into hologram data.

[0990] The terminal receives the hologram data and transmits it to the display device.

[0991] The assistive device displays a 3D model of the planets in the solar system in real time.

[0992] This system allows users to learn visually through holograms in real time using only voice input, and is expected to be used in a variety of situations, not just education, but also entertainment and business presentations.

[0993] The processing flow will be explained below.

[0994] Step 1:

[0995] A user speaks into a voice input device, for example, "Describe my dog."

[0996] Step 2:

[0997] The terminal captures the user's voice and converts it into digital voice data. The voice input device collects the voice and pre-processes it to convert it into a data format.

[0998] Step 3:

[0999] The device transmits the captured audio data to the server in real time. The audio data is sent to the server as a stream via the communication interface.

[1000] Step 4:

[1001] The server passes the received voice data to the speech recognition engine, which then decodes the voice data and provides it as input to the speech recognition engine.

[1002] Step 5:

[1003] The server's speech recognition engine converts the voice data into text data. Specifically, the speech recognition engine analyzes the voice and generates text data such as "Please describe the dog."

[1004] Step 6:

[1005] The server analyzes the text data obtained from the speech recognition engine and passes it to the generation AI to extract relevant content.

[1006] Step 7:

[1007] The server's generation AI generates a 3D image based on text data. For example, based on the data for "dog," it generates a 3D model of a dog.

[1008] Step 8:

[1009] The server sends the generated 3D model data to the hologram generation engine. The data from the generation AI is passed to the conversion engine, which prepares the hologram data for creation.

[1010] Step 9:

[1011] The server's hologram generation engine converts the 3D stereoscopic image data into hologram data, specifically, converting the 3D model into a format suitable for holographic display.

[1012] Step 10:

[1013] The server sends the converted hologram data to the terminal via a communication interface.

[1014] Step 11:

[1015] The terminal receives the hologram data, temporarily stores it in a buffer, retrieves the data through the receiving interface, and starts processing it immediately.

[1016] Step 12:

[1017] The terminal initiates the process for transmitting hologram data to an assistive device, such as a projector or hologram display.

[1018] Step 13:

[1019] The assistive device uses the received holographic data to project a 3D stereoscopic image in real time. The holographic projection device interprets the data and displays a realistic stereoscopic image in space.

[1020] Each step allows the user to seamlessly experience the audio-to-hologram conversion process.

[1021] Example 1

[1022] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1023] Many modern learning and presentation systems provide information mainly through text or 2D images, but these do not adequately support visual comprehension. Furthermore, conventional voice input systems struggle to visually display the input content in real time, requiring numerous steps and devices. This complicates the user experience and makes intuitive operation difficult.

[1024] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1025] In this invention, the server includes a voice recognition unit that converts voice data into text, an analysis unit that analyzes the text information, a generation unit that generates a 3D stereoscopic image based on the text information, and a conversion unit that converts the 3D stereoscopic image into hologram data. This makes it possible to generate a 3D stereoscopic image in real time based on a user's voice input and display the image as a hologram.

[1026] The "input means" is a device that allows the user to input voice.

[1027] "Communication means" is a mechanism for transmitting voice data from a terminal to a server.

[1028] "Speech recognition means" refers to software or algorithms that convert voice data into text on the server.

[1029] "Analysis means" refers to the process or technology by which the server analyzes text information and extracts the necessary information.

[1030] The "generation means" is a technology that allows the server to generate a 3D stereoscopic image based on text information.

[1031] The "conversion means" is the technology that the server uses to convert the 3D stereoscopic image into hologram data.

[1032] "Display means" refers to a device or technology that allows a terminal to receive hologram data and project it in real time.

[1033] The "support device" is an external device for actually projecting and displaying hologram data.

[1034] This invention relates to a system that generates a 3D stereoscopic image in real time based on a user's voice input and displays it as a hologram. Specific embodiments of this system will be described below.

[1035] The system consists of the following main components:

[1036] 1. Audio input device (terminal)

[1037] 2. Communication Devices (Terminals)

[1038] 3. Speech recognition engine (server)

[1039] 4. Generative AI (server)

[1040] 5. Hologram generation engine (server)

[1041] 6. Projection Support Devices (Terminals)

[1042] Voice input

[1043] First, a user can operate the system by speaking into a voice input device. For example, they can speak commands such as "Tell me about the planets in the solar system." The voice input device captures the voice and generates audio data. Specific voice input devices can be built-in microphones or dedicated voice input devices.

[1044] Sending audio data

[1045] The collected audio data is then transmitted to a server via the communication device, using standard wireless communication technologies such as Wi-Fi or Bluetooth.

[1046] Voice Recognition

[1047] The server then passes the received voice data to a voice recognition engine for analysis. An example of a voice recognition engine is the Google Cloud Speech-to-Text API. This engine analyzes phonemes and phonology to convert the voice data into text, and generates the corresponding text. For example, a command such as "Tell me about the planets in the solar system" is generated as text data.

[1048] Text analysis and 3D image generation

[1049] Once the text data is generated, an analysis tool on the server analyzes the text data and extracts the necessary information. Natural language processing (NLP) technology is used as the analysis tool. Based on the results of this analysis, the server's generative AI (such as OpenAI's GPT-3) creates data for generating 3D stereoscopic images. For example, a series of 3D stereoscopic images (planet models) can be generated based on data on the planets in the solar system.

[1050] Hologram Generation

[1051] The generated 3D stereoscopic image data is passed to the server's hologram generation engine, which converts it into hologram data. The hologram generation engine can use, for example, Looking Glass Factory's hologram generation technology. This engine converts the 3D model data into images from multiple viewpoints and combines them to create hologram data.

[1052] Hologram display

[1053] The hologram data is sent from the server to the terminal and displayed. The terminal receives this data and passes it to a projection assistance device, specifically a hologram projector, which uses light interference patterns to display 3D images in space.

[1054] Specific examples

[1055] For example, consider the following scenario for an educational system:

[1056] User: "Tell me about the planets in our solar system."

[1057] The device captures the audio and sends it to the server. At this stage, the built-in microphone picks up the audio signal and sends it to the server as digital data.

[1058] The server's speech recognition engine converts the text to "Tell me about the planets in the solar system," using feature extraction and statistical models of the speech signal.

[1059] The server's NLP engine parses this text and understands it as a specific information request.

[1060] The server sends the generated AI a prompt saying, "Provide information about planets in the solar system."

[1061] Generative AI generates a series of 3D stereoscopic images based on data for each planet.

[1062] The server's hologram generation engine converts these 3D models into hologram data.

[1063] The terminal receives the hologram data and transmits it to the display device, using Wi-Fi or a wired connection.

[1064] The assistive device displays a real-time stereoscopic model of the planets in the solar system. The projector uses light patterns to create 3D images in space that can be seen from multiple perspectives.

[1065] In this way, users can visually study detailed 3D holograms in real time through voice input.

[1066] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1067] Step 1:

[1068] The user speaks to the voice input device to input commands to the system. For example, the user might say, "Please describe the dog." The input is the user's voice data.

[1069] Step 2:

[1070] The terminal captures the user's voice and collects the voice data. The microphone inside the device converts the voice signal into digital data and processes the data in real time. The input is the user's voice data, and the output is the digitized voice data.

[1071] Step 3:

[1072] The voice data collected by the device is sent to a server via a communication device. Specifically, the data is transferred using wireless communication protocols such as Wi-Fi or Bluetooth. The input is digitized voice data, and the output is the voice data sent to the server.

[1073] Step 4:

[1074] The server passes the received voice data to a voice recognition engine (voice recognition means) for analysis. The voice recognition engine (for example, general voice recognition software) analyzes the voice data and converts it into text information. The input is voice data and the output is text information.

[1075] Step 5:

[1076] The server's analysis means analyzes the generated text information and extracts the necessary information. Natural language processing (NLP) technology is used for this analysis, and specific keywords and command content are extracted from the analysis results. The input is text information, and the output is the extracted keywords and command data.

[1077] Step 6:

[1078] The server passes a prompt to the generation AI based on the extracted command data. For example, a prompt such as "Please describe the dog" is input to the generation AI. The input is the command data, and the output is the prompt.

[1079] Step 7:

[1080] Based on the prompt received, the generation AI generates relevant 3D stereoscopic image data. For example, the keyword "dog" generates 3D model data of a dog. The input is the prompt, and the output is 3D stereoscopic image data.

[1081] Step 8:

[1082] The server's hologram generation engine (conversion means) converts the generated 3D stereoscopic image data into hologram data. Specifically, it uses 3D model data to generate images from multiple viewpoints and then combines them to create hologram data. The input is 3D stereoscopic image data, and the output is hologram data.

[1083] Step 9:

[1084] The server transmits the converted hologram data to the terminal via Wi-Fi or a wired connection. The input is the hologram data, and the output is the hologram data transmitted to the terminal.

[1085] Step 10:

[1086] The terminal receives the hologram data and passes it to the support device. For example, it transmits the data to a hologram projector. The input is the hologram data, and the output is the hologram data passed to the support device.

[1087] Step 11:

[1088] The assistive device uses the hologram data to project a 3D image in real time. For example, a hologram projector displays a 3D image in space based on the data it receives. The input is hologram data, and the output is a 3D hologram displayed in space.

[1089] (Application example 1)

[1090] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1091] Currently, physical stores have limited means of providing detailed product information, making it difficult for customers to fully understand product features. Furthermore, there is a lack of efficient means for providing visual product information to multiple customers simultaneously. Therefore, there is a demand for technology that can easily display product information in real time using voice input to improve the user experience and stimulate purchasing motivation.

[1092] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1093] In this invention, the server includes an input means for a user to input voice, a communication means for a terminal to transmit voice data to the server, a voice recognition means for the server to convert the voice data into text, a generation means for the server to generate a 3D stereoscopic image using a generative artificial intelligence model based on the text information, a conversion means for the server to convert the 3D stereoscopic image into hologram data, and a display means for the terminal to receive the hologram data and project it in real time. This makes it possible to display interactive 3D holograms in physical stores that allow customers to intuitively understand detailed product information through voice commands.

[1094] "Input means for inputting voice" refers to a device or technology that allows a user to give instructions to a system via voice.

[1095] "Communication means for transmitting audio data to a server" refers to a method or device for sending audio data to a server via the Internet or other communication network.

[1096] "Speech recognition means for converting voice data into text" refers to the algorithms or software used to analyze voice data and convert it into text data.

[1097] "Generative AI model" refers to an AI model or system used to generate 3D stereoscopic images based on text information.

[1098] "Means for converting a 3D stereoscopic image into hologram data" refers to an algorithm or system for converting a 3D stereoscopic image into a data format that can be projected as a hologram.

[1099] "Display means for receiving holographic data and projecting it in real time" refers to a device or technology for visually projecting the received holographic data in real time.

[1100] A "prompt" is a piece of text used to instruct a generative AI model to generate a particular output.

[1101] "Display device" is a general term for devices and hardware for visually displaying generated hologram data.

[1102] A "physical store" refers to a store located in a physical location where customers can visit in person to view and purchase products.

[1103] The present invention relates to a system that displays detailed product information in a physical store as a 3D hologram in real time. This system allows a user to input voice, and then projects a 3D stereoscopic image generated based on the voice as a hologram. Specific embodiments of this system are described below.

[1104] System configuration

[1105] The system consists of the following main components:

[1106] 1. Voice input means: A microphone for users to input voice.

[1107] 2. Communication method: The method by which the voice data is transmitted to the server via the Internet.

[1108] 3. Speech recognition means: A speech recognition engine (e.g., Google Speech-to-Text) that the server uses to convert voice data into text.

[1109] 4. Generation: The process in which the server uses a generative AI model (e.g., OpenAI's GPT-3) to generate a 3D stereoscopic image from text information.

[1110] 5. Conversion method: The algorithm used by the server to convert the 3D stereoscopic image into hologram data (e.g., Unity's hologram generation plug-in).

[1111] 6. Display means: A device that projects holographic data in real time using a smartphone or a holographic display (e.g., Looking Glass Factory's holographic display).

[1112] Processing flow

[1113] 1. Voice input: The user speaks into the smartphone, for example, "Please tell me the features of this new smartphone."

[1114] 2. Sending audio data: The smartphone sends the audio data to the server via an HTTP request over the internet.

[1115] 3. Speech recognition: The server passes the received voice data to a speech recognition engine and converts it into text. For example, the generated text is "Please tell me the features of this new smartphone."

[1116] 4. Text analysis and 3D image generation: The server uses generative AI models to generate relevant 3D images from text information, for example, a 3D model based on the detailed features and specifications of a smartphone.

[1117] 5. Hologram generation: The server's hologram generation engine analyzes the 3D stereoscopic image and converts it into data that can be displayed as a hologram.

[1118] 6. Hologram display: A smartphone receives hologram data and projects it in real time on a built-in hologram display, etc. This device can display interactive 3D images in real time.

[1119] Specific example explanation

[1120] Example 1:

[1121] When a user speaks to a smartphone sales counter in a physical store and says, "What are the features of this new smartphone?", the system instantly projects a 3D model of the smartphone and its features as a hologram, allowing the user to understand the details of the product without having to touch it.

[1122] Example prompt sentence:

[1123] User dictation: "What are the features of this new phone?"

[1124] Server prompt: "Please describe the features of the new smartphone XYZ. Generate a 3D model of it."

[1125] This system will significantly improve the customer experience in physical stores, increasing product understanding and purchasing intent.

[1126] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1127] Step 1:

[1128] A user uses the smartphone's voice input means to ask a product-related question. Specifically, the user says, "What are the features of this new smartphone?" The input is the user's voice, and the output is the voice data captured by the smartphone's microphone.

[1129] Step 2:

[1130] The device sends the captured audio data to the server via a communication means. The audio data is sent to the server as an HTTP request over the Internet. The input is the audio data captured by the smartphone's microphone, and the output is the audio data sent to the server.

[1131] Step 3:

[1132] The server receives the voice data and passes it to a voice recognition engine to convert it into text. Specifically, the server analyzes the received voice data and generates text data such as "Please tell me the features of this new smartphone." The input is voice data and the output is text data.

[1133] Step 4:

[1134] The server passes the generated text data to a generative AI model, which then generates a 3D stereoscopic image based on the text information. Specifically, the generative AI model generates a 3D model of a smartphone based on the prompt, "Features of the new smartphone." The input is text data, and the output is 3D stereoscopic image data.

[1135] Step 5:

[1136] The server passes the generated 3D stereoscopic image data to an algorithm for converting it into hologram data. Here, the server's hologram generation engine converts the 3D stereoscopic image into a data format that can be displayed as a hologram. The input is 3D stereoscopic image data, and the output is hologram data.

[1137] Step 6:

[1138] The device receives the hologram data from the server and projects it in real time using the built-in display means. Specifically, a smartphone or holographic display visually displays the hologram data. The input is the hologram data, and the output is the displayed hologram.

[1139] This will allow users to view interactive 3D holograms in physical stores that provide intuitive understanding of product details.

[1140] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1141] The system of the present invention generates 3D stereoscopic images in real time based on the user's voice input and displays them as holograms. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide a more interactive and personalized experience. Below, we will create a program for this system and explain its processing in detail.

[1142] overview

[1143] The system consists of the following main components:

[1144] 1. A voice input device (terminal) that inputs the user's voice

[1145] 2. A communication device (terminal) that transmits voice data to a server

[1146] 3. Speech recognition engine (server) that converts voice data into text

[1147] 4. Generative AI (server) that generates 3D images based on text information

[1148] 5. Hologram generation engine (server) that converts the generated 3D stereoscopic image into hologram data

[1149] 6. Support device (terminal) that receives and projects hologram data

[1150] 7. Emotion engine (server) that recognizes user emotions

[1151] Program processing flow

[1152] Voice input

[1153] 1. The user speaks into a voice input device, for example, "Please describe my dog."

[1154] 2. The terminal captures the user's voice and converts it into digital voice data. The voice input device collects the voice and performs pre-processing to convert it into a data format.

[1155] 3. Communication is established to send the audio data to the server.

[1156] Voice Recognition

[1157] 4. The server passes the received voice data to the speech recognition engine, which decodes the voice data and provides it as input to the speech recognition engine.

[1158] 5. The speech recognition engine converts the voice data into text data. For example, the text data generated is "Please describe the dog."

[1159] emotion recognition

[1160] 6. The server passes the voice data to the emotion engine, which analyzes the user's emotional state. The emotion engine analyzes the tone and tempo of the voice to identify the user's emotion (happiness, sadness, surprise, etc.).

[1161] 7. Provide the emotion engine's analysis results (user's emotional state) to the generation AI.

[1162] Text analysis and 3D image generation

[1163] 8. The server analyzes the generated text information and sentiment analysis results and passes them to the generation AI to extract relevant content.

[1164] 9. The generation AI generates a 3D image based on the input text and emotion information. For example, it generates a 3D model of a dog based on the keyword "dog," and adds movements and expressions that correspond to the user's emotions.

[1165] 10. The generated 3D stereoscopic image data is sent to the hologram generation engine.

[1166] Hologram Generation

[1167] 11. The server's hologram generation engine converts the 3D stereoscopic image data into hologram data, specifically converting the 3D model into a format suitable for holographic display.

[1168] 12. The converted hologram data is sent to the terminal.

[1169] Hologram display

[1170] 13. The terminal receives the hologram data and temporarily stores it in a buffer. It then retrieves the data through the receiving interface and starts processing it immediately.

[1171] 14. The terminal starts the process of transmitting hologram data to the assistive device, such as a projector or hologram display.

[1172] 15. The assistive device uses the received hologram data to project a 3D stereoscopic image in real time. The hologram projection device interprets the data and displays a realistic stereoscopic image in space.

[1173] Specific examples

[1174] Educational Demo Scenario

[1175] User: "Tell me about the planets in our solar system."

[1176] The device captures the user's voice and sends it to the server.

[1177] The server's speech recognition engine converts the received voice data into text such as "Tell me about the planets in the solar system."

[1178] The server's emotion engine analyzes the user's emotions from the voice data, and recognizes the emotion of joy.

[1179] The server's generation AI generates a series of 3D stereoscopic images (planet models) based on the data of each planet, adding bright colors and movement according to the user's pleasure.

[1180] The server's hologram generation engine converts these 3D models into hologram data.

[1181] The terminal receives the hologram data and transmits it to the display device.

[1182] The assistive device displays a 3D model of the planets in the solar system in real time.

[1183] The system allows users to enjoy a richer, more personalized holographic experience in real time, utilizing both voice input and emotion analysis.

[1184] The processing flow will be explained below.

[1185] Step 1:

[1186] A user speaks into a voice input device, for example, "Describe my dog."

[1187] Step 2:

[1188] The terminal captures the user's voice and converts it into digital voice data. The voice input device collects the voice and pre-processes it to convert it into a data format.

[1189] Step 3:

[1190] The device transmits the captured audio data to the server in real time. The audio data is sent to the server as a stream via the communication interface.

[1191] Step 4:

[1192] The server passes the received voice data to the speech recognition engine, which then decodes the voice data and provides it as input to the speech recognition engine.

[1193] Step 5:

[1194] The server's speech recognition engine converts the voice data into text data. Specifically, the speech recognition engine analyzes the voice and generates text data such as "Please describe the dog."

[1195] Step 6:

[1196] The server passes the voice data to the emotion engine, which analyzes the user's emotional state. The emotion engine analyzes the tone and tempo of the voice to identify the user's emotion (happiness, sadness, surprise, etc.).

[1197] Step 7:

[1198] The server's emotion engine provides the analysis results (user's emotional state) to the generation AI, which then performs the following processes based on the received emotional information.

[1199] Step 8:

[1200] The server passes the generated text information and emotion analysis results to the generation AI, which extracts relevant content. The generation AI then generates a 3D image based on the input text information and emotion information.

[1201] Step 9:

[1202] The server's AI generates 3D model data for a dog based on a keyword, such as "dog," and adds movements and facial expressions that correspond to the user's emotions.

[1203] Step 10:

[1204] The server sends the generated 3D model data to the hologram generation engine. The data from the generation AI is passed to the conversion engine, which prepares the hologram data for creation.

[1205] Step 11:

[1206] The server's hologram generation engine converts the 3D stereoscopic image data into hologram data, specifically, converting the 3D model into a format suitable for holographic display.

[1207] Step 12:

[1208] The server sends the converted hologram data to the terminal via a communication interface.

[1209] Step 13:

[1210] The terminal receives the hologram data, temporarily stores it in a buffer, retrieves the data through the receiving interface, and starts processing it immediately.

[1211] Step 14:

[1212] The terminal initiates the process for transmitting hologram data to an assistive device, such as a projector or hologram display.

[1213] Step 15:

[1214] The assistive device uses the received hologram data to project a 3D stereoscopic image in real time. The hologram projection device interprets the data and displays a realistic stereoscopic image in space.

[1215] Specific examples

[1216] Educational Demo Scenario

[1217] Step 1:

[1218] A user speaks into a voice input device, "Tell me about the planets in the solar system."

[1219] Step 2:

[1220] The device captures the user's voice and converts it into digital voice data.

[1221] Step 3:

[1222] The device transmits the captured audio data to the server in real time.

[1223] Step 4:

[1224] The server passes the received voice data to the voice recognition engine.

[1225] Step 5:

[1226] The server's voice recognition engine converts the voice data into text data such as "Tell me about the planets in the solar system."

[1227] Step 6:

[1228] The server passes the voice data to the emotion engine, which analyzes the user's emotional state (interest and pleasure).

[1229] Step 7:

[1230] The server's emotion engine provides the analysis results to the generation AI.

[1231] Step 8:

[1232] The server passes the generated text information and sentiment analysis results to the generation AI, which extracts relevant content.

[1233] Step 9:

[1234] The server's generation AI generates a series of 3D stereoscopic images (planet models) based on data for each planet, adding bright colors and movement according to the user's emotions.

[1235] Step 10:

[1236] The server sends the generated 3D model data to the hologram generation engine.

[1237] Step 11:

[1238] The server's hologram generation engine converts the generated 3D model data into hologram data.

[1239] Step 12:

[1240] The server transmits the converted hologram data to the terminal.

[1241] Step 13:

[1242] The terminal receives the hologram data and temporarily stores it in a buffer.

[1243] Step 14:

[1244] The terminal initiates the process for transmitting the hologram data to the assistive device.

[1245] Step 15:

[1246] The assistive device displays a 3D model of the planets in the solar system in real time.

[1247] Example 2

[1248] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1249] Conventional voice input systems simply recognize a user's voice, convert it into text, and display a static 3D stereoscopic image based on that text. This limited the user experience, making it difficult to provide an interactive and personalized experience. Furthermore, the lack of technology to analyze user emotions and generate content based on those emotions prevented users from achieving high levels of satisfaction.

[1250] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1251] In this invention, the server includes a means for converting voice data into text, a means for generating a 3D stereoscopic image based on the text information, a means for recognizing a user's emotion from the voice data, and a means for adjusting the movement and facial expression of the generated 3D stereoscopic image based on the user's emotion information, thereby making it possible to provide an interactive and personalized hologram experience in real time according to the user's emotion.

[1252] "Input means" refers to a device that allows a user to input voice.

[1253] "Communication means" refers to a network interface that enables a terminal to transmit audio data to a server.

[1254] "Speech recognition means" refers to the algorithm or engine that converts the voice data received by the server into text data.

[1255] "Generation means" refers to the algorithms or programs that the server uses to generate 3D stereoscopic images based on text information.

[1256] "Conversion means" refers to the algorithm or engine that the server uses to convert 3D stereoscopic images into hologram data.

[1257] "Display means" refers to a device for projecting hologram data received by the terminal in real time.

[1258] "Recognition means" refers to the algorithms or engines that the server uses to analyze the user's emotions from voice data.

[1259] "Adjustment means" refers to an algorithm or program that the generation means uses to adjust the movements and facial expressions of the 3D stereoscopic image based on the user's emotional information.

[1260] "Support equipment" refers to an external device for projecting a hologram.

[1261] This invention provides a system that generates 3D stereoscopic images in real time based on the user's voice input and displays them as holograms, and also recognizes the user's emotions and provides a personalized experience based on them.

[1262] Hardware and software used

[1263] Hardware:

[1264] Audio input devices: microphones, smartphones, etc.

[1265] Communication device: smartphone, tablet, or computer.

[1266] Server: High-performance cloud or on-premise servers.

[1267] Assistive devices: Hologram projector, VR / AR headset (e.g. Microsoft HoloLens).

[1268] software:

[1269] Speech recognition engine: Google Speech-to-Text API, Azure Speech Service, etc.

[1270] Emotion recognition engine: IBM Watson Tone Analyzer.

[1271] Generative AI: OpenAI GPT-3 or similar advanced language model.

[1272] Hologram generation engine: Unity or Unreal Engine.

[1273] Program processing explanation

[1274] 1. Voice Input

[1275] A user speaks into a voice input device, for example, "Please describe my dog."

[1276] 2. Communications

[1277] In order for the terminal to transmit the audio data to the server, the digital audio data is divided into packets and transmitted to the server via secure communication.

[1278] 3. Voice Recognition

[1279] The server passes the received voice data to a voice recognition engine, which converts the voice into text data.

[1280] 4. Emotion recognition

[1281] The server passes the voice data to an emotion recognition engine, which analyzes the user's emotional state. The emotion engine analyzes the tone and tempo of the voice to identify the user's emotion (happiness, sadness, surprise, etc.).

[1282] 5. Text Analysis and 3D Image Generation

[1283] The server inputs the generated text information and emotion analysis results into a generative AI model to generate appropriate 3D stereoscopic image data. For example, based on the text "dog" and emotion information, a 3D model of a dog is generated and its movements and facial expressions are adjusted according to the emotion.

[1284] 6. Hologram Generation

[1285] The server's hologram generation engine converts the 3D stereoscopic image data into hologram data. Specifically, it uses Unity or Unreal Engine to encode the 3D model into a format suitable for holograms.

[1286] 7. Data Transmission

[1287] The server transmits the converted hologram data to the terminal.

[1288] 8. Hologram display

[1289] The terminal receives the hologram data and transmits it to the support device, which then displays the hologram in real time.

[1290] Specific examples

[1291] Educational Demo Scenario

[1292] User: "Tell me about the planets in our solar system."

[1293] The device captures the user's voice, converts it into digital audio data, and sends it to the server.

[1294] The server's speech recognition engine converts the received voice data into text such as "Tell me about the planets in the solar system."

[1295] The server's emotion recognition engine analyzes the user's emotions from the voice and recognizes the emotion of joy.

[1296] The server's generative AI model generates a series of 3D stereoscopic images (planet models) based on data for each planet, adding bright colors and movement depending on the user's emotions.

[1297] The server's hologram generation engine converts these 3D models into hologram data.

[1298] The terminal receives the hologram data and transmits it to the support device.

[1299] The support device displays a 3D model of the planets in the solar system in real time.

[1300] Example prompt: "Tell me about the planets in the solar system."

[1301] The system allows users to utilize both voice input and emotion analysis to create a richer, more personalized holographic experience in real time.

[1302] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1303] Step 1:

[1304] The user speaks to the voice input device, saying, "Please describe my dog," and the voice input device captures the voice and converts it into digital voice data.

[1305] Input: User spoken words.

[1306] Output: Digital audio data.

[1307] What happens: An audio input device (e.g., a smartphone microphone) converts an analog audio signal into a digital signal, which is sampled at 44.1 kHz and recorded in PCM format.

[1308] Step 2:

[1309] The terminal sends voice data to the server using a communication method. The voice data is divided into packets and sent to the server using secure communication (SSL / TLS).

[1310] Input: Digital audio data.

[1311] Output: The audio data sent to the server.

[1312] What it does: The device splits the audio data into packets of 1024 bytes each and sends the series of packets to the server over an SSL / TLS connection.

[1313] Step 3:

[1314] The server receives the voice data and passes it to the voice recognition engine, which converts the received voice data into text data.

[1315] Input: Audio data sent to the server.

[1316] Output: Text data.

[1317] Specific operation: The server temporarily stores the voice data in storage and calls the Google Speech-to-Text API to convert the voice data into text data. For example, the text generated is "Please describe the dog."

[1318] Step 4:

[1319] The server passes the voice data to an emotion recognition engine, which analyzes the user's emotional state. The emotion recognition engine analyzes the tone and tempo of the voice to recognize the user's emotions.

[1320] Input: Audio data.

[1321] Output: Emotion analysis results.

[1322] How it works: The server calls IBM Watson Tone Analyzer, which analyzes the tone, pitch, speed, etc. of the voice data to identify emotions. For example, the voice may be judged to be "joyful," and the joy score may be high.

[1323] Step 5:

[1324] The server inputs text and emotion information into the generative AI model and obtains the generated 3D stereoscopic image data. The generative AI generates a 3D model based on the text and adds movements and facial expressions according to the emotion information.

[1325] Input: Text data, sentiment analysis results.

[1326] Output: 3D stereoscopic image data.

[1327] Specific operation: The server calls OpenAI GPT-3, inputs the text prompt "Please describe the dog" and the emotion "joy" as input, and generates 3D stereoscopic image data. For example, a joyful motion of a dog jumping up and down is added.

[1328] Step 6:

[1329] The server converts the 3D stereoscopic image into hologram data, which is then converted into a format suitable for holographic projection.

[1330] Input: 3D stereoscopic image data.

[1331] Output: Hologram data.

[1332] How it works: The server uses Unity or Unreal Engine to convert 3D stereoscopic image data into a volumetric hologram format, for example encoding a 3D model of a dog into a holographic format.

[1333] Step 7:

[1334] The server transmits hologram data to the terminal, which receives the data and transmits it to the support device.

[1335] Input: Hologram data.

[1336] Output: Hologram data sent to the device.

[1337] Specific operation: The server packetizes the hologram data and sends it to the device via an SSL / TLS connection. The device temporarily stores the received data in a buffer.

[1338] Step 8:

[1339] The device transmits the hologram data to the support device, which displays the hologram in real time. The support device interprets the hologram data and projects it into real space.

[1340] Input: Hologram data.

[1341] Output: The projected hologram.

[1342] How it works: The device uses Wi-Fi 6 (IEEE 802.11ax) to transmit hologram data to an assistive device (e.g., Microsoft HoloLens). The assistive device interprets the data and projects a real-time 3D hologram of a bouncing dog into space.

[1343] (Application example 2)

[1344] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1345] In today's brick-and-mortar stores, customers are limited to a single explanation method and fixed information when receiving product explanations and store guidance. It is also difficult to provide a personalized experience tailored to the customer's emotions and state, making it difficult to provide effective customer service. With such limited information provided, it is difficult to increase customer satisfaction and purchasing motivation.

[1346] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1347] In this invention, the server includes an emotion recognition means for analyzing voice data to recognize a user's emotion, a generation means for personalizing a 3D stereoscopic image based on the user's emotion, and a generation means for generating an interactive 3D stereoscopic image, thereby enabling a richer and more personalized hologram experience based on the user's voice input and emotion analysis.

[1348] The "voice input means" is a device that allows the user to input voice.

[1349] "Communication means" is a function that enables a terminal to transmit voice data to a server.

[1350] "Speech recognition means" is a function by which the server converts voice data into text.

[1351] The "generation means" is a function that allows the server to generate a 3D stereoscopic image based on text information.

[1352] The "conversion means" is a function by which the server converts a 3D stereoscopic image into hologram data.

[1353] "Emotion recognition means" is a function in which the server analyzes voice data and recognizes the user's emotions.

[1354] The "means for personalizing" is a function by which the generating means personalizes the 3D stereoscopic image based on the user's emotions.

[1355] "Display means" refers to the function of the terminal to receive hologram data and project it in real time.

[1356] As an embodiment of this invention, we will describe an interactive product explanation system for use in a brick-and-mortar store. This system provides a personalized experience by generating 3D stereoscopic images in real time based on the user's voice input and emotion analysis, and displaying the images as holograms.

[1357] System configuration

[1358] 1. Voice input means: The user uses a device to input voice. This can include a microphone on a smartphone or smart glasses.

[1359] 2. Communication means: The device has the ability to send captured audio data to the server.

[1360] 3. Speech recognition: The server uses a speech recognition engine to convert the voice data into text, for example, Google Cloud Speech-to-Text API.

[1361] 4. Emotion recognition: The server analyzes the voice data and uses an emotion engine to recognize the user's emotions, for example, IBM Watson Tone Analyzer.

[1362] 5. Generation method: The server uses a generative AI model, such as OpenAI GPT-4, to generate a 3D stereoscopic image based on text information and emotions.

[1363] 6. Conversion means: The server uses a hologram generation engine to convert the generated 3D stereoscopic image into hologram data.

[1364] 7. Display means: Includes a support device for the terminal to receive holographic data and project it in real time, such as a holographic display from Looking Glass Factory.

[1365] Program processing

[1366] The system does the following:

[1367] First, the user uses the voice input means to speak about the product or information they want to hear explained about. For example, they might say, "Please explain the features of this product." At that time, the voice input means captures the user's voice and converts it into digital voice data. This data is then sent to the server via the communication means.

[1368] The voice data that arrives at the server is converted into text data using a voice recognition means. In parallel, an emotion recognition means analyzes the voice data and recognizes the user's emotions. The user's emotional state (e.g., curiosity or surprise) is used to further refine the interaction.

[1369] Next, the text information and emotion information are passed to a generation means, which uses a generative AI model to generate a 3D stereoscopic image. The generated 3D stereoscopic image is further personalized by applying movements and facial expressions according to the user's emotions. The 3D stereoscopic image is then converted into hologram data by a conversion means.

[1370] Finally, the hologram data is transmitted to the support device through the display means and projected in real time in the physical store, allowing users to visualize the products and information as interactive holograms right in front of their eyes.

[1371] Specific examples

[1372] For example, suppose a user uses a voice input method in a store to say, "Please explain the features of this product." The voice is immediately sent to the server, where speech recognition and emotion analysis are performed. As a result, the text data "Please explain the features of this product" and the user's emotional state (curiosity) are acquired. Based on the text and emotion information, the generative AI model generates an easy-to-understand and interesting 3D stereoscopic image corresponding to the "curiosity." For example, this could include an animation or exploded view showing the internal structure of the product.

[1373] The generated 3D stereoscopic image is converted into hologram data and projected as a hologram in real time through a display device. Users can visually observe the hologram and manipulate it interactively.

[1374] Example prompt sentence:

[1375] "Please explain. Analyze the user's emotions and generate a 3D model based on those emotions. Now, the user said, 'Explain the features of this product.' The user's emotion is 'Curiosity.'"

[1376] This allows users to enjoy a more personalized experience in physical stores.

[1377] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1378] Program processing

[1379] Step 1:

[1380] The user uses a voice input means to talk about the product or information they want to hear explained. For example, they might say, "Please explain the features of this product." In this case, the input is the user's voice data. This voice data is captured by the microphone of a smartphone or smart glasses and converted into digital voice data. The output is digitized voice data.

[1381] Step 2:

[1382] The terminal transmits the captured voice data to the server via a communication means. At this time, the input is the captured digital voice data. The server receives this data and prepares to pass it to the voice recognition engine. The output is the result of transmitting the voice data to the server.

[1383] Step 3:

[1384] The server converts the voice data into text data using a speech recognition tool. The input is digital voice data. A speech recognition engine (e.g., Google Cloud Speech-to-Text) analyzes the data and generates text data such as "Describe the features of this product." The output is text data.

[1385] Step 4:

[1386] At the same time, the server passes the voice data to an emotion recognition means to analyze the user's emotions. The input is digital voice data. The emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the tone and tempo of the voice to identify the user's emotional state (e.g., data expressing curiosity or surprise). The output is the user's emotional data.

[1387] Step 5:

[1388] The server passes the generated text information and emotion analysis results to the generation means. At this time, the inputs are text data and emotion data. The generative AI model (e.g., OpenAI GPT-4) generates a 3D stereoscopic image based on the text and emotion information, corresponding to the user's emotion. The output is 3D stereoscopic image data.

[1389] Step 6:

[1390] The server passes the generated 3D stereoscopic image data to a conversion means, which converts it into hologram data. At this time, the input is the 3D stereoscopic image data. The hologram generation engine converts the 3D model data into a format suitable for hologram display. The output is hologram data.

[1391] Step 7:

[1392] The server sends the converted hologram data to the terminal. At this time, the input is the hologram data. The terminal receives the data through the receiving interface and temporarily stores it in a buffer. The output is the data storage result on the terminal.

[1393] Step 8:

[1394] The terminal transmits hologram data to the assistive device, which projects it in real time. The input is the hologram data. The assistive device (e.g., Looking Glass Factory's holographic display) interprets the received data and projects a realistic 3D image in space. The output is the user's visual holographic experience.

[1395] Specific operation example

[1396] For example, a user might say, "Explain the features of this product" using a voice input method while in a store. The voice is immediately sent to the server, where speech recognition and emotion analysis are performed. As a result, the text data "Explain the features of this product" and the user's emotional state (curiosity) are acquired. Based on the text and emotion information, a generative AI model generates an easy-to-understand and interesting 3D stereoscopic image that corresponds to the "curiosity." For example, this could include animations or exploded views that show the internal structure of the product.

[1397] Example prompt sentence:

[1398] "Please explain. Analyze the user's emotions and generate a 3D model based on those emotions. Now, the user said, 'Explain the features of this product.' The user's emotion is 'Curiosity.'"

[1399] In this way, a personalized and interactive hologram experience can be provided based on the user's voice input and emotion analysis.

[1400] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1401] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1402] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1403] [Fourth embodiment]

[1404] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1405] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1406] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1407] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1408] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1409] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1410] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1411] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1412] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1413] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1414] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1415] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1416] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1417] The system of the present invention generates a 3D stereoscopic image in real time based on the user's voice input and displays it as a hologram. Below, we will create a program for this system and explain its processing in detail.

[1418] overview

[1419] The system consists of the following main components:

[1420] 1. A voice input device (terminal) that inputs the user's voice

[1421] 2. A communication device (terminal) that transmits voice data to a server

[1422] 3. Speech recognition engine (server) that converts voice data into text

[1423] 4. Generative AI (server) that generates 3D images based on text information

[1424] 5. Hologram generation engine (server) that converts the generated 3D stereoscopic image into hologram data

[1425] 6. Support device (terminal) that receives and projects hologram data

[1426] Program processing flow

[1427] Voice input

[1428] 1. The user speaks into a voice input device, giving a command such as "Describe the dog."

[1429] 2. The device captures the user's voice and collects voice data in real time.

[1430] 3. Communication is established to send the audio data to the server.

[1431] Voice Recognition

[1432] 4. The server passes the received voice data to the voice recognition engine.

[1433] 5. The speech recognition engine analyzes the voice data and converts it into text information. For example, the text data generated is "Please describe the dog."

[1434] Text analysis and 3D image generation

[1435] 6. The server analyzes the generated text information and passes it to the generation AI to extract relevant content.

[1436] 7. Generative AI generates a 3D image based on the input text information. For example, it creates a 3D model of a dog based on the keyword "dog."

[1437] 8. The generated 3D stereoscopic image data is sent to the hologram generation engine.

[1438] Hologram Generation

[1439] 9. The server's hologram generation engine converts the 3D image data into hologram data.

[1440] 10. The converted hologram data is sent to the terminal.

[1441] Hologram display

[1442] 11. The terminal receives the hologram data and transmits it to the support device.

[1443] 12. The assistive device uses the holographic data to project a 3D image in real time.

[1444] Specific examples

[1445] Educational Demo Scenario

[1446] User: "Tell me about the planets in our solar system."

[1447] The device captures the audio and sends it to the server.

[1448] The server's speech recognition engine converts the received voice data into text such as "Tell me about the planets in the solar system."

[1449] The server's generation AI generates a series of 3D stereoscopic images (planet models) based on the data of each planet.

[1450] The server's hologram generation engine converts these 3D models into hologram data.

[1451] The terminal receives the hologram data and transmits it to the display device.

[1452] The assistive device displays a 3D model of the planets in the solar system in real time.

[1453] This system allows users to learn visually through holograms in real time using only voice input, and is expected to be used in a variety of situations, not just education, but also entertainment and business presentations.

[1454] The processing flow will be explained below.

[1455] Step 1:

[1456] A user speaks into a voice input device, for example, "Describe my dog."

[1457] Step 2:

[1458] The terminal captures the user's voice and converts it into digital voice data. The voice input device collects the voice and pre-processes it to convert it into a data format.

[1459] Step 3:

[1460] The device transmits the captured audio data to the server in real time. The audio data is sent to the server as a stream via the communication interface.

[1461] Step 4:

[1462] The server passes the received voice data to the speech recognition engine, which then decodes the voice data and provides it as input to the speech recognition engine.

[1463] Step 5:

[1464] The server's speech recognition engine converts the voice data into text data. Specifically, the speech recognition engine analyzes the voice and generates text data such as "Please describe the dog."

[1465] Step 6:

[1466] The server analyzes the text data obtained from the speech recognition engine and passes it to the generation AI to extract relevant content.

[1467] Step 7:

[1468] The server's generation AI generates a 3D image based on text data. For example, based on the data for "dog," it generates a 3D model of a dog.

[1469] Step 8:

[1470] The server sends the generated 3D model data to the hologram generation engine. The data from the generation AI is passed to the conversion engine, which prepares the hologram data for creation.

[1471] Step 9:

[1472] The server's hologram generation engine converts the 3D stereoscopic image data into hologram data, specifically, converting the 3D model into a format suitable for holographic display.

[1473] Step 10:

[1474] The server sends the converted hologram data to the terminal via a communication interface.

[1475] Step 11:

[1476] The terminal receives the hologram data, temporarily stores it in a buffer, retrieves the data through the receiving interface, and starts processing it immediately.

[1477] Step 12:

[1478] The terminal initiates the process for transmitting hologram data to an assistive device, such as a projector or hologram display.

[1479] Step 13:

[1480] The assistive device uses the received holographic data to project a 3D stereoscopic image in real time. The holographic projection device interprets the data and displays a realistic stereoscopic image in space.

[1481] Each step allows the user to seamlessly experience the audio-to-hologram conversion process.

[1482] Example 1

[1483] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1484] Many modern learning and presentation systems provide information mainly through text or 2D images, but these do not adequately support visual comprehension. Furthermore, conventional voice input systems struggle to visually display the input content in real time, requiring numerous steps and devices. This complicates the user experience and makes intuitive operation difficult.

[1485] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1486] In this invention, the server includes a voice recognition unit that converts voice data into text, an analysis unit that analyzes the text information, a generation unit that generates a 3D stereoscopic image based on the text information, and a conversion unit that converts the 3D stereoscopic image into hologram data. This makes it possible to generate a 3D stereoscopic image in real time based on a user's voice input and display the image as a hologram.

[1487] The "input means" is a device that allows the user to input voice.

[1488] "Communication means" is a mechanism for transmitting voice data from a terminal to a server.

[1489] "Speech recognition means" refers to software or algorithms that convert voice data into text on the server.

[1490] "Analysis means" refers to the process or technology by which the server analyzes text information and extracts the necessary information.

[1491] The "generation means" is a technology that allows the server to generate a 3D stereoscopic image based on text information.

[1492] The "conversion means" is the technology that the server uses to convert the 3D stereoscopic image into hologram data.

[1493] "Display means" refers to a device or technology that allows a terminal to receive hologram data and project it in real time.

[1494] The "support device" is an external device for actually projecting and displaying hologram data.

[1495] This invention relates to a system that generates a 3D stereoscopic image in real time based on a user's voice input and displays it as a hologram. Specific embodiments of this system will be described below.

[1496] The system consists of the following main components:

[1497] 1. Audio input device (terminal)

[1498] 2. Communication Devices (Terminals)

[1499] 3. Speech recognition engine (server)

[1500] 4. Generative AI (server)

[1501] 5. Hologram generation engine (server)

[1502] 6. Projection Support Devices (Terminals)

[1503] Voice input

[1504] First, a user can operate the system by speaking into a voice input device. For example, they can speak commands such as "Tell me about the planets in the solar system." The voice input device captures the voice and generates audio data. Specific voice input devices can be built-in microphones or dedicated voice input devices.

[1505] Sending audio data

[1506] The collected audio data is then transmitted to a server via the communication device, using standard wireless communication technologies such as Wi-Fi or Bluetooth.

[1507] Voice Recognition

[1508] The server then passes the received voice data to a voice recognition engine for analysis. An example of a voice recognition engine is the Google Cloud Speech-to-Text API. This engine analyzes phonemes and phonology to convert the voice data into text, and generates the corresponding text. For example, a command such as "Tell me about the planets in the solar system" is generated as text data.

[1509] Text analysis and 3D image generation

[1510] Once the text data is generated, an analysis tool on the server analyzes the text data and extracts the necessary information. Natural language processing (NLP) technology is used as the analysis tool. Based on the results of this analysis, the server's generative AI (such as OpenAI's GPT-3) creates data for generating 3D stereoscopic images. For example, a series of 3D stereoscopic images (planet models) can be generated based on data on the planets in the solar system.

[1511] Hologram Generation

[1512] The generated 3D stereoscopic image data is passed to the server's hologram generation engine, which converts it into hologram data. The hologram generation engine can use, for example, Looking Glass Factory's hologram generation technology. This engine converts the 3D model data into images from multiple viewpoints and combines them to create hologram data.

[1513] Hologram display

[1514] The hologram data is sent from the server to the terminal and displayed. The terminal receives this data and passes it to a projection assistance device, specifically a hologram projector, which uses light interference patterns to display 3D images in space.

[1515] Specific examples

[1516] For example, consider the following scenario for an educational system:

[1517] User: "Tell me about the planets in our solar system."

[1518] The device captures the audio and sends it to the server. At this stage, the built-in microphone picks up the audio signal and sends it to the server as digital data.

[1519] The server's speech recognition engine converts the text to "Tell me about the planets in the solar system," using feature extraction and statistical models of the speech signal.

[1520] The server's NLP engine parses this text and understands it as a specific information request.

[1521] The server sends the generated AI a prompt saying, "Provide information about planets in the solar system."

[1522] Generative AI generates a series of 3D stereoscopic images based on data for each planet.

[1523] The server's hologram generation engine converts these 3D models into hologram data.

[1524] The terminal receives the hologram data and transmits it to the display device, using Wi-Fi or a wired connection.

[1525] The assistive device displays a real-time stereoscopic model of the planets in the solar system. The projector uses light patterns to create 3D images in space that can be seen from multiple perspectives.

[1526] In this way, users can visually study detailed 3D holograms in real time through voice input.

[1527] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1528] Step 1:

[1529] The user speaks to the voice input device to input commands to the system. For example, the user might say, "Please describe the dog." The input is the user's voice data.

[1530] Step 2:

[1531] The terminal captures the user's voice and collects the voice data. The microphone inside the device converts the voice signal into digital data and processes the data in real time. The input is the user's voice data, and the output is the digitized voice data.

[1532] Step 3:

[1533] The voice data collected by the device is sent to a server via a communication device. Specifically, the data is transferred using wireless communication protocols such as Wi-Fi or Bluetooth. The input is digitized voice data, and the output is the voice data sent to the server.

[1534] Step 4:

[1535] The server passes the received voice data to a voice recognition engine (voice recognition means) for analysis. The voice recognition engine (for example, general voice recognition software) analyzes the voice data and converts it into text information. The input is voice data and the output is text information.

[1536] Step 5:

[1537] The server's analysis means analyzes the generated text information and extracts the necessary information. Natural language processing (NLP) technology is used for this analysis, and specific keywords and command content are extracted from the analysis results. The input is text information, and the output is the extracted keywords and command data.

[1538] Step 6:

[1539] The server passes a prompt to the generation AI based on the extracted command data. For example, a prompt such as "Please describe the dog" is input to the generation AI. The input is the command data, and the output is the prompt.

[1540] Step 7:

[1541] Based on the prompt received, the generation AI generates relevant 3D stereoscopic image data. For example, the keyword "dog" generates 3D model data of a dog. The input is the prompt, and the output is 3D stereoscopic image data.

[1542] Step 8:

[1543] The server's hologram generation engine (conversion means) converts the generated 3D stereoscopic image data into hologram data. Specifically, it uses 3D model data to generate images from multiple viewpoints and then combines them to create hologram data. The input is 3D stereoscopic image data, and the output is hologram data.

[1544] Step 9:

[1545] The server transmits the converted hologram data to the terminal via Wi-Fi or a wired connection. The input is the hologram data, and the output is the hologram data transmitted to the terminal.

[1546] Step 10:

[1547] The terminal receives the hologram data and passes it to the support device. For example, it transmits the data to a hologram projector. The input is the hologram data, and the output is the hologram data passed to the support device.

[1548] Step 11:

[1549] The assistive device uses the hologram data to project a 3D image in real time. For example, a hologram projector displays a 3D image in space based on the data it receives. The input is hologram data, and the output is a 3D hologram displayed in space.

[1550] (Application example 1)

[1551] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1552] Currently, physical stores have limited means of providing detailed product information, making it difficult for customers to fully understand product features. Furthermore, there is a lack of efficient means for providing visual product information to multiple customers simultaneously. Therefore, there is a demand for technology that can easily display product information in real time using voice input to improve the user experience and stimulate purchasing motivation.

[1553] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1554] In this invention, the server includes an input means for a user to input voice, a communication means for a terminal to transmit voice data to the server, a voice recognition means for the server to convert the voice data into text, a generation means for the server to generate a 3D stereoscopic image using a generative artificial intelligence model based on the text information, a conversion means for the server to convert the 3D stereoscopic image into hologram data, and a display means for the terminal to receive the hologram data and project it in real time. This makes it possible to display interactive 3D holograms in physical stores that allow customers to intuitively understand detailed product information through voice commands.

[1555] "Input means for inputting voice" refers to a device or technology that allows a user to give instructions to a system via voice.

[1556] "Communication means for transmitting audio data to a server" refers to a method or device for sending audio data to a server via the Internet or other communication network.

[1557] "Speech recognition means for converting voice data into text" refers to the algorithms or software used to analyze voice data and convert it into text data.

[1558] "Generative AI model" refers to an AI model or system used to generate 3D stereoscopic images based on text information.

[1559] "Means for converting a 3D stereoscopic image into hologram data" refers to an algorithm or system for converting a 3D stereoscopic image into a data format that can be projected as a hologram.

[1560] "Display means for receiving holographic data and projecting it in real time" refers to a device or technology for visually projecting the received holographic data in real time.

[1561] A "prompt" is a piece of text used to instruct a generative AI model to generate a particular output.

[1562] "Display device" is a general term for devices and hardware for visually displaying generated hologram data.

[1563] A "physical store" refers to a store located in a physical location where customers can visit in person to view and purchase products.

[1564] The present invention relates to a system that displays detailed product information in a physical store as a 3D hologram in real time. This system allows a user to input voice, and then projects a 3D stereoscopic image generated based on the voice as a hologram. Specific embodiments of this system are described below.

[1565] System configuration

[1566] The system consists of the following main components:

[1567] 1. Voice input means: A microphone for users to input voice.

[1568] 2. Communication method: The method by which the voice data is transmitted to the server via the Internet.

[1569] 3. Speech recognition means: A speech recognition engine (e.g., Google Speech-to-Text) that the server uses to convert voice data into text.

[1570] 4. Generation: The process in which the server uses a generative AI model (e.g., OpenAI's GPT-3) to generate a 3D stereoscopic image from text information.

[1571] 5. Conversion method: The algorithm used by the server to convert the 3D stereoscopic image into hologram data (e.g., Unity's hologram generation plug-in).

[1572] 6. Display means: A device that projects holographic data in real time using a smartphone or a holographic display (e.g., Looking Glass Factory's holographic display).

[1573] Processing flow

[1574] 1. Voice input: The user speaks into the smartphone, for example, "Please tell me the features of this new smartphone."

[1575] 2. Sending audio data: The smartphone sends the audio data to the server via an HTTP request over the internet.

[1576] 3. Speech recognition: The server passes the received voice data to a speech recognition engine and converts it into text. For example, the generated text is "Please tell me the features of this new smartphone."

[1577] 4. Text analysis and 3D image generation: The server uses generative AI models to generate relevant 3D images from text information, for example, a 3D model based on the detailed features and specifications of a smartphone.

[1578] 5. Hologram generation: The server's hologram generation engine analyzes the 3D stereoscopic image and converts it into data that can be displayed as a hologram.

[1579] 6. Hologram display: A smartphone receives hologram data and projects it in real time on a built-in hologram display, etc. This device can display interactive 3D images in real time.

[1580] Specific example explanation

[1581] Example 1:

[1582] When a user speaks to a smartphone sales counter in a physical store and says, "What are the features of this new smartphone?", the system instantly projects a 3D model of the smartphone and its features as a hologram, allowing the user to understand the details of the product without having to touch it.

[1583] Example prompt sentence:

[1584] User dictation: "What are the features of this new phone?"

[1585] Server prompt: "Please describe the features of the new smartphone XYZ. Generate a 3D model of it."

[1586] This system will significantly improve the customer experience in physical stores, increasing product understanding and purchasing intent.

[1587] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1588] Step 1:

[1589] A user uses the smartphone's voice input means to ask a product-related question. Specifically, the user says, "What are the features of this new smartphone?" The input is the user's voice, and the output is the voice data captured by the smartphone's microphone.

[1590] Step 2:

[1591] The device sends the captured audio data to the server via a communication means. The audio data is sent to the server as an HTTP request over the Internet. The input is the audio data captured by the smartphone's microphone, and the output is the audio data sent to the server.

[1592] Step 3:

[1593] The server receives the voice data and passes it to a voice recognition engine to convert it into text. Specifically, the server analyzes the received voice data and generates text data such as "Please tell me the features of this new smartphone." The input is voice data and the output is text data.

[1594] Step 4:

[1595] The server passes the generated text data to a generative AI model, which then generates a 3D stereoscopic image based on the text information. Specifically, the generative AI model generates a 3D model of a smartphone based on the prompt, "Features of the new smartphone." The input is text data, and the output is 3D stereoscopic image data.

[1596] Step 5:

[1597] The server passes the generated 3D stereoscopic image data to an algorithm for converting it into hologram data. Here, the server's hologram generation engine converts the 3D stereoscopic image into a data format that can be displayed as a hologram. The input is 3D stereoscopic image data, and the output is hologram data.

[1598] Step 6:

[1599] The device receives the hologram data from the server and projects it in real time using the built-in display means. Specifically, a smartphone or holographic display visually displays the hologram data. The input is the hologram data, and the output is the displayed hologram.

[1600] This will allow users to view interactive 3D holograms in physical stores that provide intuitive understanding of product details.

[1601] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1602] The system of the present invention generates 3D stereoscopic images in real time based on the user's voice input and displays them as holograms. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide a more interactive and personalized experience. Below, we will create a program for this system and explain its processing in detail.

[1603] overview

[1604] The system consists of the following main components:

[1605] 1. A voice input device (terminal) that inputs the user's voice

[1606] 2. A communication device (terminal) that transmits voice data to a server

[1607] 3. Speech recognition engine (server) that converts voice data into text

[1608] 4. Generative AI (server) that generates 3D images based on text information

[1609] 5. Hologram generation engine (server) that converts the generated 3D stereoscopic image into hologram data

[1610] 6. Support device (terminal) that receives and projects hologram data

[1611] 7. Emotion engine (server) that recognizes user emotions

[1612] Program processing flow

[1613] Voice input

[1614] 1. The user speaks into a voice input device, for example, "Please describe my dog."

[1615] 2. The terminal captures the user's voice and converts it into digital voice data. The voice input device collects the voice and performs pre-processing to convert it into a data format.

[1616] 3. Communication is established to send the audio data to the server.

[1617] Voice Recognition

[1618] 4. The server passes the received voice data to the speech recognition engine, which decodes the voice data and provides it as input to the speech recognition engine.

[1619] 5. The speech recognition engine converts the voice data into text data. For example, the text data generated is "Please describe the dog."

[1620] emotion recognition

[1621] 6. The server passes the voice data to the emotion engine, which analyzes the user's emotional state. The emotion engine analyzes the tone and tempo of the voice to identify the user's emotion (happiness, sadness, surprise, etc.).

[1622] 7. Provide the emotion engine's analysis results (user's emotional state) to the generation AI.

[1623] Text analysis and 3D image generation

[1624] 8. The server analyzes the generated text information and sentiment analysis results and passes them to the generation AI to extract relevant content.

[1625] 9. The generation AI generates a 3D image based on the input text and emotion information. For example, it generates a 3D model of a dog based on the keyword "dog," and adds movements and expressions that correspond to the user's emotions.

[1626] 10. The generated 3D stereoscopic image data is sent to the hologram generation engine.

[1627] Hologram Generation

[1628] 11. The server's hologram generation engine converts the 3D stereoscopic image data into hologram data, specifically converting the 3D model into a format suitable for holographic display.

[1629] 12. The converted hologram data is sent to the terminal.

[1630] Hologram display

[1631] 13. The terminal receives the hologram data and temporarily stores it in a buffer. It then retrieves the data through the receiving interface and starts processing it immediately.

[1632] 14. The terminal starts the process of transmitting hologram data to the assistive device, such as a projector or hologram display.

[1633] 15. The assistive device uses the received hologram data to project a 3D stereoscopic image in real time. The hologram projection device interprets the data and displays a realistic stereoscopic image in space.

[1634] Specific examples

[1635] Educational Demo Scenario

[1636] User: "Tell me about the planets in our solar system."

[1637] The device captures the user's voice and sends it to the server.

[1638] The server's speech recognition engine converts the received voice data into text such as "Tell me about the planets in the solar system."

[1639] The server's emotion engine analyzes the user's emotions from the voice data, and recognizes the emotion of joy.

[1640] The server's generation AI generates a series of 3D stereoscopic images (planet models) based on the data of each planet, adding bright colors and movement according to the user's pleasure.

[1641] The server's hologram generation engine converts these 3D models into hologram data.

[1642] The terminal receives the hologram data and transmits it to the display device.

[1643] The assistive device displays a 3D model of the planets in the solar system in real time.

[1644] The system allows users to enjoy a richer, more personalized holographic experience in real time, utilizing both voice input and emotion analysis.

[1645] The processing flow will be explained below.

[1646] Step 1:

[1647] A user speaks into a voice input device, for example, "Describe my dog."

[1648] Step 2:

[1649] The terminal captures the user's voice and converts it into digital voice data. The voice input device collects the voice and pre-processes it to convert it into a data format.

[1650] Step 3:

[1651] The device transmits the captured audio data to the server in real time. The audio data is sent to the server as a stream via the communication interface.

[1652] Step 4:

[1653] The server passes the received voice data to the speech recognition engine, which then decodes the voice data and provides it as input to the speech recognition engine.

[1654] Step 5:

[1655] The server's speech recognition engine converts the voice data into text data. Specifically, the speech recognition engine analyzes the voice and generates text data such as "Please describe the dog."

[1656] Step 6:

[1657] The server passes the voice data to the emotion engine, which analyzes the user's emotional state. The emotion engine analyzes the tone and tempo of the voice to identify the user's emotion (happiness, sadness, surprise, etc.).

[1658] Step 7:

[1659] The server's emotion engine provides the analysis results (user's emotional state) to the generation AI, which then performs the following processes based on the received emotional information.

[1660] Step 8:

[1661] The server passes the generated text information and emotion analysis results to the generation AI, which extracts relevant content. The generation AI then generates a 3D image based on the input text information and emotion information.

[1662] Step 9:

[1663] The server's AI generates 3D model data for a dog based on a keyword, such as "dog," and adds movements and facial expressions that correspond to the user's emotions.

[1664] Step 10:

[1665] The server sends the generated 3D model data to the hologram generation engine. The data from the generation AI is passed to the conversion engine, which prepares the hologram data for creation.

[1666] Step 11:

[1667] The server's hologram generation engine converts the 3D stereoscopic image data into hologram data, specifically, converting the 3D model into a format suitable for holographic display.

[1668] Step 12:

[1669] The server sends the converted hologram data to the terminal via a communication interface.

[1670] Step 13:

[1671] The terminal receives the hologram data, temporarily stores it in a buffer, retrieves the data through the receiving interface, and starts processing it immediately.

[1672] Step 14:

[1673] The terminal initiates the process for transmitting hologram data to an assistive device, such as a projector or hologram display.

[1674] Step 15:

[1675] The assistive device uses the received holographic data to project a 3D stereoscopic image in real time. The holographic projection device interprets the data and displays a realistic stereoscopic image in space.

[1676] Specific examples

[1677] Educational Demo Scenario

[1678] Step 1:

[1679] A user speaks into a voice input device, "Tell me about the planets in the solar system."

[1680] Step 2:

[1681] The device captures the user's voice and converts it into digital voice data.

[1682] Step 3:

[1683] The device transmits the captured audio data to the server in real time.

[1684] Step 4:

[1685] The server passes the received voice data to the voice recognition engine.

[1686] Step 5:

[1687] The server's voice recognition engine converts the voice data into text data such as "Tell me about the planets in the solar system."

[1688] Step 6:

[1689] The server passes the voice data to the emotion engine, which analyzes the user's emotional state (interest and pleasure).

[1690] Step 7:

[1691] The server's emotion engine provides the analysis results to the generation AI.

[1692] Step 8:

[1693] The server passes the generated text information and sentiment analysis results to the generation AI, which extracts relevant content.

[1694] Step 9:

[1695] The server's generation AI generates a series of 3D stereoscopic images (planet models) based on data for each planet, adding bright colors and movement according to the user's emotions.

[1696] Step 10:

[1697] The server sends the generated 3D model data to the hologram generation engine.

[1698] Step 11:

[1699] The server's hologram generation engine converts the generated 3D model data into hologram data.

[1700] Step 12:

[1701] The server transmits the converted hologram data to the terminal.

[1702] Step 13:

[1703] The terminal receives the hologram data and temporarily stores it in a buffer.

[1704] Step 14:

[1705] The terminal initiates the process for transmitting the hologram data to the assistive device.

[1706] Step 15:

[1707] The assistive device displays a 3D model of the planets in the solar system in real time.

[1708] Example 2

[1709] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1710] Conventional voice input systems simply recognize a user's voice, convert it into text, and display a static 3D stereoscopic image based on that text. This limited the user experience, making it difficult to provide an interactive and personalized experience. Furthermore, the lack of technology to analyze user emotions and generate content based on those emotions prevented users from achieving high levels of satisfaction.

[1711] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1712] In this invention, the server includes a means for converting voice data into text, a means for generating a 3D stereoscopic image based on the text information, a means for recognizing a user's emotion from the voice data, and a means for adjusting the movement and facial expression of the generated 3D stereoscopic image based on the user's emotion information, thereby making it possible to provide an interactive and personalized hologram experience in real time according to the user's emotion.

[1713] "Input means" refers to a device that allows a user to input voice.

[1714] "Communication means" refers to a network interface that enables a terminal to transmit audio data to a server.

[1715] "Speech recognition means" refers to the algorithm or engine that converts the voice data received by the server into text data.

[1716] "Generation means" refers to the algorithms or programs that the server uses to generate 3D stereoscopic images based on text information.

[1717] "Conversion means" refers to the algorithm or engine that the server uses to convert 3D stereoscopic images into hologram data.

[1718] "Display means" refers to a device for projecting hologram data received by the terminal in real time.

[1719] "Recognition means" refers to the algorithms or engines that the server uses to analyze the user's emotions from voice data.

[1720] "Adjustment means" refers to an algorithm or program that the generation means uses to adjust the movements and facial expressions of the 3D stereoscopic image based on the user's emotional information.

[1721] "Support equipment" refers to an external device for projecting a hologram.

[1722] This invention provides a system that generates 3D stereoscopic images in real time based on the user's voice input and displays them as holograms, and also recognizes the user's emotions and provides a personalized experience based on them.

[1723] Hardware and software used

[1724] Hardware:

[1725] Audio input devices: microphones, smartphones, etc.

[1726] Communication device: smartphone, tablet, or computer.

[1727] Server: High-performance cloud or on-premise servers.

[1728] Assistive devices: Hologram projector, VR / AR headset (e.g. Microsoft HoloLens).

[1729] software:

[1730] Speech recognition engine: Google Speech-to-Text API, Azure Speech Service, etc.

[1731] Emotion recognition engine: IBM Watson Tone Analyzer.

[1732] Generative AI: OpenAI GPT-3 or similar advanced language model.

[1733] Hologram generation engine: Unity or Unreal Engine.

[1734] Program processing explanation

[1735] 1. Voice Input

[1736] A user speaks into a voice input device, for example, "Please describe my dog."

[1737] 2. Communications

[1738] In order for the terminal to transmit the audio data to the server, the digital audio data is divided into packets and transmitted to the server via secure communication.

[1739] 3. Voice Recognition

[1740] The server passes the received voice data to a voice recognition engine, which converts the voice into text data.

[1741] 4. Emotion recognition

[1742] The server passes the voice data to an emotion recognition engine, which analyzes the user's emotional state. The emotion engine analyzes the tone and tempo of the voice to identify the user's emotion (happiness, sadness, surprise, etc.).

[1743] 5. Text Analysis and 3D Image Generation

[1744] The server inputs the generated text information and emotion analysis results into a generative AI model to generate appropriate 3D stereoscopic image data. For example, based on the text "dog" and emotion information, a 3D model of a dog is generated and its movements and facial expressions are adjusted according to the emotion.

[1745] 6. Hologram Generation

[1746] The server's hologram generation engine converts the 3D stereoscopic image data into hologram data. Specifically, it uses Unity or Unreal Engine to encode the 3D model into a format suitable for holograms.

[1747] 7. Data Transmission

[1748] The server transmits the converted hologram data to the terminal.

[1749] 8. Hologram display

[1750] The terminal receives the hologram data and transmits it to the support device, which then displays the hologram in real time.

[1751] Specific examples

[1752] Educational Demo Scenario

[1753] User: "Tell me about the planets in our solar system."

[1754] The device captures the user's voice, converts it into digital audio data, and sends it to the server.

[1755] The server's speech recognition engine converts the received voice data into text such as "Tell me about the planets in the solar system."

[1756] The server's emotion recognition engine analyzes the user's emotions from the voice and recognizes the emotion of joy.

[1757] The server's generative AI model generates a series of 3D stereoscopic images (planet models) based on data for each planet, adding bright colors and movement depending on the user's emotions.

[1758] The server's hologram generation engine converts these 3D models into hologram data.

[1759] The terminal receives the hologram data and transmits it to the support device.

[1760] The support device displays a 3D model of the planets in the solar system in real time.

[1761] Example prompt: "Tell me about the planets in the solar system."

[1762] The system allows users to enjoy a richer, more personalized holographic experience in real time, utilizing both voice input and emotion analysis.

[1763] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1764] Step 1:

[1765] The user speaks to the voice input device, saying, "Please describe my dog," and the voice input device captures the voice and converts it into digital voice data.

[1766] Input: User spoken words.

[1767] Output: Digital audio data.

[1768] What happens: An audio input device (e.g., a smartphone microphone) converts an analog audio signal into a digital signal, which is sampled at 44.1 kHz and recorded in PCM format.

[1769] Step 2:

[1770] The terminal sends voice data to the server using a communication method. The voice data is divided into packets and sent to the server using secure communication (SSL / TLS).

[1771] Input: Digital audio data.

[1772] Output: The audio data sent to the server.

[1773] What it does: The device splits the audio data into packets of 1024 bytes each and sends the series of packets to the server over an SSL / TLS connection.

[1774] Step 3:

[1775] The server receives the voice data and passes it to the voice recognition engine, which converts the received voice data into text data.

[1776] Input: Audio data sent to the server.

[1777] Output: Text data.

[1778] Specific operation: The server temporarily stores the voice data in storage and calls the Google Speech-to-Text API to convert the voice data into text data. For example, the text generated is "Please describe the dog."

[1779] Step 4:

[1780] The server passes the voice data to an emotion recognition engine, which analyzes the user's emotional state. The emotion recognition engine analyzes the tone and tempo of the voice to recognize the user's emotions.

[1781] Input: Audio data.

[1782] Output: Emotion analysis results.

[1783] How it works: The server calls IBM Watson Tone Analyzer, which analyzes the tone, pitch, speed, etc. of the voice data to identify emotions. For example, the voice may be judged to be "joyful," and the joy score may be high.

[1784] Step 5:

[1785] The server inputs text and emotion information into the generative AI model and obtains the generated 3D stereoscopic image data. The generative AI generates a 3D model based on the text and adds movements and facial expressions according to the emotion information.

[1786] Input: Text data, sentiment analysis results.

[1787] Output: 3D stereoscopic image data.

[1788] Specific operation: The server calls OpenAI GPT-3, inputs the text prompt "Please describe the dog" and the emotion "joy" as input, and generates 3D stereoscopic image data. For example, a joyful motion of a dog jumping up and down is added.

[1789] Step 6:

[1790] The server converts the 3D stereoscopic image into hologram data, which is then converted into a format suitable for holographic projection.

[1791] Input: 3D stereoscopic image data.

[1792] Output: Hologram data.

[1793] How it works: The server uses Unity or Unreal Engine to convert 3D stereoscopic image data into a volumetric hologram format, for example encoding a 3D model of a dog into a holographic format.

[1794] Step 7:

[1795] The server transmits hologram data to the terminal, which receives the data and transmits it to the support device.

[1796] Input: Hologram data.

[1797] Output: Hologram data sent to the device.

[1798] Specific operation: The server packetizes the hologram data and sends it to the device via an SSL / TLS connection. The device temporarily stores the received data in a buffer.

[1799] Step 8:

[1800] The device transmits the hologram data to the support device, which displays the hologram in real time. The support device interprets the hologram data and projects it into real space.

[1801] Input: Hologram data.

[1802] Output: The projected hologram.

[1803] How it works: The device uses Wi-Fi 6 (IEEE 802.11ax) to transmit hologram data to an assistive device (e.g., Microsoft HoloLens). The assistive device interprets the data and projects a real-time 3D hologram of a bouncing dog into space.

[1804] (Application example 2)

[1805] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1806] In today's brick-and-mortar stores, customers are limited to a single explanation method and fixed information when receiving product explanations and store guidance. It is also difficult to provide a personalized experience tailored to the customer's emotions and state, making it difficult to provide effective customer service. With such limited information provided, it is difficult to increase customer satisfaction and purchasing motivation.

[1807] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1808] In this invention, the server includes an emotion recognition means for analyzing voice data to recognize a user's emotion, a generation means for personalizing a 3D stereoscopic image based on the user's emotion, and a generation means for generating an interactive 3D stereoscopic image, thereby enabling a richer and more personalized hologram experience based on the user's voice input and emotion analysis.

[1809] The "voice input means" is a device that allows the user to input voice.

[1810] "Communication means" is a function that enables a terminal to transmit voice data to a server.

[1811] "Speech recognition means" is a function by which the server converts voice data into text.

[1812] The "generation means" is a function that allows the server to generate a 3D stereoscopic image based on text information.

[1813] The "conversion means" is a function by which the server converts a 3D stereoscopic image into hologram data.

[1814] "Emotion recognition means" is a function in which the server analyzes voice data and recognizes the user's emotions.

[1815] The "means for personalizing" is a function by which the generating means personalizes the 3D stereoscopic image based on the user's emotions.

[1816] "Display means" refers to the function of the terminal to receive hologram data and project it in real time.

[1817] As an embodiment of this invention, we will describe an interactive product explanation system for use in a brick-and-mortar store. This system provides a personalized experience by generating 3D stereoscopic images in real time based on the user's voice input and emotion analysis, and displaying the images as holograms.

[1818] System configuration

[1819] 1. Voice input means: The user uses a device to input voice. This can include a microphone on a smartphone or smart glasses.

[1820] 2. Communication means: The device has the ability to send captured audio data to the server.

[1821] 3. Speech recognition: The server uses a speech recognition engine to convert the voice data into text, for example, Google Cloud Speech-to-Text API.

[1822] 4. Emotion recognition: The server analyzes the voice data and uses an emotion engine to recognize the user's emotions, for example, IBM Watson Tone Analyzer.

[1823] 5. Generation method: The server uses a generative AI model, such as OpenAI GPT-4, to generate a 3D stereoscopic image based on text information and emotions.

[1824] 6. Conversion means: The server uses a hologram generation engine to convert the generated 3D stereoscopic image into hologram data.

[1825] 7. Display means: Includes a support device for the terminal to receive holographic data and project it in real time, such as a holographic display from Looking Glass Factory.

[1826] Program processing

[1827] The system does the following:

[1828] First, the user uses the voice input means to speak about the product or information they want to hear explained about. For example, they might say, "Please explain the features of this product." At that time, the voice input means captures the user's voice and converts it into digital voice data. This data is then sent to the server via the communication means.

[1829] The voice data that arrives at the server is converted into text data using a voice recognition means. In parallel, an emotion recognition means analyzes the voice data and recognizes the user's emotions. The user's emotional state (e.g., curiosity or surprise) is used to further refine the interaction.

[1830] Next, the text information and emotion information are passed to a generation means, which uses a generative AI model to generate a 3D stereoscopic image. The generated 3D stereoscopic image is further personalized by applying movements and facial expressions according to the user's emotions. The 3D stereoscopic image is then converted into hologram data by a conversion means.

[1831] Finally, the hologram data is transmitted to the support device through the display means and projected in real time in the physical store, allowing users to visualize the products and information as interactive holograms right in front of their eyes.

[1832] Specific examples

[1833] For example, suppose a user uses a voice input method in a store to say, "Please explain the features of this product." The voice is immediately sent to the server, where speech recognition and emotion analysis are performed. As a result, the text data "Please explain the features of this product" and the user's emotional state (curiosity) are acquired. Based on the text and emotion information, the generative AI model generates an easy-to-understand and interesting 3D stereoscopic image corresponding to the "curiosity." For example, this could include an animation or exploded view showing the internal structure of the product.

[1834] The generated 3D stereoscopic image is converted into hologram data and projected as a hologram in real time through a display device. Users can visually observe the hologram and manipulate it interactively.

[1835] Example prompt sentence:

[1836] "Please explain. Analyze the user's emotions and generate a 3D model based on those emotions. Now, the user said, 'Explain the features of this product.' The user's emotion is 'Curiosity.'"

[1837] This allows users to enjoy a more personalized experience in-store.

[1838] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1839] Program processing

[1840] Step 1:

[1841] The user uses a voice input means to talk about the product or information they want to hear explained. For example, they might say, "Please explain the features of this product." In this case, the input is the user's voice data. This voice data is captured by the microphone of a smartphone or smart glasses and converted into digital voice data. The output is digitized voice data.

[1842] Step 2:

[1843] The terminal transmits the captured voice data to the server via a communication means. At this time, the input is the captured digital voice data. The server receives this data and prepares to pass it to the voice recognition engine. The output is the result of transmitting the voice data to the server.

[1844] Step 3:

[1845] The server converts the voice data into text data using a speech recognition tool. The input is digital voice data. A speech recognition engine (e.g., Google Cloud Speech-to-Text) analyzes the data and generates text data such as "Describe the features of this product." The output is text data.

[1846] Step 4:

[1847] At the same time, the server passes the voice data to an emotion recognition means to analyze the user's emotions. The input is digital voice data. The emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the tone and tempo of the voice to identify the user's emotional state (e.g., data expressing curiosity or surprise). The output is the user's emotional data.

[1848] Step 5:

[1849] The server passes the generated text information and emotion analysis results to the generation means. At this time, the inputs are text data and emotion data. The generative AI model (e.g., OpenAI GPT-4) generates a 3D stereoscopic image based on the text and emotion information, corresponding to the user's emotion. The output is 3D stereoscopic image data.

[1850] Step 6:

[1851] The server passes the generated 3D stereoscopic image data to a conversion means, which converts it into hologram data. At this time, the input is the 3D stereoscopic image data. The hologram generation engine converts the 3D model data into a format suitable for hologram display. The output is hologram data.

[1852] Step 7:

[1853] The server sends the converted hologram data to the terminal. At this time, the input is the hologram data. The terminal receives the data through the receiving interface and temporarily stores it in a buffer. The output is the data storage result on the terminal.

[1854] Step 8:

[1855] The terminal transmits hologram data to the assistive device, which projects it in real time. The input is the hologram data. The assistive device (e.g., Looking Glass Factory's holographic display) interprets the received data and projects a realistic 3D image in space. The output is the user's visual holographic experience.

[1856] Specific operation example

[1857] For example, a user might say, "Explain the features of this product" using a voice input method while in a store. The voice is immediately sent to the server, where speech recognition and emotion analysis are performed. As a result, the text data "Explain the features of this product" and the user's emotional state (curiosity) are acquired. Based on the text and emotion information, a generative AI model generates an easy-to-understand and interesting 3D stereoscopic image that corresponds to the "curiosity." For example, this could include animations or exploded views that show the internal structure of the product.

[1858] Example prompt sentence:

[1859] "Please explain. Analyze the user's emotions and generate a 3D model based on those emotions. Now, the user said, 'Explain the features of this product.' The user's emotion is 'Curiosity.'"

[1860] In this way, a personalized and interactive hologram experience can be provided based on the user's voice input and emotion analysis.

[1861] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1862] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1863] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1864] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1865] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1866] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1867] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1868] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1869] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1870] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1871] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1872] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1873] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1874] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1875] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1876] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1877] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1878] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1879] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1880] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1881] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1882] The following is further disclosed regarding the above embodiment.

[1883] (Claim 1)

[1884] an input means for a user to input voice;

[1885] A communication means for the terminal to transmit voice data to a server;

[1886] a speech recognition means for converting speech data into text by the server;

[1887] A generating means for generating a 3D stereoscopic image based on text information by the server;

[1888] A conversion means for converting the 3D stereoscopic image into hologram data by the server;

[1889] A system in which a terminal receives holographic data and includes a display means for projecting it in real time.

[1890] (Claim 2)

[1891] 10. The system of claim 1, wherein the generating means generates an interactive 3D image based on text information.

[1892] (Claim 3)

[1893] 2. The system according to claim 1, wherein the display means is compatible with a plurality of support devices.

[1894] "Example 1"

[1895] (Claim 1)

[1896] an input means for a user to input voice;

[1897] A communication means for the terminal to transmit voice data to a server;

[1898] a speech recognition means for converting speech data into text by the server;

[1899] an analysis means for the server to analyze the text information;

[1900] A generating means for generating a 3D stereoscopic image based on text information by the server;

[1901] A conversion means for converting the 3D stereoscopic image into hologram data by the server;

[1902] A system in which a terminal receives holographic data and includes a display means for projecting it in real time.

[1903] (Claim 2)

[1904] 2. The system of claim 1, wherein the generating means generates 3D stereoscopic image data based on the text information and passes it to the hologram generating engine.

[1905] (Claim 3)

[1906] 10. The system of claim 1, wherein the display means passes the hologram data to a supporting device for projection.

[1907] "Application Example 1"

[1908] (Claim 1)

[1909] an input means for a user to input voice;

[1910] A communication means for the terminal to transmit voice data to a server;

[1911] a speech recognition means for converting speech data into text by the server;

[1912] A generating means for generating a 3D stereoscopic image using a generating artificial intelligence model based on text information by the server;

[1913] A conversion means for converting the 3D stereoscopic image into hologram data by the server;

[1914] A system in which a terminal receives holographic data and includes a display means for projecting it in real time.

[1915] (Claim 2)

[1916] 10. The system of claim 1, wherein the generating means generates an interactive 3D image based on text information and further parses the information using prompt sentences.

[1917] (Claim 3)

[1918] 2. The system according to claim 1, wherein the display means is compatible with a plurality of display devices and is particularly applicable to a device for providing product information in a physical store.

[1919] "Example 2: Combining Emotion Engines"

[1920] (Claim 1)

[1921] an input means for a user to input voice;

[1922] A communication means for the terminal to transmit voice data to a server;

[1923] a speech recognition means for converting speech data into text by the server;

[1924] A generating means for generating a 3D stereoscopic image based on text information by the server;

[1925] A conversion means for converting the 3D stereoscopic image into hologram data by the server;

[1926] a display means for receiving the hologram data and projecting it in real time;

[1927] A recognition means for the server to recognize the user's emotion from the voice data;

[1928] The system includes an adjustment means for adjusting the movements and facial expressions of the 3D stereoscopic image based on emotional information.

[1929] (Claim 2)

[1930] 2. The system of claim 1, wherein the generating means generates an interactive 3D image based on text information and emotion information.

[1931] (Claim 3)

[1932] 2. The system according to claim 1, wherein the display means is compatible with a plurality of support devices.

[1933] "Application example 2 when combining emotion engines"

[1934] (Claim 1)

[1935] an input means for a user to input voice;

[1936] A communication means for the terminal to transmit voice data to a server;

[1937] a speech recognition means for converting speech data into text by the server;

[1938] A generating means for generating a 3D stereoscopic image based on text information by the server;

[1939] A conversion means for converting the 3D stereoscopic image into hologram data by the server;

[1940] An emotion recognition means for the server to analyze the voice data and recognize the emotion of the user;

[1941] The generating means personalizes the 3D stereoscopic image based on the user's emotion;

[1942] A system in which a terminal receives holographic data and includes a display means for projecting it in real time.

[1943] (Claim 2)

[1944] 10. The system of claim 1, wherein the generating means generates an interactive 3D stereoscopic image based on text information.

[1945] (Claim 3)

[1946] 2. The system according to claim 1, wherein the display means is compatible with a plurality of display devices. [Explanation of symbols]

[1947] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. an input means for a user to input voice; A communication means for the terminal to transmit voice data to a server; a speech recognition means for converting speech data into text by the server; A generating means for generating a 3D stereoscopic image based on text information by the server; A conversion means for converting the 3D stereoscopic image into hologram data by the server; A system in which a terminal receives holographic data and includes a display means for projecting it in real time.

2. 10. The system of claim 1, wherein the generating means generates an interactive 3D image based on text information.

3. 2. The system according to claim 1, wherein the display means is compatible with a plurality of support devices.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A