system

The system addresses the lack of engaging entertainment by using AR glasses to transform daily environments with virtual characters and interiors, enhancing user experience through real-time object replacement.

JP2026036168APending Publication Date: 2026-03-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024138683
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Conventional entertainment methods fail to adequately alleviate stress and boredom in daily life, particularly during commuting or working from home, due to limited ways to refresh oneself.

Method used

A system that uses AR glasses to receive voice instructions, convert them into text, capture real-world images, identify objects, select replacement data based on user input, and overlay virtual characters or interiors onto the real-world environment in real-time.

Benefits of technology

Enhances daily life experiences by transforming the real world into an entertaining virtual environment, providing new stimulation and enjoyment through the integration of virtual characters and interiors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026036168000001_ABST
    Figure 2026036168000001_ABST
Patent Text Reader

Abstract

Provide a system. A means for receiving a voice instruction from a user; means for converting the acquired voice instruction into text data and analyzing the instruction content; A means for capturing camera images, identifying real-world objects, and acquiring their position and orientation information; A means for selecting replacement data based on a designated character or interior; a means for generating selected replacement data using the position and orientation information of the object; a means for displaying the generated data by superimposing it on a camera image; A system including:
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, routine tasks such as daily life and commuting cause stress and boredom for many people. Conventional entertainment methods are not sufficient to fully alleviate these problems. Since there are limited ways to refresh oneself, especially while commuting or working from home, there is a need to improve people's quality of life. This invention aims to transform the real world into a highly entertaining virtual world, providing new stimulation and enjoyment to everyday life. [Means for solving the problem]

[0005] The system of this invention includes means for receiving voice instructions from a user, means for converting the received voice instructions into text data and analyzing the instructions, means for capturing camera images to identify real-world objects and obtain their position and orientation information, means for selecting replacement data based on the specified character or interior, means for generating the selected replacement data using the object's position and orientation information, and means for displaying the generated data by overlaying it on the camera image. This allows users to transform their own environment into their favorite characters and scenery, even while commuting or at home, thereby improving the quality of their daily lives.

[0006] "User" refers to the person who operates the system and issues voice commands.

[0007] "Voice instructions" refer to verbal instructions given by the user to the system.

[0008] The term "means" refers to elements that constitute various functions and devices included in the system of the present invention.

[0009] "Device" refers to the device responsible for receiving voice commands and capturing camera images. In this context, this typically refers to AR glasses.

[0010] "Server" refers to the central processing unit that analyzes voice instructions, performs image recognition, and generates AR data.

[0011] "Text data" refers to character information obtained as a result of analyzing voice instructions.

[0012] "Camera footage" refers to real-world video data acquired in real time from a user's device.

[0013] An "object" is an object that exists in the real world and needs to be recognized (such as a person, vehicle, or piece of furniture).

[0014] "Location information" is information that indicates the spatial position of an object recognized in a camera image.

[0015] "Posture information" is information that indicates the spatial posture, such as the direction and angle, of an object recognized in a camera image.

[0016] A "character" is a virtual entity (such as an anime character) that is displayed based on the user's instructions.

[0017] "Interior" refers to the visual decoration of the room, such as the walls and furniture, and is subject to change based on the user's instructions.

[0018] "Replacement data" is data used to replace real-world objects with virtual characters or interiors.

[0019] "Generation" refers to creating new data based on the position and orientation information of an object.

[0020] "Overlay" refers to overlaying virtual data onto camera footage.

[0021] "Real-time" means immediate processing with minimal delay. [Brief explanation of the drawings]

[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0024] First, the terms used in the following description will be explained.

[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0030] [First embodiment]

[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0043] This invention is a system that uses AR glasses worn by the user to replace real-world objects with virtual characters and interiors. This system allows users to enjoy their daily commute and their home life even more. The program processing is explained below in natural language.

[0044] Program processing explanation

[0045] 1. Recognition of voice commands

[0046] The user issues voice commands to the AR glasses. Possible commands include "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street on my commute into characters." The device captures the voice input and converts it into text data. The text data is then sent to a server for analysis.

[0047] 2. Parsing the instructions

[0048] The server analyzes the acquired text data. It uses natural language processing technology to understand the meaning of the instructions and identify which objects to replace. This analysis clarifies the "target (people, room interior, etc.)" and "changes (characters, Swiss lodge style, etc.)."

[0049] 3. Video Capture and Object Recognition

[0050] The device uses a camera to capture images within the user's field of view. This video data is sent in real time to a server, which uses image recognition technology to identify real-world objects (such as pedestrians, vehicles, and furniture) and obtain the position and orientation information of each object.

[0051] 4. Selection of replacement data

[0052] Based on the parsed instructions, the server selects replacement character and interior data, which is retrieved from a library containing the necessary 3D models and animation data.

[0053] 5. AR Data Generation

[0054] The server generates AR data based on the acquired object position and orientation information, placing the selected character and interior data in the real world at the appropriate position and orientation. This data also includes information necessary for real-time rendering, such as lighting and shadows.

[0055] 6. Displaying Data

[0056] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image. From the user's perspective, this makes it appear as if real-world objects have been replaced with the specified characters or interior.

[0057] Specific examples

[0058] Example 1: Change the room interior to Swiss lodge style

[0059] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[0060] 2. The device converts the voice into text data and sends it to the server.

[0061] 3. The server analyzes the data and identifies instructions for interior modifications.

[0062] 4. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[0063] 5. The server selects Swiss lodge-style interior data.

[0064] 6. The server combines the interior data with the room's location information to generate AR data.

[0065] 7. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. To the user, the interior of the room appears to have been transformed into a Swiss lodge.

[0066] Example 2: Change passersby in the city into characters during your commute

[0067] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[0068] 2. The device converts the voice into text data and sends it to the server.

[0069] 3. The server analyzes the data and identifies the passerby's replacement instructions.

[0070] 4. The device's camera captures images of the city, and the server recognizes passersby.

[0071] 5. The server selects the character data.

[0072] 6. The server combines the character data with the location information of passersby to generate AR data.

[0073] 7. The server sends the generated AR data to the device, which then displays it over the camera image. To the user, it appears as if passersby on their way to work have been transformed into characters.

[0074] This system allows users to enjoy a variety of entertainment experiences in their daily lives, improving their quality of life.

[0075] The processing flow will be explained below.

[0076] Step 1:

[0077] The user issues voice commands to the AR glasses, such as "Turn people on the street into characters on my commute."

[0078] Step 2:

[0079] The device recognizes voice commands and converts the voice data into digital form. An internal voice recognition engine analyzes this digital voice data and converts it into text data.

[0080] Step 3:

[0081] The terminal transmits the converted text data to the server, which includes the content of the voice instructions.

[0082] Step 4:

[0083] The server analyzes the received text data, using natural language processing technology to understand the meaning of the voice commands and extract the "target (people, room interior, etc.)" and "change content (character, Swiss lodge style, etc.)."

[0084] Step 5:

[0085] The device's camera captures images of the real world from the user's point of view, and the image data is sent to a server in real time.

[0086] Step 6:

[0087] The server analyzes the received video data and uses image recognition technology to identify objects in the video (passersby, vehicles, furniture, etc.), and identifies the position and posture information of each object and stores it in a database.

[0088] Step 7:

[0089] Based on the analysis results, the server selects the requested character and interior data, and retrieves the corresponding 3D model and animation data from the library.

[0090] Step 8:

[0091] The server uses the position and orientation information of the identified objects to generate AR data for displaying the selected characters and interiors in the real world, including the necessary rendering information.

[0092] Step 9:

[0093] The server sends the generated AR data to the device.

[0094] Step 10:

[0095] The device processes the AR data received in real time, and selected characters and interior designs are overlaid on the camera image.

[0096] Step 11:

[0097] When the user observes the real world through the AR glasses, the characters and interiors that have been changed as instructed are superimposed on the real world, allowing the user to experience passersby on their commute or the interior of their room as if they were replaced with virtual characters and designs.

[0098] Example 1

[0099] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0100] Systems that provide users with new entertainment experiences by replacing real-world objects with virtual characters and interiors require technology that can accurately understand user instructions and reflect them in real-time video. Furthermore, high accuracy and speed are required when systems integrate real-world video with virtual data, and technology is needed to provide users with a natural, seamless experience.

[0101] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0102] In this invention, the server includes means for receiving voice instructions from a user, means for converting the received voice instructions into text data and analyzing the instructions, means for capturing camera images, identifying real-world objects and obtaining position and orientation information, means for selecting replacement data based on the specified virtual character or interior, means for generating the selected replacement data using the object's position and orientation information, and means for displaying the generated data superimposed on the camera image. This makes it possible to quickly and accurately analyze the user's voice instructions and display virtual data that is highly accurately integrated with real-world images in real time.

[0103] "User" refers to an individual who uses the system to issue voice commands and enjoy the experience of replacing real-world objects with virtual characters and interiors.

[0104] "Voice instruction" refers to a command made by voice that a user issues to request some action or change to the system.

[0105] "Text data" refers to character string information that is generated by capturing voice instructions and converting them using voice recognition software.

[0106] "Camera footage" refers to real-world video data captured by a camera built into a device.

[0107] "Real-world objects" refer to physical objects that are within the user's field of view (e.g., pedestrians, vehicles, furniture, etc.).

[0108] "Position and orientation information" refers to information about the spatial position of an object in the real world, as well as its direction and orientation.

[0109] "Virtual character" refers to a computer-generated character that appears in place of a real-world object.

[0110] "Interior" refers to virtual data about the decoration and design of the space the user is in.

[0111] "Replacement Data" refers to information, including 3D models and animation data, used to replace real-world objects with virtual characters or interiors.

[0112] "Generated data" refers to AR data that is generated by the server based on the object's position and orientation information, and is used to apply virtual characters and interiors to the real world.

[0113] "Means for overlaying and displaying on camera images" refers to the method or technology by which the device receives the generated data, synthesizes it on camera images from the real world, and displays it to the user.

[0114] "Lens technology" refers to the optical techniques and devices used to integrate real-world images with generated data.

[0115] This invention is a system that uses AR glasses worn by the user to display real-world objects replaced with virtual characters and interiors. This system allows users to enjoy their daily commute and home life even more. Below, we will explain the specific program processing method and the hardware and software used.

[0116] System Overview

[0117] This system is primarily comprised of three components: a server, a device (including AR glasses), and a user. It combines various technologies to realize the process in which the user issues voice commands, real-world images are processed according to those commands, and virtual data is synthesized and displayed.

[0118] Hardware and Software

[0119] Device: AR glasses have built-in high-performance microphones and cameras, as well as a processor and necessary memory to process data in real time.

[0120] Server: A server with high processing power that uses natural language processing technology (e.g., GPT-4 (registered trademark) by OpenAI (registered trademark)) and image recognition technology (e.g., YOLO and OpenCV).

[0121] Speech Recognition Software: Speech recognition technology such as Google® Cloud Speech-to-Text is used to convert voice instructions into text data.

[0122] Rendering engine: An engine such as Unity or Unreal Engine that integrates real-world images with virtual data for display.

[0123] Specific processing of the program

[0124] Recognizing voice commands

[0125] The user issues voice commands into the AR glasses. For example, they can give specific instructions such as "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street to characters on my commute." The device uses a built-in microphone to capture the voice and obtains the voice data in real time. The obtained voice data is then converted into text data using voice recognition software.

[0126] Analysis of voice instructions

[0127] The device sends the converted text data to a server, which then uses natural language processing technology to analyze the text data. From the analysis results, the server identifies the "subject" (people, room interior, etc.) and the "changes" (characters, Swiss lodge style, etc.).

[0128] Video Capture and Object Recognition

[0129] The device uses a built-in camera to capture images within the user's field of view. This video data is sent in real time to a server, which uses image recognition technology to identify real-world objects (such as pedestrians, vehicles, and furniture) and obtain the position and orientation information of each object.

[0130] Selection of replacement data

[0131] Based on the parsed instructions, the server selects replacement character and interior data, which is retrieved from a library containing the necessary 3D models and animation data.

[0132] AR data generation

[0133] The server generates AR data based on the acquired object position and orientation information, placing the selected character and interior data in the real world at the appropriate position and orientation. This data also includes information necessary for real-time rendering, such as lighting and shadows.

[0134] Viewing Data

[0135] The generated AR data is sent from the server to the device, which processes the received AR data in real time and displays it overlaid on the camera image. Through the AR glasses, the user can see that real-world objects have been replaced with the specified characters or interior.

[0136] Specific examples

[0137] Example 1: Change the room interior to Swiss lodge style

[0138] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[0139] 2. The device converts the voice into text data and sends it to the server.

[0140] 3. The server analyzes the data and identifies instructions for interior modifications.

[0141] 4. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[0142] 5. The server selects Swiss lodge-style interior data.

[0143] 6. The server combines the interior data with the room's location information to generate AR data.

[0144] 7. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. To the user, the interior of the room appears to have been transformed into a Swiss lodge.

[0145] Example 2: Change passersby in the city into characters during your commute

[0146] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[0147] 2. The device converts the voice into text data and sends it to the server.

[0148] 3. The server analyzes the data and identifies the passerby's replacement instructions.

[0149] 4. The device's camera captures images of the city, and the server recognizes passersby.

[0150] 5. The server selects the character data.

[0151] 6. The server combines the character data with the location information of passersby to generate AR data.

[0152] 7. The server sends the generated AR data to the device, which then displays it over the camera image. To the user, it appears as if passersby on their way to work have been transformed into characters.

[0153] Prompt Sentence Examples

[0154] 1. "How can I transform my room into a Swiss lodge?"

[0155] 2. "How do I direct the people on the street to become characters?"

[0156] This system allows users to enjoy a variety of entertainment experiences in their daily lives, improving their quality of life.

[0157] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0158] Step 1: Getting voice instructions

[0159] The user issues voice commands to the AR glasses. For example, they can give specific instructions such as "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street to characters on my commute." The device captures the voice using a built-in microphone. The input is the user's voice commands, and the output is voice data.

[0160] Step 2: Convert the audio data

[0161] The device converts the captured voice data into text data using voice recognition software (e.g., Google Cloud Speech-to-Text). It uses voice data as input and natural language processing technology to output text data. This conversion process is performed in real time.

[0162] Step 3: Analyzing voice commands

[0163] The device sends the converted text data to the server. The server uses natural language processing technology (for example, OpenAI's GPT-4) to analyze the text data. The input is the text data, and the output is the analysis results divided into the referent and the changes. From the analysis results, the "target" (for example, the room's interior) and the "changes" (for example, Swiss lodge style) are identified.

[0164] Step 4: Capture footage

[0165] The device uses the built-in camera to capture images within the user's field of view. The camera is activated and images are collected in real time. The input is the real-world image captured by the camera, and the output is the image data.

[0166] Step 5: Object Recognition

[0167] The device sends the captured video data to a server, which uses image recognition technology (e.g., YOLO or OpenCV) to identify real-world objects (e.g., furniture or passersby). The input is the video data, and the output is the object's position and orientation. The server recognizes the spatial coordinates and orientation of each object and collects the information.

[0168] Step 6: Select replacement data

[0169] The server selects replacement character and interior data based on the analyzed instructions, and obtains appropriate 3D models and animation data from the data library. The input is the analysis results and object position and orientation information, and the output is the selected 3D model and animation data.

[0170] Step 7: Generate AR data

[0171] The server generates AR data for placing the selected character and interior data in the real world based on the acquired object position and orientation information. This data also includes information necessary for real-time rendering, such as lighting and shadows. The input is the object position and orientation information, as well as the selected 3D model and animation data, and the output is the generated AR data.

[0172] Step 8: Send AR data

[0173] The server sends the generated AR data to the device. The data is compressed in an efficient format and transmitted with low latency. The input is the generated AR data, and the output is the transmitted AR data.

[0174] Step 9: Overlay on camera image

[0175] The device processes the received AR data in real time and overlays it on the camera image using a rendering engine (e.g., Unity or Unreal Engine). Through the AR glasses, the user can see that real-world objects have been replaced with virtual characters and interiors. The input is the transmitted AR data and real-world camera images, and the output is the combined image displayed in the user's field of view.

[0176] (Application example 1)

[0177] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0178] Systems that can quickly and effectively replace real-world interiors and objects with virtual characters and decorations are extremely important in the entertainment and marketing fields. However, existing systems have issues with not being able to change interiors or optimize product displays in real time according to seasons or events. There is also a lack of flexible systems that allow users to easily issue commands and have changes reflected instantly. This makes it difficult to provide an engaging customer experience in physical stores.

[0179] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0180] In this invention, the server includes means for acquiring voice instructions from a user, means for converting the acquired voice instructions into text data and analyzing the instruction content, means for capturing camera images, identifying real-world objects, and acquiring position and orientation information, means for selecting replacement data based on the designated character or interior, means for generating the selected replacement data using the object position and orientation information, means for displaying the generated data by overlaying it on the camera image, and means for generating and changing AR data of the interior and products of a physical store according to the season or an event. This enables physical stores to change the interior and product displays in real time to match the season or an event, which is expected to improve the customer experience.

[0181] "Voice instructions from the user" refers to a form in which the user inputs commands or requests by voice.

[0182] "Text data" is a data format in which voice instructions are converted into character string information.

[0183] The "instruction content" is a specific request or command given by the user through voice instructions.

[0184] "Camera footage" is real-world video data captured by a camera.

[0185] "Real-world objects" are actual objects such as furniture, products, and walls that appear in the camera image.

[0186] "Position and orientation information" is data on where an object is currently located and in what orientation it is facing.

[0187] "Replacement data" is virtual character and interior data that is roughly matched to real-world objects based on user instructions.

[0188] "Generated data" refers to virtual data that is processed and generated to be overlaid on camera images.

[0189] "Overlaying and displaying on camera images" means overlaying and displaying virtual data on images captured by a camera.

[0190] "According to the season or event" means changing the theme or decorations to suit the time of year or a particular event.

[0191] "Brick and mortar store interior" refers to the internal layout and decoration of an actual store.

[0192] "AR data of a product" is virtual data of a product displayed using augmented reality technology.

[0193] "Processing in real time" means processing data in immediate response to user instructions or changes in the environment.

[0194] "Lens technology" is a term that refers to the technology for integrating real-world images with virtual data.

[0195] "Integration" means bringing together different data and information.

[0196] This invention is a system that uses smart glasses or a smartphone worn by the user to display real-world objects replaced with virtual characters and interiors. This system enables real-world stores to change their interiors and product displays in real time according to the season or events, which is expected to improve the customer experience.

[0197] Specifically, it consists of the following elements:

[0198] Server: This receives voice instructions from the user and converts them into text data. A speech recognition API (e.g., Google Speech-to-Text API) is used. The acquired text data is then analyzed using natural language processing (NLP) technology to understand the instructions. An example of a natural language processing API is OpenAI's GPT-4. Based on the analyzed instructions, appropriate replacement data is selected. The server also receives camera images sent from the smart glasses or smartphone, identifies real-world objects using an image recognition API (e.g., Amazon Rekognition), and obtains information about the object's position and orientation. Based on this information, the server generates the selected replacement data and generates the augmented reality (AR) data needed to place the data in the real world at the optimal position and orientation.

[0199] Terminal (smart glasses or smartphone): The device allows the user to give voice commands, captures the voice, and sends it to the server. It also has the function of capturing camera images and sending them to the server. Furthermore, it overlays the AR data received from the server on the camera image in real time. A real-time rendering engine (e.g., Unity) is used for the display.

[0200] Examples:

[0201] Example 1: To change the interior of a store to suit the season, the user (store manager or staff member) issues a voice command to the smart glasses saying, "Decorate the interior of the store to suit the spring season." The smart glasses capture the voice and send it to the server. The server converts the voice into text data, analyzes it, and identifies it as "spring decoration." It receives camera footage, identifies objects in the store (shelves, products, walls, etc.), and acquires their position and orientation information. It selects spring decoration data, processes it in real time, and generates AR data. The smart glasses receive the generated data and display it overlaid on the camera footage from inside the store. In a similar manner, it is possible to change the interior of the store to suit the season or event.

[0202] Example prompt sentence:

[0203] "The interior of the store is decorated in the style of a cherry blossom garden."

[0204] "Change to Christmas decorations for the event."

[0205] This allows virtual changes to store interiors and merchandise to be made in real time, providing customers with an engaging experience.

[0206] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0207] Step 1:

[0208] The user issues a voice command. The user issues a voice command into the smart glasses or smartphone, such as "Decorate the interior of the store to match the spring season." The input is voice, and the output is voice data.

[0209] Step 2:

[0210] The device captures the audio and converts it to text data using a speech recognition API (e.g., Google Speech-to-Text API). The input here is audio data, and the output is text data. During the process, the audio is analyzed and converted into a string.

[0211] Step 3:

[0212] The terminal sends text data to the server. The input is text data, and the output is data sent to the server. The data is sent to the server using the terminal's communication function.

[0213] Step 4:

[0214] The server receives the text data and analyzes it using a natural language processing API (e.g., OpenAI's GPT-4). The input is text data, and the output is instructions as a result of the analysis (e.g., "Decorate the interior of the store in a spring-like style"). Natural language processing technology is used to understand the meaning of the text.

[0215] Step 5:

[0216] The server selects appropriate replacement data (spring decoration data) based on the analysis results. The input is the analysis results, and the output is the selected replacement data. Appropriate 3D models and animation data are selected from the library.

[0217] Step 6:

[0218] The device's camera captures images of the real world and sends the data to the server. The input is the camera image, and the output is the image data sent to the server. Images are captured and sent in real time.

[0219] Step 7:

[0220] The server receives the video data and uses an image recognition API (e.g., Amazon Rekognition) to identify real-world objects. The input is the video data, and the output is information about the identified objects (position and posture information). Objects are recognized through image recognition processing.

[0221] Step 8:

[0222] The server generates AR data based on the object's position and orientation information to apply the selected replacement data to the real world. The input is object information and replacement data, and the output is AR data. Rendering calculations are performed to place the data in the appropriate position.

[0223] Step 9:

[0224] The server sends the generated AR data to the device. The input is the generated AR data, and the output is the data sent to the device. The data is sent using the server's communication function.

[0225] Step 10:

[0226] The AR data received by the device is overlaid on the camera image in real time. The input is the AR data and the camera image, and the output is an augmented reality image that the user can see. The display is done using a real-time rendering engine (e.g., Unity).

[0227] The above steps make it possible to change the interior and product displays of physical stores in real time according to the season or events.

[0228] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0229] This invention is a system that uses AR glasses worn by the user to replace real-world objects with virtual characters and interiors. This invention includes a function that, by combining it with an emotion engine that recognizes the user's emotions, makes more adaptive changes to the display according to the user's emotional state. The program processing is explained below in natural language.

[0230] Program processing explanation

[0231] 1. Recognition of voice commands

[0232] The user issues voice commands to the AR glasses. Possible commands include "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street on my commute into characters." The device captures the voice input and converts it into text data. The text data is then sent to a server for analysis.

[0233] 2. Emotion Recognition by Emotion Engine

[0234] The device's built-in emotion engine recognizes emotions from the user's voice and facial expressions. For example, it analyzes the user's tone of voice and facial muscle movements to identify emotions such as joy, sadness, anger, and surprise.

[0235] 3. Instruction and emotion analysis

[0236] The server analyzes the acquired text data. It uses natural language processing technology to understand the meaning of the voice commands. It also analyzes emotion recognition data from the emotion engine to understand the user's emotional state. Based on the results of this analysis, it selects the most appropriate target (character or interior design).

[0237] 4. Video Capture and Object Recognition

[0238] The device's camera captures video from the user's point of view. The video data is sent in real time to a server. The server uses image recognition technology to identify real-world objects (passersby, vehicles, furniture, etc.). The server then determines the location and orientation of each object and stores them in a database.

[0239] 5. Selection of replacement data

[0240] The server selects the best replacement character and interior design data based on the analyzed instructions and the recognized emotion. For example, if the user is feeling sad, it selects an uplifting character and interior design. This data is retrieved from a library and includes the necessary 3D models and animation data.

[0241] 6. AR Data Generation

[0242] The server generates AR data based on the acquired object position and orientation information to display the selected character and interior in the real world. This data also includes information necessary for real-time rendering, such as lighting and shadows.

[0243] 7. Displaying Data

[0244] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image. From the user's perspective, this makes it appear as if real-world objects have been replaced with the specified characters or interior.

[0245] Specific examples

[0246] Example 1: Change the room interior to Swiss lodge style and adjust it according to the emotion.

[0247] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[0248] 2. The device converts the voice into text data and sends it to the server.

[0249] 3. The device's emotion engine recognizes the user's emotions and determines that they want to relax.

[0250] 4. The server analyzes the data and identifies instructions for interior modifications.

[0251] 5. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[0252] 6. Based on the emotional data, the server selects interior design data that is more relaxing, such as a Swiss lodge.

[0253] 7. The server combines the interior data with the room's location information to generate AR data.

[0254] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. The user feels as if the interior of the room has been transformed into a Swiss lodge, creating an even more relaxing atmosphere.

[0255] Example 2: Transforming pedestrians into characters during your commute and adjusting their emotions

[0256] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[0257] 2. The device converts the voice into text data and sends it to the server.

[0258] 3. The device's emotion engine recognizes the user's emotions and determines that they are "feeling stressed."

[0259] 4. The server analyzes the data and identifies the passerby's replacement instructions.

[0260] 5. The device's camera captures images of the city, and the server recognizes passersby.

[0261] 6. The server selects character data that has the effect of reducing stress based on the emotional data.

[0262] 7. The server combines the character data with the location information of passersby to generate AR data.

[0263] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. Passersby on the way to work are transformed into characters that help reduce stress, making the commute more enjoyable for the user.

[0264] This system will enable users to enjoy a variety of entertainment experiences tailored to their emotions in their daily lives, further improving their quality of life.

[0265] The processing flow will be explained below.

[0266] Step 1:

[0267] The user issues voice commands to the AR glasses, such as "Turn people on the street into characters on my commute."

[0268] Step 2:

[0269] The device captures voice commands and converts the voice data into digital form, which is then analyzed by an internal speech recognition engine and converted into text data.

[0270] Step 3:

[0271] The terminal transmits the converted text data to the server, which includes the content of the voice instructions.

[0272] Step 4:

[0273] The server analyzes the received text data, using natural language processing technology to understand the meaning of the voice commands and extract the "target (people, room interior, etc.)" and "change content (character, Swiss lodge style, etc.)."

[0274] Step 5:

[0275] The device's emotion engine recognizes emotions from the user's voice and facial expressions, analyzing the tone of voice and facial muscle movements to identify emotions such as joy, sadness, anger, and surprise.

[0276] Step 6:

[0277] The device transmits the recognized emotion data to the server.

[0278] Step 7:

[0279] The server analyzes the voice commands and emotional data and then selects the character and interior design that best suits the user's emotions and instructions. For example, if the user is feeling sad, it will select an uplifting character.

[0280] Step 8:

[0281] The device's camera captures images of the real world from the user's point of view, and the image data is sent to a server in real time.

[0282] Step 9:

[0283] The server analyzes the received video data and uses image recognition technology to identify objects in the video (passersby, vehicles, furniture, etc.), and determines the position and posture of each object and stores them in a database.

[0284] Step 10:

[0285] The server uses the object's position and orientation information to generate AR data for displaying the selected character and interior in the real world. This data also includes the necessary rendering information.

[0286] Step 11:

[0287] The server sends the generated AR data to the device.

[0288] Step 12:

[0289] The device processes the AR data received in real time, and selected characters and interior designs are overlaid on the camera image.

[0290] Step 13:

[0291] The app also adjusts based on emotions. For example, if the user is feeling stressed, characters and interior design that will help alleviate stress will be selected and reflected in the display.

[0292] Step 14:

[0293] When the user observes the real world through the AR glasses, characters and interiors that change according to voice commands and emotions are superimposed on the real world, allowing the user to experience passersby on their commute or the interior of their room as if they were replaced with virtual characters and designs.

[0294] Example 2

[0295] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0296] Conventional augmented reality (AR) systems can replace real-world objects with virtual characters or interiors based on user instructions, but they lack the ability to adaptively change the display according to the user's emotional state. This makes it difficult to provide an experience that is in line with the user's emotions. It is also difficult to process in real time and realize a display that includes natural lighting and shadows.

[0297] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0298] In this invention, the server includes means for receiving voice instructions from a user, means for converting the received voice instructions into text data and analyzing the instructions, means for recognizing emotions from the user's voice and facial expressions, means for capturing camera images, identifying real-world objects, and acquiring position and orientation information, means for selecting optimal replacement data based on the voice instructions and the recognized emotions, means for generating the selected replacement data using the object's position and orientation information, and means for displaying the generated data overlaid on the camera image, thereby enabling adaptive AR display according to the user's emotional state in real time.

[0299] A "user" is a person who uses the system to issue voice commands and replace real-world objects with virtual characters and interiors.

[0300] "Voice instructions" are voice requests or commands given by the user to the system.

[0301] "Text data" is character information converted from voice instructions using voice recognition technology.

[0302] "Emotion recognition means" is a technology that analyzes the user's voice and facial expressions to identify the user's emotional state (joy, sadness, anger, etc.).

[0303] "Camera footage" refers to real-world video data captured in real time by a camera mounted on a device.

[0304] "Object" refers to a concrete object that exists in the real world (e.g., a passerby, a vehicle, furniture, etc.).

[0305] "Position and orientation information" refers to data on the physical position and orientation of an object identified in a camera image.

[0306] "Replacement data" is digital data used to replace an object with a virtual character or virtual interior.

[0307] "Rendering information" is data that includes the representation of lighting, shadows, and other elements required for AR display.

[0308] "Real-time processing" refers to the process of analyzing, generating, and displaying data instantly, without delay.

[0309] MODE FOR CARRYING OUT THE INVENTION

[0310] This invention is a system that uses augmented reality (AR) glasses worn by the user to display real-world objects replaced with virtual characters and interiors. This system incorporates an emotion engine that recognizes the user's emotions and includes a function to adaptively change the display according to the user's emotional state.

[0311] System Configuration

[0312] Hardware and software used

[0313] Device: AR glasses

[0314] Camera: A camera for capturing images of the real world.

[0315] Microphone: A microphone for capturing user voice commands.

[0316] Display: A display for overlaying virtual objects onto real-world images.

[0317] Server: Responsible for calculation processing

[0318] Natural language processing technology: For example, Google Cloud Natural Language API

[0319] Speech recognition software: for example, Google Cloud Speech-to-Text

[0320] Emotion recognition software: Examples include IBM Watson® Tone Analyzer and Microsoft® Azure® Face API

[0321] Image recognition technology: For example, Google Cloud Vision API

[0322] AR data generation platform: Examples include Unity and Unreal Engine

[0323] Program processing

[0324] Acquiring and analyzing voice instructions

[0325] The user issues voice commands to the AR glasses, such as "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street to characters on my commute." The device captures these commands and converts them into text using voice recognition software. The converted text is then sent to a server where it is analyzed using natural language processing technology.

[0326] emotion recognition

[0327] The device's built-in emotion engine recognizes emotions such as joy, sadness, anger, and surprise from the user's voice and facial expressions. Speech and facial recognition software analyzes this data to identify the user's emotional state.

[0328] Video Capture and Object Recognition

[0329] The device's camera captures images from the user's point of view in real time and transmits the image data to a server, which uses image recognition technology to identify real-world objects and determine their position and orientation.

[0330] Selecting and generating replacement data

[0331] The server selects the most suitable virtual character and interior design data based on the analyzed voice commands and emotion recognition data. This data is selected from a library stored in advance, and lighting and shadow information required for real-time rendering is also synchronized.

[0332] Viewing Data

[0333] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image. For example, if the user is feeling sad, an uplifting character or interior design will be displayed overlaid on the real world.

[0334] Specific examples

[0335] Example 1: Change the room interior to Swiss lodge style and adjust it according to the emotion.

[0336] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[0337] 2. The device converts the voice into text data and sends it to the server.

[0338] 3. The device's emotion engine recognizes the user's emotion as "I want to relax."

[0339] 4. The server analyzes the data and identifies instructions for interior modifications.

[0340] 5. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[0341] 6. The server selects interior design data that is more relaxing, such as a Swiss lodge.

[0342] 7. The server combines the interior data with the room's location information to generate AR data.

[0343] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. The user feels as if the interior of the room has been transformed into a Swiss lodge, creating an even more relaxing atmosphere.

[0344] Example 2: Transforming pedestrians into characters during your commute and adjusting their emotions

[0345] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[0346] 2. The device converts the voice into text data and sends it to the server.

[0347] 3. The device's emotion engine recognizes the user's emotion as "feeling stressed."

[0348] 4. The server analyzes the data and identifies the passerby's replacement instructions.

[0349] 5. The device's camera captures images of the city, and the server recognizes passersby.

[0350] 6. The server selects character data that has the effect of reducing stress based on the emotional data.

[0351] 7. The server combines the character data with the location information of passersby to generate AR data.

[0352] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. Passersby on the user's commute are transformed into characters that help reduce stress, making the commute more enjoyable.

[0353] In this way, the system of the present invention provides a variety of entertainment experiences that correspond to the user's emotions in their daily lives, contributing to improving the quality of life.

[0354] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0355] Step 1:

[0356] The user issues voice commands to the AR glasses. For example, they might say, "Decorate the room in a Swiss lodge style." The device uses a microphone to capture the user's voice. This voice data becomes the input.

[0357] Step 2:

[0358] The device converts the captured voice data into text data using voice recognition software (e.g., voice recognition API). The input is voice data and the output is text data. The converted text data is sent to the server.

[0359] Step 3:

[0360] The server analyzes the received text data using natural language processing technology (e.g., natural language processing API). The input is text data, and the output is the analysis result. Here, the content of the voice instruction is understood.

[0361] Step 4:

[0362] The device uses an emotion engine to recognize emotions from the user's voice and facial expressions. The input is the user's voice data and camera footage, and the output is emotion data. Emotion recognition software (e.g., emotion analysis API) is used to identify emotions from voice tone and facial muscle movements.

[0363] Step 5:

[0364] The device's camera captures the image from the user's point of view in real time. The input is the real-world camera image, and the output is the captured image data. The image data is sent to the server where it is processed.

[0365] Step 6:

[0366] The server analyzes the captured video data using image recognition technology (e.g., image analysis API). The input is the video data, and the output is the position and posture information of the identified objects. This information is stored in a database.

[0367] Step 7:

[0368] The server selects the optimal replacement character and interior data based on the analyzed voice instructions and emotion recognition data. The input is voice instruction data and emotion data, and the output is the selected replacement data. It also obtains appropriate 3D models and animation data from the library.

[0369] Step 8:

[0370] The server generates AR data based on the object's position and orientation information to display the selected character and interior in real-world footage. The input is the object's position and the selected data, and the output is the generated AR data. Lighting and shadow information is also incorporated.

[0371] Step 9:

[0372] The server sends the generated AR data to the device. The device displays the received AR data overlaid on the camera image. The input is AR data, and the output is an AR display overlaid on the image of the real world. Through this, users can experience virtual characters and interiors as if they were in the real world.

[0373] The system utilizes generative AI models and prompts to provide a rich and adaptive entertainment experience that responds to the user's emotions, while also achieving natural AR display in real time, improving quality of life.

[0374] (Application example 2)

[0375] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0376] In recent years, advances in augmented reality (AR) technology have enabled users to combine the real world with virtual elements for enjoyment. However, current systems simply replace real-world objects with virtual characters or interiors, and are unable to adaptively change the display in response to the user's emotions. Therefore, there is a need for the development of systems that can take user emotions into account and provide greater satisfaction and relaxation.

[0377] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring voice instructions from the user, means for converting the acquired voice instructions into text data and analyzing the instruction content, means including an emotion engine for recognizing the user's emotions, means for capturing camera footage, identifying real-world objects, and acquiring position and orientation information, means for selecting replacement data based on the emotion data and the specified character or interior, means for generating the selected replacement data using the object's position and orientation information, and means for overlaying and displaying the generated data on the camera footage. This enables adaptive virtual replacement of real-world objects based on the user's emotions.

[0378] "Means for acquiring voice instructions from the user" refers to devices or software that detect commands spoken by the user to a wearable device such as AR glasses and incorporate them into the system.

[0379] "Means for converting acquired voice instructions into text data and analyzing the content of the instructions" refers to technology or devices that use voice recognition technology to convert voice data into character data, and then analyze the character data to understand the content of the instructions.

[0380] The "means including an emotion engine for recognizing the user's emotions" refers to a software or hardware system for recognizing the user's current emotional state in real time through analysis of voice and facial expressions.

[0381] "Means for capturing camera images, identifying real-world objects, and acquiring position and orientation information" refers to technology and devices that use a camera to capture real-world images, identify objects in the images using image recognition technology, and determine their position and orientation.

[0382] "Means for selecting replacement data based on emotional data and specified character or interior" refers to an algorithm or device for selecting appropriate character or interior data in accordance with the user's emotional state and voice instructions.

[0383] "Means for generating selected replacement data using object position and orientation information" refers to technology or devices that generate data for appropriately placing selected character and interior data in the real world based on the position and orientation information of identified objects.

[0384] "Means for displaying generated data by overlaying it on camera images" refers to technology or devices for displaying generated virtual content by overlaying it on real-world images captured by a camera.

[0385] System Overview

[0386] This system uses AR glasses worn by the user to replace real-world objects with virtual characters and interiors. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it also includes a function to adaptively change the display according to the user's emotional state.

[0387] Hardware and software used

[0388] Hardware: AR glasses (e.g., Microsoft HoloLens (registered trademark)), smart devices with cameras, servers

[0389] Software: Speech recognition engine (e.g., Google Cloud Speech-to-Text), emotion recognition engine (e.g., Microsoft Azure Emotion API), image recognition technology (e.g., YOLO, TENSORFLOW (registered trademark))

[0390] Program Processing Overview

[0391] Obtaining voice commands from the user

[0392] A microphone installed on the device captures voice commands from the user. For example, if a user says, "Make my kitchen look like a Parisian cafe," that voice input is captured.

[0393] Analysis of voice instructions

[0394] The captured voice data is converted into text data by a voice recognition engine, and the converted text data is sent to a server where the instruction content is analyzed by an analysis engine.

[0395] Emotion recognition

[0396] The device's built-in emotion recognition engine analyzes the user's voice and facial expressions in real time to recognize their emotions. For example, if it recognizes that the user is in a relaxed mood, appropriate data will be selected.

[0397] Camera image capture and object recognition

[0398] The device's camera captures video within the user's field of view, and the video data is sent in real time to a server, which uses image recognition technology to identify real-world objects and determine their position and orientation.

[0399] Selection of replacement data

[0400] The server selects the optimal replacement data (characters and interior design) based on the analyzed voice commands and the recognized emotions. For example, it may select a Parisian cafe-style interior design that has a relaxing effect.

[0401] AR data generation

[0402] The server generates AR data based on the acquired object position and orientation information, including lighting and shadow information, to display the selected character and interior in the real world.

[0403] Viewing Data

[0404] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image, making it appear to the user that real-world objects have been replaced with the specified characters or interior.

[0405] Specific examples

[0406] Example 1: Remodeling your kitchen to look like a Parisian cafe

[0407] 1. The user issues a voice command such as, "Make my kitchen look like a Parisian cafe."

[0408] 2. The device converts the voice into text data and sends it to the server.

[0409] 3. The emotion engine determines that the user wants to relax.

[0410] 4. The server analyzes the data and identifies instructions for interior modifications.

[0411] 5. The device's camera captures images of the kitchen, and the server recognizes the location of the furniture.

[0412] 6. The server selects Parisian cafe-style interior data that has a relaxing effect.

[0413] 7. The server combines the interior data with the kitchen location information to generate AR data.

[0414] 8. The server sends the generated AR data to the device, which then overlays it on the camera image, allowing the user to enjoy a relaxing Parisian cafe-style kitchen.

[0415] Examples of prompt statements

[0416] You are a programmer tasked with designing a system for a food delivery AR glasses application that changes the kitchen interior to a virtual theme (e.g., Parisian cafe) and automatically adjusts to the user's emotions. Write a program procedure that meets the following requirements:

[0417] Requirements:

[0418] 1. The user issues a voice command (e.g., "Make my kitchen look like a Parisian cafe").

[0419] 2. Convert the voice instructions into text data and send it to the server.

[0420] 3. Recognize user emotions and automatically select appropriate themes.

[0421] 4. The AR glasses' camera captures images of the kitchen, and the server recognizes the position of the furniture.

[0422] 5. The server selects appropriate theme data (e.g., interior data of a Parisian cafe).

[0423] 6. The server generates the AR data and sends it to the device.

[0424] 7. AR data is overlaid on the camera image on the device.

[0425] Follow this procedure to design your program.

[0426] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0427] Step 1:

[0428] The user utters the voice command "Make my kitchen look like a Parisian cafe." The input is the user's voice, and the output is the captured voice data. The device's microphone captures this voice and stores it in its internal memory.

[0429] Step 2:

[0430] To convert the voice data into text data, the device uses the Google Cloud Speech-to-Text engine. The input is the captured voice data, and the output is the converted text data, which is sent from the device to the server for the next analysis step.

[0431] Step 3:

[0432] The server parses the received text data. Using a natural language processing engine (e.g., various NLP libraries), the input is the transformed text data and the output is the parsed results, which include instructions for the user to transform their kitchen into a Parisian cafe style.

[0433] Step 4:

[0434] The device's built-in emotion recognition engine recognizes the user's emotions in real time. The input is the user's voice tone and facial expression data, and the output is the user's emotional data. This emotional data is sent to the server and processed together with the analysis results.

[0435] Step 5:

[0436] The device's camera captures images of the kitchen and transmits the image data to the server in real time. The input is the captured camera image, and the output is the transmitted image data.

[0437] Step 6:

[0438] The server analyzes the transmitted video data using image recognition technology (e.g., YOLO, TensorFlow). The input is the camera's video data, and the output is the position and orientation information of objects in the kitchen. This allows the server to identify the positions of furniture and other objects in the kitchen.

[0439] Step 7:

[0440] The server selects the optimal replacement data (e.g., interior design data of a Parisian cafe) based on the analyzed voice instructions and the recognized emotion data. The input is the analysis results of the voice instructions and the emotion data, and the output is the selected replacement data.

[0441] Step 8:

[0442] The server uses the object's position and orientation information to generate AR data for displaying the selected interior data in the real world. The input is the object's position and orientation information and the selected replacement data, and the output is the generated AR data. This data includes, for example, 3D models and lighting information.

[0443] Step 9:

[0444] The generated AR data is sent from the server to the device. The input is the generated AR data, and the output is the data sent to the device. The device processes the received AR data in real time and displays it overlaid on the camera image. This makes the kitchen appear to have been transformed into a Parisian cafe in the user's field of vision.

[0445] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0446] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0447] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0448] [Second embodiment]

[0449] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0450] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0451] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0452] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0453] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0455] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0456] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0457] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0458] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0459] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0460] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0461] This invention is a system that uses AR glasses worn by the user to replace real-world objects with virtual characters and interiors. This system allows users to enjoy their daily commute and their home life even more. The program processing is explained below in natural language.

[0462] Program processing explanation

[0463] 1. Recognition of voice commands

[0464] The user issues voice commands to the AR glasses. Possible commands include "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street on my commute into characters." The device captures the voice input and converts it into text data. The text data is then sent to a server for analysis.

[0465] 2. Parsing the instructions

[0466] The server analyzes the acquired text data. It uses natural language processing technology to understand the meaning of the instructions and identify which objects to replace. This analysis clarifies the "target (people, room interior, etc.)" and "changes (characters, Swiss lodge style, etc.)."

[0467] 3. Video Capture and Object Recognition

[0468] The device uses a camera to capture images within the user's field of view. This video data is sent in real time to a server, which uses image recognition technology to identify real-world objects (such as pedestrians, vehicles, and furniture) and obtain the position and orientation information of each object.

[0469] 4. Selection of replacement data

[0470] Based on the parsed instructions, the server selects replacement character and interior data, which is retrieved from a library containing the necessary 3D models and animation data.

[0471] 5. AR Data Generation

[0472] The server generates AR data based on the acquired object position and orientation information, placing the selected character and interior data in the real world at the appropriate position and orientation. This data also includes information necessary for real-time rendering, such as lighting and shadows.

[0473] 6. Displaying Data

[0474] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image. From the user's perspective, this makes it appear as if real-world objects have been replaced with the specified characters or interior.

[0475] Specific examples

[0476] Example 1: Change the room interior to Swiss lodge style

[0477] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[0478] 2. The device converts the voice into text data and sends it to the server.

[0479] 3. The server analyzes the data and identifies instructions for interior modifications.

[0480] 4. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[0481] 5. The server selects Swiss lodge-style interior data.

[0482] 6. The server combines the interior data with the room's location information to generate AR data.

[0483] 7. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. To the user, the interior of the room appears to have been transformed into a Swiss lodge.

[0484] Example 2: Change passersby in the city into characters during your commute

[0485] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[0486] 2. The device converts the voice into text data and sends it to the server.

[0487] 3. The server analyzes the data and identifies the passerby's replacement instructions.

[0488] 4. The device's camera captures images of the city, and the server recognizes passersby.

[0489] 5. The server selects the character data.

[0490] 6. The server combines the character data with the location information of passersby to generate AR data.

[0491] 7. The server sends the generated AR data to the device, which then displays it over the camera image. To the user, it appears as if passersby on their way to work have been transformed into characters.

[0492] This system allows users to enjoy a variety of entertainment experiences in their daily lives, improving their quality of life.

[0493] The processing flow will be explained below.

[0494] Step 1:

[0495] The user issues voice commands to the AR glasses, such as "Turn people on the street into characters on my commute."

[0496] Step 2:

[0497] The device recognizes voice commands and converts the voice data into digital form. An internal voice recognition engine analyzes this digital voice data and converts it into text data.

[0498] Step 3:

[0499] The terminal transmits the converted text data to the server, which includes the content of the voice instructions.

[0500] Step 4:

[0501] The server analyzes the received text data, using natural language processing technology to understand the meaning of the voice commands and extract the "target (people, room interior, etc.)" and "change content (character, Swiss lodge style, etc.)."

[0502] Step 5:

[0503] The device's camera captures images of the real world from the user's point of view, and the image data is sent to a server in real time.

[0504] Step 6:

[0505] The server analyzes the received video data and uses image recognition technology to identify objects in the video (passersby, vehicles, furniture, etc.), and identifies the position and posture information of each object and stores it in a database.

[0506] Step 7:

[0507] Based on the analysis results, the server selects the requested character and interior data, and retrieves the corresponding 3D model and animation data from the library.

[0508] Step 8:

[0509] The server uses the position and orientation information of the identified objects to generate AR data for displaying the selected characters and interiors in the real world, including the necessary rendering information.

[0510] Step 9:

[0511] The server sends the generated AR data to the device.

[0512] Step 10:

[0513] The device processes the AR data received in real time, and selected characters and interior designs are overlaid on the camera image.

[0514] Step 11:

[0515] When the user observes the real world through the AR glasses, the characters and interiors that have been changed as instructed are superimposed on the real world, allowing the user to experience passersby on their commute or the interior of their room as if they were replaced with virtual characters and designs.

[0516] Example 1

[0517] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0518] Systems that provide users with new entertainment experiences by replacing real-world objects with virtual characters and interiors require technology that can accurately understand user instructions and reflect them in real-time video. Furthermore, high accuracy and speed are required when systems integrate real-world video with virtual data, and technology is needed to provide users with a natural, seamless experience.

[0519] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0520] In this invention, the server includes means for receiving voice instructions from a user, means for converting the received voice instructions into text data and analyzing the instructions, means for capturing camera images, identifying real-world objects and obtaining position and orientation information, means for selecting replacement data based on the specified virtual character or interior, means for generating the selected replacement data using the object's position and orientation information, and means for displaying the generated data superimposed on the camera image. This makes it possible to quickly and accurately analyze the user's voice instructions and display virtual data that is highly accurately integrated with real-world images in real time.

[0521] "User" refers to an individual who uses the system to issue voice commands and enjoy the experience of replacing real-world objects with virtual characters and interiors.

[0522] "Voice instruction" refers to a command made by voice that a user issues to request some action or change to the system.

[0523] "Text data" refers to character string information that is generated by capturing voice instructions and converting them using voice recognition software.

[0524] "Camera footage" refers to real-world video data captured by a camera built into a device.

[0525] "Real-world objects" refer to physical objects that are within the user's field of view (e.g., pedestrians, vehicles, furniture, etc.).

[0526] "Position and orientation information" refers to information about the spatial position of an object in the real world, as well as its direction and orientation.

[0527] "Virtual character" refers to a computer-generated character that appears in place of a real-world object.

[0528] "Interior" refers to virtual data about the decoration and design of the space the user is in.

[0529] "Replacement Data" refers to information, including 3D models and animation data, used to replace real-world objects with virtual characters or interiors.

[0530] "Generated data" refers to AR data that is generated by the server based on the object's position and orientation information, and is used to apply virtual characters and interiors to the real world.

[0531] "Means for overlaying and displaying on camera images" refers to the method or technology by which the device receives the generated data, synthesizes it on camera images from the real world, and displays it to the user.

[0532] "Lens technology" refers to the optical techniques and devices used to integrate real-world images with generated data.

[0533] This invention is a system that uses AR glasses worn by the user to display real-world objects replaced with virtual characters and interiors. This system allows users to enjoy their daily commute and home life even more. Below, we will explain the specific program processing method and the hardware and software used.

[0534] System Overview

[0535] This system is primarily comprised of three components: a server, a device (including AR glasses), and a user. It combines various technologies to realize the process in which the user issues voice commands, real-world images are processed according to those commands, and virtual data is synthesized and displayed.

[0536] Hardware and Software

[0537] Device: AR glasses have built-in high-performance microphones and cameras, as well as a processor and necessary memory to process data in real time.

[0538] Server: A server with high processing power that uses natural language processing technology (e.g., OpenAI's GPT-4) and image recognition technology (e.g., YOLO and OpenCV).

[0539] Speech Recognition Software: Uses speech recognition technology such as Google Cloud Speech-to-Text to convert voice commands into text data.

[0540] Rendering engine: An engine such as Unity or Unreal Engine that integrates real-world images with virtual data for display.

[0541] Specific processing of the program

[0542] Recognizing voice commands

[0543] The user issues voice commands into the AR glasses. For example, they can give specific instructions such as "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street to characters on my commute." The device uses a built-in microphone to capture the voice and obtains the voice data in real time. The obtained voice data is then converted into text data using voice recognition software.

[0544] Analysis of voice instructions

[0545] The device sends the converted text data to a server, which then uses natural language processing technology to analyze the text data. From the analysis results, the server identifies the "subject" (people, room interior, etc.) and the "changes" (characters, Swiss lodge style, etc.).

[0546] Video Capture and Object Recognition

[0547] The device uses a built-in camera to capture images within the user's field of view. This video data is sent in real time to a server, which uses image recognition technology to identify real-world objects (such as pedestrians, vehicles, and furniture) and obtain the position and orientation information of each object.

[0548] Selection of replacement data

[0549] Based on the parsed instructions, the server selects replacement character and interior data, which is retrieved from a library containing the necessary 3D models and animation data.

[0550] AR data generation

[0551] The server generates AR data based on the acquired object position and orientation information, placing the selected character and interior data in the real world at the appropriate position and orientation. This data also includes information necessary for real-time rendering, such as lighting and shadows.

[0552] Viewing Data

[0553] The generated AR data is sent from the server to the device, which processes the received AR data in real time and displays it overlaid on the camera image. Through the AR glasses, the user can see that real-world objects have been replaced with the specified characters or interior.

[0554] Specific examples

[0555] Example 1: Change the room interior to Swiss lodge style

[0556] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[0557] 2. The device converts the voice into text data and sends it to the server.

[0558] 3. The server analyzes the data and identifies instructions for interior modifications.

[0559] 4. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[0560] 5. The server selects Swiss lodge-style interior data.

[0561] 6. The server combines the interior data with the room's location information to generate AR data.

[0562] 7. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. To the user, the interior of the room appears to have been transformed into a Swiss lodge.

[0563] Example 2: Change passersby in the city into characters during your commute

[0564] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[0565] 2. The device converts the voice into text data and sends it to the server.

[0566] 3. The server analyzes the data and identifies the passerby's replacement instructions.

[0567] 4. The device's camera captures images of the city, and the server recognizes passersby.

[0568] 5. The server selects the character data.

[0569] 6. The server combines the character data with the location information of passersby to generate AR data.

[0570] 7. The server sends the generated AR data to the device, which then displays it over the camera image. To the user, it appears as if passersby on their way to work have been transformed into characters.

[0571] Prompt Sentence Examples

[0572] 1. "How can I transform my room into a Swiss lodge?"

[0573] 2. "How do I direct the people on the street to become characters?"

[0574] This system allows users to enjoy a variety of entertainment experiences in their daily lives, improving their quality of life.

[0575] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0576] Step 1: Getting voice instructions

[0577] The user issues voice commands to the AR glasses. For example, they can give specific instructions such as "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street to characters on my commute." The device captures the voice using a built-in microphone. The input is the user's voice commands, and the output is voice data.

[0578] Step 2: Convert the audio data

[0579] The device converts the captured voice data into text data using voice recognition software (e.g., Google Cloud Speech-to-Text). It uses voice data as input and natural language processing technology to output text data. This conversion process is performed in real time.

[0580] Step 3: Analyzing voice commands

[0581] The device sends the converted text data to the server. The server uses natural language processing technology (for example, OpenAI's GPT-4) to analyze the text data. The input is the text data, and the output is the analysis results divided into the referent and the changes. From the analysis results, the "target" (for example, the room's interior) and the "changes" (for example, Swiss lodge style) are identified.

[0582] Step 4: Capture footage

[0583] The device uses the built-in camera to capture images within the user's field of view. The camera is activated and images are collected in real time. The input is the real-world image captured by the camera, and the output is the image data.

[0584] Step 5: Object Recognition

[0585] The device sends the captured video data to a server, which uses image recognition technology (e.g., YOLO or OpenCV) to identify real-world objects (e.g., furniture or passersby). The input is the video data, and the output is the object's position and orientation. The server recognizes the spatial coordinates and orientation of each object and collects the information.

[0586] Step 6: Select replacement data

[0587] The server selects replacement character and interior data based on the analyzed instructions, and obtains appropriate 3D models and animation data from the data library. The input is the analysis results and object position and orientation information, and the output is the selected 3D model and animation data.

[0588] Step 7: Generate AR data

[0589] The server generates AR data for placing the selected character and interior data in the real world based on the acquired object position and orientation information. This data also includes information necessary for real-time rendering, such as lighting and shadows. The input is the object position and orientation information, as well as the selected 3D model and animation data, and the output is the generated AR data.

[0590] Step 8: Send AR data

[0591] The server sends the generated AR data to the device. The data is compressed in an efficient format and transmitted with low latency. The input is the generated AR data, and the output is the transmitted AR data.

[0592] Step 9: Overlay on camera image

[0593] The device processes the received AR data in real time and overlays it on the camera image using a rendering engine (e.g., Unity or Unreal Engine). Through the AR glasses, the user can see that real-world objects have been replaced with virtual characters and interiors. The input is the transmitted AR data and real-world camera images, and the output is the combined image displayed in the user's field of view.

[0594] (Application example 1)

[0595] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0596] Systems that can quickly and effectively replace real-world interiors and objects with virtual characters and decorations are extremely important in the entertainment and marketing fields. However, existing systems have issues with not being able to change interiors or optimize product displays in real time according to seasons or events. There is also a lack of flexible systems that allow users to easily issue commands and have changes reflected instantly. This makes it difficult to provide an engaging customer experience in physical stores.

[0597] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0598] In this invention, the server includes means for acquiring voice instructions from a user, means for converting the acquired voice instructions into text data and analyzing the instruction content, means for capturing camera images, identifying real-world objects, and acquiring position and orientation information, means for selecting replacement data based on the designated character or interior, means for generating the selected replacement data using the object position and orientation information, means for displaying the generated data by overlaying it on the camera image, and means for generating and changing AR data of the interior and products of a physical store according to the season or an event. This enables physical stores to change the interior and product displays in real time to match the season or an event, which is expected to improve the customer experience.

[0599] "Voice instructions from the user" refers to a form in which the user inputs commands or requests by voice.

[0600] "Text data" is a data format in which voice instructions are converted into character string information.

[0601] The "instruction content" is a specific request or command given by the user through voice instructions.

[0602] "Camera footage" is real-world video data captured by a camera.

[0603] "Real-world objects" are actual objects such as furniture, products, and walls that appear in the camera image.

[0604] "Position and orientation information" is data on where an object is currently located and in what orientation it is facing.

[0605] "Replacement data" is virtual character and interior data that is roughly matched to real-world objects based on user instructions.

[0606] "Generated data" refers to virtual data that is processed and generated to be overlaid on camera images.

[0607] "Overlaying and displaying on camera images" means overlaying and displaying virtual data on images captured by a camera.

[0608] "According to the season or event" means changing the theme or decorations to suit the time of year or a particular event.

[0609] "Brick and mortar store interior" refers to the internal layout and decoration of an actual store.

[0610] "AR data of a product" is virtual data of a product displayed using augmented reality technology.

[0611] "Processing in real time" means processing data in immediate response to user instructions or changes in the environment.

[0612] "Lens technology" is a term that refers to the technology for integrating real-world images with virtual data.

[0613] "Integration" means bringing together different data and information.

[0614] This invention is a system that uses smart glasses or a smartphone worn by the user to display real-world objects replaced with virtual characters and interiors. This system enables real-world stores to change their interiors and product displays in real time according to the season or events, which is expected to improve the customer experience.

[0615] Specifically, it consists of the following elements:

[0616] Server: This receives voice instructions from the user and converts them into text data. A speech recognition API (e.g., Google Speech-to-Text API) is used. The acquired text data is then analyzed using natural language processing (NLP) technology to understand the instructions. An example of a natural language processing API is OpenAI's GPT-4. Based on the analyzed instructions, appropriate replacement data is selected. The server also receives camera images sent from the smart glasses or smartphone, identifies real-world objects using an image recognition API (e.g., Amazon Rekognition), and obtains information about the object's position and orientation. Based on this information, the server generates the selected replacement data and generates the augmented reality (AR) data needed to place the data in the real world at the optimal position and orientation.

[0617] Terminal (smart glasses or smartphone): The device allows the user to give voice commands, captures the voice, and sends it to the server. It also has the function of capturing camera images and sending them to the server. Furthermore, it overlays the AR data received from the server on the camera image in real time. A real-time rendering engine (e.g., Unity) is used for the display.

[0618] Examples:

[0619] Example 1: To change the interior of a store to suit the season, the user (store manager or staff member) issues a voice command to the smart glasses saying, "Decorate the interior of the store to suit the spring season." The smart glasses capture the voice and send it to the server. The server converts the voice into text data, analyzes it, and identifies it as "spring decoration." It receives camera footage, identifies objects in the store (shelves, products, walls, etc.), and acquires their position and orientation information. It selects spring decoration data, processes it in real time, and generates AR data. The smart glasses receive the generated data and display it overlaid on the camera footage from inside the store. In a similar manner, it is possible to change the interior of the store to suit the season or event.

[0620] Example prompt sentence:

[0621] "The interior of the store is decorated in the style of a cherry blossom garden."

[0622] "Change to Christmas decorations for the event."

[0623] This allows virtual changes to store interiors and merchandise to be made in real time, providing customers with an engaging experience.

[0624] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0625] Step 1:

[0626] The user issues a voice command. The user issues a voice command into the smart glasses or smartphone, such as "Decorate the interior of the store to match the spring season." The input is voice, and the output is voice data.

[0627] Step 2:

[0628] The device captures the audio and converts it to text data using a speech recognition API (e.g., Google Speech-to-Text API). The input here is audio data, and the output is text data. During the process, the audio is analyzed and converted into a string.

[0629] Step 3:

[0630] The terminal sends text data to the server. The input is text data, and the output is data sent to the server. The data is sent to the server using the terminal's communication function.

[0631] Step 4:

[0632] The server receives the text data and analyzes it using a natural language processing API (e.g., OpenAI's GPT-4). The input is text data, and the output is instructions as a result of the analysis (e.g., "Decorate the interior of the store in a spring-like style"). Natural language processing technology is used to understand the meaning of the text.

[0633] Step 5:

[0634] The server selects appropriate replacement data (spring decoration data) based on the analysis results. The input is the analysis results, and the output is the selected replacement data. Appropriate 3D models and animation data are selected from the library.

[0635] Step 6:

[0636] The device's camera captures images of the real world and sends the data to the server. The input is the camera image, and the output is the image data sent to the server. Images are captured and sent in real time.

[0637] Step 7:

[0638] The server receives the video data and uses an image recognition API (e.g., Amazon Rekognition) to identify real-world objects. The input is the video data, and the output is information about the identified objects (position and posture information). Objects are recognized through image recognition processing.

[0639] Step 8:

[0640] The server generates AR data based on the object's position and orientation information to apply the selected replacement data to the real world. The input is object information and replacement data, and the output is AR data. Rendering calculations are performed to place the data in the appropriate position.

[0641] Step 9:

[0642] The server sends the generated AR data to the device. The input is the generated AR data, and the output is the data sent to the device. The data is sent using the server's communication function.

[0643] Step 10:

[0644] The AR data received by the device is overlaid on the camera image in real time. The input is the AR data and the camera image, and the output is an augmented reality image that the user can see. The display is done using a real-time rendering engine (e.g., Unity).

[0645] The above steps make it possible to change the interior and product displays of physical stores in real time according to the season or events.

[0646] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0647] This invention is a system that uses AR glasses worn by the user to replace real-world objects with virtual characters and interiors. This invention includes a function that, by combining it with an emotion engine that recognizes the user's emotions, makes more adaptive changes to the display according to the user's emotional state. The program processing is explained below in natural language.

[0648] Program processing explanation

[0649] 1. Recognition of voice commands

[0650] The user issues voice commands to the AR glasses. Possible commands include "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street on my commute into characters." The device captures the voice input and converts it into text data. The text data is then sent to a server for analysis.

[0651] 2. Emotion Recognition by Emotion Engine

[0652] The device's built-in emotion engine recognizes emotions from the user's voice and facial expressions. For example, it analyzes the user's tone of voice and facial muscle movements to identify emotions such as joy, sadness, anger, and surprise.

[0653] 3. Instruction and emotion analysis

[0654] The server analyzes the acquired text data. It uses natural language processing technology to understand the meaning of the voice commands. It also analyzes emotion recognition data from the emotion engine to understand the user's emotional state. Based on the results of this analysis, it selects the most appropriate target (character or interior design).

[0655] 4. Video Capture and Object Recognition

[0656] The device's camera captures video from the user's point of view. The video data is sent in real time to a server. The server uses image recognition technology to identify real-world objects (passersby, vehicles, furniture, etc.). The server then determines the location and orientation of each object and stores them in a database.

[0657] 5. Selection of replacement data

[0658] The server selects the best replacement character and interior design data based on the analyzed instructions and the recognized emotion. For example, if the user is feeling sad, it selects an uplifting character and interior design. This data is retrieved from a library and includes the necessary 3D models and animation data.

[0659] 6. AR Data Generation

[0660] The server generates AR data based on the acquired object position and orientation information to display the selected character and interior in the real world. This data also includes information necessary for real-time rendering, such as lighting and shadows.

[0661] 7. Displaying Data

[0662] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image. From the user's perspective, this makes it appear as if real-world objects have been replaced with the specified characters or interior.

[0663] Specific examples

[0664] Example 1: Change the room interior to Swiss lodge style and adjust it according to the emotion.

[0665] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[0666] 2. The device converts the voice into text data and sends it to the server.

[0667] 3. The device's emotion engine recognizes the user's emotions and determines that they want to relax.

[0668] 4. The server analyzes the data and identifies instructions for interior modifications.

[0669] 5. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[0670] 6. Based on the emotional data, the server selects interior design data that is more relaxing, such as a Swiss lodge.

[0671] 7. The server combines the interior data with the room's location information to generate AR data.

[0672] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. The user feels as if the interior of the room has been transformed into a Swiss lodge, creating an even more relaxing atmosphere.

[0673] Example 2: Transforming pedestrians into characters during your commute and adjusting their emotions

[0674] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[0675] 2. The device converts the voice into text data and sends it to the server.

[0676] 3. The device's emotion engine recognizes the user's emotions and determines that they are "feeling stressed."

[0677] 4. The server analyzes the data and identifies the passerby's replacement instructions.

[0678] 5. The device's camera captures images of the city, and the server recognizes passersby.

[0679] 6. The server selects character data that has the effect of reducing stress based on the emotional data.

[0680] 7. The server combines the character data with the location information of passersby to generate AR data.

[0681] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. Passersby on the way to work are transformed into characters that help reduce stress, making the commute more enjoyable for the user.

[0682] This system will enable users to enjoy a variety of entertainment experiences tailored to their emotions in their daily lives, further improving their quality of life.

[0683] The processing flow will be explained below.

[0684] Step 1:

[0685] The user issues voice commands to the AR glasses, such as "Turn people on the street into characters on my commute."

[0686] Step 2:

[0687] The device captures voice commands and converts the voice data into digital form, which is then analyzed by an internal speech recognition engine and converted into text data.

[0688] Step 3:

[0689] The terminal transmits the converted text data to the server, which includes the content of the voice instructions.

[0690] Step 4:

[0691] The server analyzes the received text data, using natural language processing technology to understand the meaning of the voice commands and extract the "target (people, room interior, etc.)" and "change content (character, Swiss lodge style, etc.)."

[0692] Step 5:

[0693] The device's emotion engine recognizes emotions from the user's voice and facial expressions, analyzing the tone of voice and facial muscle movements to identify emotions such as joy, sadness, anger, and surprise.

[0694] Step 6:

[0695] The device transmits the recognized emotion data to the server.

[0696] Step 7:

[0697] The server analyzes the voice commands and emotional data and then selects the character and interior design that best suits the user's emotions and instructions. For example, if the user is feeling sad, it will select an uplifting character.

[0698] Step 8:

[0699] The device's camera captures images of the real world from the user's point of view, and the image data is sent to a server in real time.

[0700] Step 9:

[0701] The server analyzes the received video data and uses image recognition technology to identify objects in the video (passersby, vehicles, furniture, etc.), and determines the position and posture of each object and stores them in a database.

[0702] Step 10:

[0703] The server uses the object's position and orientation information to generate AR data for displaying the selected character and interior in the real world. This data also includes the necessary rendering information.

[0704] Step 11:

[0705] The server sends the generated AR data to the device.

[0706] Step 12:

[0707] The device processes the AR data received in real time, and selected characters and interior designs are overlaid on the camera image.

[0708] Step 13:

[0709] The app also adjusts based on emotions. For example, if the user is feeling stressed, characters and interior design that will help alleviate stress will be selected and reflected in the display.

[0710] Step 14:

[0711] When the user observes the real world through the AR glasses, characters and interiors that change according to voice commands and emotions are superimposed on the real world, allowing the user to experience passersby on their commute or the interior of their room as if they were replaced with virtual characters and designs.

[0712] Example 2

[0713] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0714] Conventional augmented reality (AR) systems can replace real-world objects with virtual characters or interiors based on user instructions, but they lack the ability to adaptively change the display according to the user's emotional state. This makes it difficult to provide an experience that is in line with the user's emotions. It is also difficult to process in real time and realize a display that includes natural lighting and shadows.

[0715] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0716] In this invention, the server includes means for receiving voice instructions from a user, means for converting the received voice instructions into text data and analyzing the instructions, means for recognizing emotions from the user's voice and facial expressions, means for capturing camera images, identifying real-world objects, and acquiring position and orientation information, means for selecting optimal replacement data based on the voice instructions and the recognized emotions, means for generating the selected replacement data using the object's position and orientation information, and means for displaying the generated data overlaid on the camera image, thereby enabling adaptive AR display according to the user's emotional state in real time.

[0717] A "user" is a person who uses the system to issue voice commands and replace real-world objects with virtual characters and interiors.

[0718] "Voice instructions" are voice requests or commands given by the user to the system.

[0719] "Text data" is character information converted from voice instructions using voice recognition technology.

[0720] "Emotion recognition means" is a technology that analyzes the user's voice and facial expressions to identify the user's emotional state (joy, sadness, anger, etc.).

[0721] "Camera footage" refers to real-world video data captured in real time by a camera mounted on a device.

[0722] "Object" refers to a concrete object that exists in the real world (e.g., a passerby, a vehicle, furniture, etc.).

[0723] "Position and orientation information" refers to data on the physical position and orientation of an object identified in a camera image.

[0724] "Replacement data" is digital data used to replace an object with a virtual character or virtual interior.

[0725] "Rendering information" is data that includes the representation of lighting, shadows, and other elements required for AR display.

[0726] "Real-time processing" refers to the process of analyzing, generating, and displaying data instantly, without delay.

[0727] MODE FOR CARRYING OUT THE INVENTION

[0728] This invention is a system that uses augmented reality (AR) glasses worn by the user to display real-world objects replaced with virtual characters and interiors. This system incorporates an emotion engine that recognizes the user's emotions and includes a function to adaptively change the display according to the user's emotional state.

[0729] System Configuration

[0730] Hardware and software used

[0731] Device: AR glasses

[0732] Camera: A camera for capturing images of the real world.

[0733] Microphone: A microphone for capturing user voice commands.

[0734] Display: A display for overlaying virtual objects onto real-world images.

[0735] Server: Responsible for calculation processing

[0736] Natural language processing technology: For example, Google Cloud Natural Language API

[0737] Speech recognition software: for example, Google Cloud Speech-to-Text

[0738] Emotion recognition software: Examples include IBM Watson Tone Analyzer and Microsoft Azure Face API

[0739] Image recognition technology: For example, Google Cloud Vision API

[0740] AR data generation platform: Examples include Unity and Unreal Engine

[0741] Program processing

[0742] Acquiring and analyzing voice instructions

[0743] The user issues voice commands to the AR glasses, such as "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street to characters on my commute." The device captures these commands and converts them into text using voice recognition software. The converted text is then sent to a server where it is analyzed using natural language processing technology.

[0744] emotion recognition

[0745] The device's built-in emotion engine recognizes emotions such as joy, sadness, anger, and surprise from the user's voice and facial expressions. Speech and facial recognition software analyzes this data to identify the user's emotional state.

[0746] Video Capture and Object Recognition

[0747] The device's camera captures images from the user's point of view in real time and transmits the image data to a server, which uses image recognition technology to identify real-world objects and determine their position and orientation.

[0748] Selecting and generating replacement data

[0749] The server selects the most suitable virtual character and interior design data based on the analyzed voice commands and emotion recognition data. This data is selected from a library stored in advance, and lighting and shadow information required for real-time rendering is also synchronized.

[0750] Viewing Data

[0751] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image. For example, if the user is feeling sad, an uplifting character or interior design will be displayed overlaid on the real world.

[0752] Specific examples

[0753] Example 1: Change the room interior to Swiss lodge style and adjust it according to the emotion.

[0754] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[0755] 2. The device converts the voice into text data and sends it to the server.

[0756] 3. The device's emotion engine recognizes the user's emotion as "I want to relax."

[0757] 4. The server analyzes the data and identifies instructions for interior modifications.

[0758] 5. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[0759] 6. The server selects interior design data that is more relaxing, such as a Swiss lodge.

[0760] 7. The server combines the interior data with the room's location information to generate AR data.

[0761] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. The user feels as if the interior of the room has been transformed into a Swiss lodge, creating an even more relaxing atmosphere.

[0762] Example 2: Transforming pedestrians into characters during your commute and adjusting their emotions

[0763] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[0764] 2. The device converts the voice into text data and sends it to the server.

[0765] 3. The device's emotion engine recognizes the user's emotion as "feeling stressed."

[0766] 4. The server analyzes the data and identifies the passerby's replacement instructions.

[0767] 5. The device's camera captures images of the city, and the server recognizes passersby.

[0768] 6. The server selects character data that has the effect of reducing stress based on the emotional data.

[0769] 7. The server combines the character data with the location information of passersby to generate AR data.

[0770] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. Passersby on the user's commute are transformed into characters that help reduce stress, making the commute more enjoyable.

[0771] In this way, the system of the present invention provides a variety of entertainment experiences that correspond to the user's emotions in their daily lives, contributing to improving the quality of life.

[0772] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0773] Step 1:

[0774] The user issues voice commands to the AR glasses. For example, they might say, "Decorate the room in a Swiss lodge style." The device uses a microphone to capture the user's voice. This voice data becomes the input.

[0775] Step 2:

[0776] The device converts the captured voice data into text data using voice recognition software (e.g., voice recognition API). The input is voice data and the output is text data. The converted text data is sent to the server.

[0777] Step 3:

[0778] The server analyzes the received text data using natural language processing technology (e.g., natural language processing API). The input is text data, and the output is the analysis result. Here, the content of the voice instruction is understood.

[0779] Step 4:

[0780] The device uses an emotion engine to recognize emotions from the user's voice and facial expressions. The input is the user's voice data and camera footage, and the output is emotion data. Emotion recognition software (e.g., emotion analysis API) is used to identify emotions from voice tone and facial muscle movements.

[0781] Step 5:

[0782] The device's camera captures the image from the user's point of view in real time. The input is the real-world camera image, and the output is the captured image data. The image data is sent to the server where it is processed.

[0783] Step 6:

[0784] The server analyzes the captured video data using image recognition technology (e.g., image analysis API). The input is the video data, and the output is the position and posture information of the identified objects. This information is stored in a database.

[0785] Step 7:

[0786] The server selects the optimal replacement character and interior data based on the analyzed voice instructions and emotion recognition data. The input is voice instruction data and emotion data, and the output is the selected replacement data. It also obtains appropriate 3D models and animation data from the library.

[0787] Step 8:

[0788] The server generates AR data based on the object's position and orientation information to display the selected character and interior in real-world footage. The input is the object's position and the selected data, and the output is the generated AR data. Lighting and shadow information is also incorporated.

[0789] Step 9:

[0790] The server sends the generated AR data to the device. The device displays the received AR data overlaid on the camera image. The input is AR data, and the output is an AR display overlaid on the image of the real world. Through this, users can experience virtual characters and interiors as if they were in the real world.

[0791] The system utilizes generative AI models and prompts to provide a rich and adaptive entertainment experience that responds to the user's emotions, while also achieving natural AR display in real time, improving quality of life.

[0792] (Application example 2)

[0793] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0794] In recent years, advances in augmented reality (AR) technology have enabled users to combine the real world with virtual elements for enjoyment. However, current systems simply replace real-world objects with virtual characters or interiors, and are unable to adaptively change the display in response to the user's emotions. Therefore, there is a need for the development of systems that can take user emotions into account and provide greater satisfaction and relaxation.

[0795] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring voice instructions from the user, means for converting the acquired voice instructions into text data and analyzing the instruction content, means including an emotion engine for recognizing the user's emotions, means for capturing camera footage, identifying real-world objects, and acquiring position and orientation information, means for selecting replacement data based on the emotion data and the specified character or interior, means for generating the selected replacement data using the object's position and orientation information, and means for overlaying and displaying the generated data on the camera footage. This enables adaptive virtual replacement of real-world objects based on the user's emotions.

[0796] "Means for acquiring voice instructions from the user" refers to devices or software that detect commands spoken by the user to a wearable device such as AR glasses and incorporate them into the system.

[0797] "Means for converting acquired voice instructions into text data and analyzing the content of the instructions" refers to technology or devices that use voice recognition technology to convert voice data into character data, and then analyze the character data to understand the content of the instructions.

[0798] The "means including an emotion engine for recognizing the user's emotions" refers to a software or hardware system for recognizing the user's current emotional state in real time through analysis of voice and facial expressions.

[0799] "Means for capturing camera images, identifying real-world objects, and acquiring position and orientation information" refers to technology and devices that use a camera to capture real-world images, identify objects in the images using image recognition technology, and determine their position and orientation.

[0800] "Means for selecting replacement data based on emotional data and specified character or interior" refers to an algorithm or device for selecting appropriate character or interior data in accordance with the user's emotional state and voice instructions.

[0801] "Means for generating selected replacement data using object position and orientation information" refers to technology or devices that generate data for appropriately placing selected character and interior data in the real world based on the position and orientation information of identified objects.

[0802] "Means for displaying generated data by overlaying it on camera images" refers to technology or devices for displaying generated virtual content by overlaying it on real-world images captured by a camera.

[0803] System Overview

[0804] This system uses AR glasses worn by the user to replace real-world objects with virtual characters and interiors. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it also includes a function to adaptively change the display according to the user's emotional state.

[0805] Hardware and software used

[0806] Hardware: AR glasses (e.g. Microsoft HoloLens), smart devices with cameras, servers

[0807] Software: Speech recognition engine (e.g., Google Cloud Speech-to-Text), emotion recognition engine (e.g., Microsoft Azure Emotion API), image recognition technology (e.g., YOLO, TensorFlow)

[0808] Program Processing Overview

[0809] Obtaining voice commands from the user

[0810] A microphone installed on the device captures voice commands from the user. For example, if a user says, "Make my kitchen look like a Parisian cafe," that voice input is captured.

[0811] Analysis of voice instructions

[0812] The captured voice data is converted into text data by a voice recognition engine, and the converted text data is sent to a server where the instruction content is analyzed by an analysis engine.

[0813] Emotion recognition

[0814] The device's built-in emotion recognition engine analyzes the user's voice and facial expressions in real time to recognize their emotions. For example, if it recognizes that the user is in a relaxed mood, appropriate data will be selected.

[0815] Camera image capture and object recognition

[0816] The device's camera captures video within the user's field of view, and the video data is sent in real time to a server, which uses image recognition technology to identify real-world objects and determine their position and orientation.

[0817] Selection of replacement data

[0818] The server selects the optimal replacement data (characters and interior design) based on the analyzed voice commands and the recognized emotions. For example, it may select a Parisian cafe-style interior design that has a relaxing effect.

[0819] AR data generation

[0820] The server generates AR data based on the acquired object position and orientation information, including lighting and shadow information, to display the selected character and interior in the real world.

[0821] Viewing Data

[0822] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image, making it appear to the user that real-world objects have been replaced with the specified characters or interior.

[0823] Specific examples

[0824] Example 1: Remodeling your kitchen to look like a Parisian cafe

[0825] 1. The user issues a voice command such as, "Make my kitchen look like a Parisian cafe."

[0826] 2. The device converts the voice into text data and sends it to the server.

[0827] 3. The emotion engine determines that the user wants to relax.

[0828] 4. The server analyzes the data and identifies instructions for interior modifications.

[0829] 5. The device's camera captures images of the kitchen, and the server recognizes the location of the furniture.

[0830] 6. The server selects Parisian cafe-style interior data that has a relaxing effect.

[0831] 7. The server combines the interior data with the kitchen location information to generate AR data.

[0832] 8. The server sends the generated AR data to the device, which then overlays it on the camera image, allowing the user to enjoy a relaxing Parisian cafe-style kitchen.

[0833] Examples of prompt statements

[0834] You are a programmer tasked with designing a system for a food delivery AR glasses application that changes the kitchen interior to a virtual theme (e.g., Parisian cafe) and automatically adjusts to the user's emotions. Write a program procedure that meets the following requirements:

[0835] Requirements:

[0836] 1. The user issues a voice command (e.g., "Make my kitchen look like a Parisian cafe").

[0837] 2. Convert the voice instructions into text data and send it to the server.

[0838] 3. Recognize user emotions and automatically select appropriate themes.

[0839] 4. The AR glasses' camera captures images of the kitchen, and the server recognizes the position of the furniture.

[0840] 5. The server selects appropriate theme data (e.g., interior data of a Parisian cafe).

[0841] 6. The server generates the AR data and sends it to the device.

[0842] 7. AR data is overlaid on the camera image on the device.

[0843] Follow this procedure to design your program.

[0844] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0845] Step 1:

[0846] The user utters the voice command "Make my kitchen look like a Parisian cafe." The input is the user's voice, and the output is the captured voice data. The device's microphone captures this voice and stores it in its internal memory.

[0847] Step 2:

[0848] To convert the voice data into text data, the device uses the Google Cloud Speech-to-Text engine. The input is the captured voice data, and the output is the converted text data, which is sent from the device to the server for the next analysis step.

[0849] Step 3:

[0850] The server parses the received text data. Using a natural language processing engine (e.g., various NLP libraries), the input is the transformed text data and the output is the parsed results, which include instructions for the user to transform their kitchen into a Parisian cafe style.

[0851] Step 4:

[0852] The device's built-in emotion recognition engine recognizes the user's emotions in real time. The input is the user's voice tone and facial expression data, and the output is the user's emotional data. This emotional data is sent to the server and processed together with the analysis results.

[0853] Step 5:

[0854] The device's camera captures images of the kitchen and transmits the image data to the server in real time. The input is the captured camera image, and the output is the transmitted image data.

[0855] Step 6:

[0856] The server analyzes the transmitted video data using image recognition technology (e.g., YOLO, TensorFlow). The input is the camera's video data, and the output is the position and orientation information of objects in the kitchen. This allows the server to identify the positions of furniture and other objects in the kitchen.

[0857] Step 7:

[0858] The server selects the optimal replacement data (e.g., interior design data of a Parisian cafe) based on the analyzed voice instructions and the recognized emotion data. The input is the analysis results of the voice instructions and the emotion data, and the output is the selected replacement data.

[0859] Step 8:

[0860] The server uses the object's position and orientation information to generate AR data for displaying the selected interior data in the real world. The input is the object's position and orientation information and the selected replacement data, and the output is the generated AR data. This data includes, for example, 3D models and lighting information.

[0861] Step 9:

[0862] The generated AR data is sent from the server to the device. The input is the generated AR data, and the output is the data sent to the device. The device processes the received AR data in real time and displays it overlaid on the camera image. This makes the kitchen appear to have been transformed into a Parisian cafe in the user's field of vision.

[0863] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0864] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0865] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0866] [Third embodiment]

[0867] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0868] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0869] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0870] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0871] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0872] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0873] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0874] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0875] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0876] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0877] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0878] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0879] This invention is a system that uses AR glasses worn by the user to replace real-world objects with virtual characters and interiors. This system allows users to enjoy their daily commute and their home life even more. The program processing is explained below in natural language.

[0880] Program processing explanation

[0881] 1. Recognition of voice commands

[0882] The user issues voice commands to the AR glasses. Possible commands include "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street on my commute into characters." The device captures the voice input and converts it into text data. The text data is then sent to a server for analysis.

[0883] 2. Parsing the instructions

[0884] The server analyzes the acquired text data. It uses natural language processing technology to understand the meaning of the instructions and identify which objects to replace. This analysis clarifies the "target (people, room interior, etc.)" and "changes (characters, Swiss lodge style, etc.)."

[0885] 3. Video Capture and Object Recognition

[0886] The device uses a camera to capture images within the user's field of view. This video data is sent in real time to a server, which uses image recognition technology to identify real-world objects (such as pedestrians, vehicles, and furniture) and obtain the position and orientation information of each object.

[0887] 4. Selection of replacement data

[0888] Based on the parsed instructions, the server selects replacement character and interior data, which is retrieved from a library containing the necessary 3D models and animation data.

[0889] 5. AR Data Generation

[0890] The server generates AR data based on the acquired object position and orientation information, placing the selected character and interior data in the real world at the appropriate position and orientation. This data also includes information necessary for real-time rendering, such as lighting and shadows.

[0891] 6. Displaying Data

[0892] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image. From the user's perspective, this makes it appear as if real-world objects have been replaced with the specified characters or interior.

[0893] Specific examples

[0894] Example 1: Change the room interior to Swiss lodge style

[0895] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[0896] 2. The device converts the voice into text data and sends it to the server.

[0897] 3. The server analyzes the data and identifies instructions for interior modifications.

[0898] 4. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[0899] 5. The server selects Swiss lodge-style interior data.

[0900] 6. The server combines the interior data with the room's location information to generate AR data.

[0901] 7. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. To the user, the interior of the room appears to have been transformed into a Swiss lodge.

[0902] Example 2: Change passersby in the city into characters during your commute

[0903] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[0904] 2. The device converts the voice into text data and sends it to the server.

[0905] 3. The server analyzes the data and identifies the passerby's replacement instructions.

[0906] 4. The device's camera captures images of the city, and the server recognizes passersby.

[0907] 5. The server selects the character data.

[0908] 6. The server combines the character data with the location information of passersby to generate AR data.

[0909] 7. The server sends the generated AR data to the device, which then displays it over the camera image. To the user, it appears as if passersby on their way to work have been transformed into characters.

[0910] This system allows users to enjoy a variety of entertainment experiences in their daily lives, improving their quality of life.

[0911] The processing flow will be explained below.

[0912] Step 1:

[0913] The user issues voice commands to the AR glasses, such as "Turn people on the street into characters on my commute."

[0914] Step 2:

[0915] The device recognizes voice commands and converts the voice data into digital form. An internal voice recognition engine analyzes this digital voice data and converts it into text data.

[0916] Step 3:

[0917] The terminal transmits the converted text data to the server, which includes the content of the voice instructions.

[0918] Step 4:

[0919] The server analyzes the received text data, using natural language processing technology to understand the meaning of the voice commands and extract the "target (people, room interior, etc.)" and "change content (character, Swiss lodge style, etc.)."

[0920] Step 5:

[0921] The device's camera captures images of the real world from the user's point of view, and the image data is sent to a server in real time.

[0922] Step 6:

[0923] The server analyzes the received video data and uses image recognition technology to identify objects in the video (passersby, vehicles, furniture, etc.), and identifies the position and posture information of each object and stores it in a database.

[0924] Step 7:

[0925] Based on the analysis results, the server selects the requested character and interior data, and retrieves the corresponding 3D model and animation data from the library.

[0926] Step 8:

[0927] The server uses the position and orientation information of the identified objects to generate AR data for displaying the selected characters and interiors in the real world, including the necessary rendering information.

[0928] Step 9:

[0929] The server sends the generated AR data to the device.

[0930] Step 10:

[0931] The device processes the AR data received in real time, and selected characters and interior designs are overlaid on the camera image.

[0932] Step 11:

[0933] When the user observes the real world through the AR glasses, the characters and interiors that have been changed as instructed are superimposed on the real world, allowing the user to experience passersby on their commute or the interior of their room as if they were replaced with virtual characters and designs.

[0934] Example 1

[0935] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0936] Systems that provide users with new entertainment experiences by replacing real-world objects with virtual characters and interiors require technology that can accurately understand user instructions and reflect them in real-time video. Furthermore, high accuracy and speed are required when systems integrate real-world video with virtual data, and technology is needed to provide users with a natural, seamless experience.

[0937] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0938] In this invention, the server includes means for receiving voice instructions from a user, means for converting the received voice instructions into text data and analyzing the instructions, means for capturing camera images, identifying real-world objects and obtaining position and orientation information, means for selecting replacement data based on the specified virtual character or interior, means for generating the selected replacement data using the object's position and orientation information, and means for displaying the generated data superimposed on the camera image. This makes it possible to quickly and accurately analyze the user's voice instructions and display virtual data that is highly accurately integrated with real-world images in real time.

[0939] "User" refers to an individual who uses the system to issue voice commands and enjoy the experience of replacing real-world objects with virtual characters and interiors.

[0940] "Voice instruction" refers to a command made by voice that a user issues to request some action or change to the system.

[0941] "Text data" refers to character string information that is generated by capturing voice instructions and converting them using voice recognition software.

[0942] "Camera footage" refers to real-world video data captured by a camera built into a device.

[0943] "Real-world objects" refer to physical objects that are within the user's field of view (e.g., pedestrians, vehicles, furniture, etc.).

[0944] "Position and orientation information" refers to information about the spatial position of an object in the real world, as well as its direction and orientation.

[0945] "Virtual character" refers to a computer-generated character that appears in place of a real-world object.

[0946] "Interior" refers to virtual data about the decoration and design of the space the user is in.

[0947] "Replacement Data" refers to information, including 3D models and animation data, used to replace real-world objects with virtual characters or interiors.

[0948] "Generated data" refers to AR data that is generated by the server based on the object's position and orientation information, and is used to apply virtual characters and interiors to the real world.

[0949] "Means for overlaying and displaying on camera images" refers to the method or technology by which the device receives the generated data, synthesizes it on camera images from the real world, and displays it to the user.

[0950] "Lens technology" refers to the optical techniques and devices used to integrate real-world images with generated data.

[0951] This invention is a system that uses AR glasses worn by the user to display real-world objects replaced with virtual characters and interiors. This system allows users to enjoy their daily commute and home life even more. Below, we will explain the specific program processing method and the hardware and software used.

[0952] System Overview

[0953] This system is primarily comprised of three components: a server, a device (including AR glasses), and a user. It combines various technologies to realize the process in which the user issues voice commands, real-world images are processed according to those commands, and virtual data is synthesized and displayed.

[0954] Hardware and Software

[0955] Device: AR glasses have built-in high-performance microphones and cameras, as well as a processor and necessary memory to process data in real time.

[0956] Server: A server with high processing power that uses natural language processing technology (e.g., OpenAI's GPT-4) and image recognition technology (e.g., YOLO and OpenCV).

[0957] Speech Recognition Software: Uses speech recognition technology such as Google Cloud Speech-to-Text to convert voice commands into text data.

[0958] Rendering engine: An engine such as Unity or Unreal Engine that integrates real-world images with virtual data for display.

[0959] Specific processing of the program

[0960] Recognizing voice commands

[0961] The user issues voice commands into the AR glasses. For example, they can give specific instructions such as "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street to characters on my commute." The device uses a built-in microphone to capture the voice and obtains the voice data in real time. The obtained voice data is then converted into text data using voice recognition software.

[0962] Analysis of voice instructions

[0963] The device sends the converted text data to a server, which then uses natural language processing technology to analyze the text data. From the analysis results, the server identifies the "subject" (people, room interior, etc.) and the "changes" (characters, Swiss lodge style, etc.).

[0964] Video Capture and Object Recognition

[0965] The device uses a built-in camera to capture images within the user's field of view. This video data is sent in real time to a server, which uses image recognition technology to identify real-world objects (such as pedestrians, vehicles, and furniture) and obtain the position and orientation information of each object.

[0966] Selection of replacement data

[0967] Based on the parsed instructions, the server selects replacement character and interior data, which is retrieved from a library containing the necessary 3D models and animation data.

[0968] AR data generation

[0969] The server generates AR data based on the acquired object position and orientation information, placing the selected character and interior data in the real world at the appropriate position and orientation. This data also includes information necessary for real-time rendering, such as lighting and shadows.

[0970] Viewing Data

[0971] The generated AR data is sent from the server to the device, which processes the received AR data in real time and displays it overlaid on the camera image. Through the AR glasses, the user can see that real-world objects have been replaced with the specified characters or interior.

[0972] Specific examples

[0973] Example 1: Change the room interior to Swiss lodge style

[0974] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[0975] 2. The device converts the voice into text data and sends it to the server.

[0976] 3. The server analyzes the data and identifies instructions for interior modifications.

[0977] 4. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[0978] 5. The server selects Swiss lodge-style interior data.

[0979] 6. The server combines the interior data with the room's location information to generate AR data.

[0980] 7. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. To the user, the interior of the room appears to have been transformed into a Swiss lodge.

[0981] Example 2: Change passersby in the city into characters during your commute

[0982] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[0983] 2. The device converts the voice into text data and sends it to the server.

[0984] 3. The server analyzes the data and identifies the passerby's replacement instructions.

[0985] 4. The device's camera captures images of the city, and the server recognizes passersby.

[0986] 5. The server selects the character data.

[0987] 6. The server combines the character data with the location information of passersby to generate AR data.

[0988] 7. The server sends the generated AR data to the device, which then displays it over the camera image. To the user, it appears as if passersby on their way to work have been transformed into characters.

[0989] Prompt Sentence Examples

[0990] 1. "How can I transform my room into a Swiss lodge?"

[0991] 2. "How do I direct the people on the street to become characters?"

[0992] This system allows users to enjoy a variety of entertainment experiences in their daily lives, improving their quality of life.

[0993] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0994] Step 1: Getting voice instructions

[0995] The user issues voice commands to the AR glasses. For example, they can give specific instructions such as "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street to characters on my commute." The device captures the voice using a built-in microphone. The input is the user's voice commands, and the output is voice data.

[0996] Step 2: Convert the audio data

[0997] The device converts the captured voice data into text data using voice recognition software (e.g., Google Cloud Speech-to-Text). It uses voice data as input and natural language processing technology to output text data. This conversion process is performed in real time.

[0998] Step 3: Analyzing voice commands

[0999] The device sends the converted text data to the server. The server uses natural language processing technology (for example, OpenAI's GPT-4) to analyze the text data. The input is the text data, and the output is the analysis results divided into the referent and the changes. From the analysis results, the "target" (for example, the room's interior) and the "changes" (for example, Swiss lodge style) are identified.

[1000] Step 4: Capture footage

[1001] The device uses the built-in camera to capture images within the user's field of view. The camera is activated and images are collected in real time. The input is the real-world image captured by the camera, and the output is the image data.

[1002] Step 5: Object Recognition

[1003] The device sends the captured video data to a server, which uses image recognition technology (e.g., YOLO or OpenCV) to identify real-world objects (e.g., furniture or passersby). The input is the video data, and the output is the object's position and orientation. The server recognizes the spatial coordinates and orientation of each object and collects the information.

[1004] Step 6: Select replacement data

[1005] The server selects replacement character and interior data based on the analyzed instructions, and obtains appropriate 3D models and animation data from the data library. The input is the analysis results and object position and orientation information, and the output is the selected 3D model and animation data.

[1006] Step 7: Generate AR data

[1007] The server generates AR data for placing the selected character and interior data in the real world based on the acquired object position and orientation information. This data also includes information necessary for real-time rendering, such as lighting and shadows. The input is the object position and orientation information, as well as the selected 3D model and animation data, and the output is the generated AR data.

[1008] Step 8: Send AR data

[1009] The server sends the generated AR data to the device. The data is compressed in an efficient format and transmitted with low latency. The input is the generated AR data, and the output is the transmitted AR data.

[1010] Step 9: Overlay on camera image

[1011] The device processes the received AR data in real time and overlays it on the camera image using a rendering engine (e.g., Unity or Unreal Engine). Through the AR glasses, the user can see that real-world objects have been replaced with virtual characters and interiors. The input is the transmitted AR data and real-world camera images, and the output is the combined image displayed in the user's field of view.

[1012] (Application example 1)

[1013] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1014] Systems that can quickly and effectively replace real-world interiors and objects with virtual characters and decorations are extremely important in the entertainment and marketing fields. However, existing systems have issues with not being able to change interiors or optimize product displays in real time according to seasons or events. There is also a lack of flexible systems that allow users to easily issue commands and have changes reflected instantly. This makes it difficult to provide an engaging customer experience in physical stores.

[1015] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1016] In this invention, the server includes means for acquiring voice instructions from a user, means for converting the acquired voice instructions into text data and analyzing the instruction content, means for capturing camera images, identifying real-world objects, and acquiring position and orientation information, means for selecting replacement data based on the designated character or interior, means for generating the selected replacement data using the object position and orientation information, means for displaying the generated data by overlaying it on the camera image, and means for generating and changing AR data of the interior and products of a physical store according to the season or an event. This enables physical stores to change the interior and product displays in real time to match the season or an event, which is expected to improve the customer experience.

[1017] "Voice instructions from the user" refers to a form in which the user inputs commands or requests by voice.

[1018] "Text data" is a data format in which voice instructions are converted into character string information.

[1019] The "instruction content" is a specific request or command given by the user through voice instructions.

[1020] "Camera footage" is real-world video data captured by a camera.

[1021] "Real-world objects" are actual objects such as furniture, products, and walls that appear in the camera image.

[1022] "Position and orientation information" is data on where an object is currently located and in what orientation it is facing.

[1023] "Replacement data" is virtual character and interior data that is roughly matched to real-world objects based on user instructions.

[1024] "Generated data" refers to virtual data that is processed and generated to be overlaid on camera images.

[1025] "Overlaying and displaying on camera images" means overlaying and displaying virtual data on images captured by a camera.

[1026] "According to the season or event" means changing the theme or decorations to suit the time of year or a particular event.

[1027] "Brick and mortar store interior" refers to the internal layout and decoration of an actual store.

[1028] "AR data of a product" is virtual data of a product displayed using augmented reality technology.

[1029] "Processing in real time" means processing data in immediate response to user instructions or changes in the environment.

[1030] "Lens technology" is a term that refers to the technology for integrating real-world images with virtual data.

[1031] "Integration" means bringing together different data and information.

[1032] This invention is a system that uses smart glasses or a smartphone worn by the user to display real-world objects replaced with virtual characters and interiors. This system enables real-world stores to change their interiors and product displays in real time according to the season or events, which is expected to improve the customer experience.

[1033] Specifically, it consists of the following elements:

[1034] Server: This receives voice instructions from the user and converts them into text data. A speech recognition API (e.g., Google Speech-to-Text API) is used. The acquired text data is then analyzed using natural language processing (NLP) technology to understand the instructions. An example of a natural language processing API is OpenAI's GPT-4. Based on the analyzed instructions, appropriate replacement data is selected. The server also receives camera images sent from the smart glasses or smartphone, identifies real-world objects using an image recognition API (e.g., Amazon Rekognition), and obtains information about the object's position and orientation. Based on this information, the server generates the selected replacement data and generates the augmented reality (AR) data needed to place the data in the real world at the optimal position and orientation.

[1035] Terminal (smart glasses or smartphone): The device allows the user to give voice commands, captures the voice, and sends it to the server. It also has the function of capturing camera images and sending them to the server. Furthermore, it overlays the AR data received from the server on the camera image in real time. A real-time rendering engine (e.g., Unity) is used for the display.

[1036] Examples:

[1037] Example 1: To change the interior of a store to suit the season, the user (store manager or staff member) issues a voice command to the smart glasses saying, "Decorate the interior of the store to suit the spring season." The smart glasses capture the voice and send it to the server. The server converts the voice into text data, analyzes it, and identifies it as "spring decoration." It receives camera footage, identifies objects in the store (shelves, products, walls, etc.), and acquires their position and orientation information. It selects spring decoration data, processes it in real time, and generates AR data. The smart glasses receive the generated data and display it overlaid on the camera footage from inside the store. In a similar manner, it is possible to change the interior of the store to suit the season or event.

[1038] Example prompt sentence:

[1039] "The interior of the store is decorated in the style of a cherry blossom garden."

[1040] "Change to Christmas decorations for the event."

[1041] This allows virtual changes to store interiors and merchandise to be made in real time, providing customers with an engaging experience.

[1042] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1043] Step 1:

[1044] The user issues a voice command. The user issues a voice command into the smart glasses or smartphone, such as "Decorate the interior of the store to match the spring season." The input is voice, and the output is voice data.

[1045] Step 2:

[1046] The device captures the audio and converts it to text data using a speech recognition API (e.g., Google Speech-to-Text API). The input here is audio data, and the output is text data. During the process, the audio is analyzed and converted into a string.

[1047] Step 3:

[1048] The terminal sends text data to the server. The input is text data, and the output is data sent to the server. The data is sent to the server using the terminal's communication function.

[1049] Step 4:

[1050] The server receives the text data and analyzes it using a natural language processing API (e.g., OpenAI's GPT-4). The input is text data, and the output is instructions as a result of the analysis (e.g., "Decorate the interior of the store in a spring-like style"). Natural language processing technology is used to understand the meaning of the text.

[1051] Step 5:

[1052] The server selects appropriate replacement data (spring decoration data) based on the analysis results. The input is the analysis results, and the output is the selected replacement data. Appropriate 3D models and animation data are selected from the library.

[1053] Step 6:

[1054] The device's camera captures images of the real world and sends the data to the server. The input is the camera image, and the output is the image data sent to the server. Images are captured and sent in real time.

[1055] Step 7:

[1056] The server receives the video data and uses an image recognition API (e.g., Amazon Rekognition) to identify real-world objects. The input is the video data, and the output is information about the identified objects (position and posture information). Objects are recognized through image recognition processing.

[1057] Step 8:

[1058] The server generates AR data based on the object's position and orientation information to apply the selected replacement data to the real world. The input is object information and replacement data, and the output is AR data. Rendering calculations are performed to place the data in the appropriate position.

[1059] Step 9:

[1060] The server sends the generated AR data to the device. The input is the generated AR data, and the output is the data sent to the device. The data is sent using the server's communication function.

[1061] Step 10:

[1062] The AR data received by the device is overlaid on the camera image in real time. The input is the AR data and the camera image, and the output is an augmented reality image that the user can see. The display is done using a real-time rendering engine (e.g., Unity).

[1063] The above steps make it possible to change the interior and product displays of physical stores in real time according to the season or events.

[1064] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1065] This invention is a system that uses AR glasses worn by the user to replace real-world objects with virtual characters and interiors. This invention includes a function that, by combining it with an emotion engine that recognizes the user's emotions, makes more adaptive changes to the display according to the user's emotional state. The program processing is explained below in natural language.

[1066] Program processing explanation

[1067] 1. Recognition of voice commands

[1068] The user issues voice commands to the AR glasses. Possible commands include "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street on my commute into characters." The device captures the voice input and converts it into text data. The text data is then sent to a server for analysis.

[1069] 2. Emotion Recognition by Emotion Engine

[1070] The device's built-in emotion engine recognizes emotions from the user's voice and facial expressions. For example, it analyzes the user's tone of voice and facial muscle movements to identify emotions such as joy, sadness, anger, and surprise.

[1071] 3. Instruction and emotion analysis

[1072] The server analyzes the acquired text data. It uses natural language processing technology to understand the meaning of the voice commands. It also analyzes emotion recognition data from the emotion engine to understand the user's emotional state. Based on the results of this analysis, it selects the most appropriate target (character or interior design).

[1073] 4. Video Capture and Object Recognition

[1074] The device's camera captures video from the user's point of view. The video data is sent in real time to a server. The server uses image recognition technology to identify real-world objects (passersby, vehicles, furniture, etc.). The server then determines the location and orientation of each object and stores them in a database.

[1075] 5. Selection of replacement data

[1076] The server selects the best replacement character and interior design data based on the analyzed instructions and the recognized emotion. For example, if the user is feeling sad, it selects an uplifting character and interior design. This data is retrieved from a library and includes the necessary 3D models and animation data.

[1077] 6. AR Data Generation

[1078] The server generates AR data based on the acquired object position and orientation information to display the selected character and interior in the real world. This data also includes information necessary for real-time rendering, such as lighting and shadows.

[1079] 7. Displaying Data

[1080] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image. From the user's perspective, this makes it appear as if real-world objects have been replaced with the specified characters or interior.

[1081] Specific examples

[1082] Example 1: Change the room interior to Swiss lodge style and adjust it according to the emotion.

[1083] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[1084] 2. The device converts the voice into text data and sends it to the server.

[1085] 3. The device's emotion engine recognizes the user's emotions and determines that they want to relax.

[1086] 4. The server analyzes the data and identifies instructions for interior modifications.

[1087] 5. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[1088] 6. Based on the emotional data, the server selects interior design data that is more relaxing, such as a Swiss lodge.

[1089] 7. The server combines the interior data with the room's location information to generate AR data.

[1090] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. The user feels as if the interior of the room has been transformed into a Swiss lodge, creating an even more relaxing atmosphere.

[1091] Example 2: Transforming pedestrians into characters during your commute and adjusting their emotions

[1092] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[1093] 2. The device converts the voice into text data and sends it to the server.

[1094] 3. The device's emotion engine recognizes the user's emotions and determines that they are "feeling stressed."

[1095] 4. The server analyzes the data and identifies the passerby's replacement instructions.

[1096] 5. The device's camera captures images of the city, and the server recognizes passersby.

[1097] 6. The server selects character data that has the effect of reducing stress based on the emotional data.

[1098] 7. The server combines the character data with the location information of passersby to generate AR data.

[1099] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. Passersby on the way to work are transformed into characters that help reduce stress, making the commute more enjoyable for the user.

[1100] This system will enable users to enjoy a variety of entertainment experiences tailored to their emotions in their daily lives, further improving their quality of life.

[1101] The processing flow will be explained below.

[1102] Step 1:

[1103] The user issues voice commands to the AR glasses, such as "Turn people on the street into characters on my commute."

[1104] Step 2:

[1105] The device captures voice commands and converts the voice data into digital form, which is then analyzed by an internal speech recognition engine and converted into text data.

[1106] Step 3:

[1107] The terminal transmits the converted text data to the server, which includes the content of the voice instructions.

[1108] Step 4:

[1109] The server analyzes the received text data, using natural language processing technology to understand the meaning of the voice commands and extract the "target (people, room interior, etc.)" and "change content (character, Swiss lodge style, etc.)."

[1110] Step 5:

[1111] The device's emotion engine recognizes emotions from the user's voice and facial expressions, analyzing the tone of voice and facial muscle movements to identify emotions such as joy, sadness, anger, and surprise.

[1112] Step 6:

[1113] The device transmits the recognized emotion data to the server.

[1114] Step 7:

[1115] The server analyzes the voice commands and emotional data and then selects the character and interior design that best suits the user's emotions and instructions. For example, if the user is feeling sad, it will select an uplifting character.

[1116] Step 8:

[1117] The device's camera captures images of the real world from the user's point of view, and the image data is sent to a server in real time.

[1118] Step 9:

[1119] The server analyzes the received video data and uses image recognition technology to identify objects in the video (passersby, vehicles, furniture, etc.), and determines the position and posture of each object and stores them in a database.

[1120] Step 10:

[1121] The server uses the object's position and orientation information to generate AR data for displaying the selected character and interior in the real world. This data also includes the necessary rendering information.

[1122] Step 11:

[1123] The server sends the generated AR data to the device.

[1124] Step 12:

[1125] The device processes the AR data received in real time, and selected characters and interior designs are overlaid on the camera image.

[1126] Step 13:

[1127] The app also adjusts based on emotions. For example, if the user is feeling stressed, characters and interior design that will help alleviate stress will be selected and reflected in the display.

[1128] Step 14:

[1129] When the user observes the real world through the AR glasses, characters and interiors that change according to voice commands and emotions are superimposed on the real world, allowing the user to experience passersby on their commute or the interior of their room as if they were replaced with virtual characters and designs.

[1130] Example 2

[1131] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1132] Conventional augmented reality (AR) systems can replace real-world objects with virtual characters or interiors based on user instructions, but they lack the ability to adaptively change the display according to the user's emotional state. This makes it difficult to provide an experience that is in line with the user's emotions. It is also difficult to process in real time and realize a display that includes natural lighting and shadows.

[1133] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1134] In this invention, the server includes means for receiving voice instructions from a user, means for converting the received voice instructions into text data and analyzing the instructions, means for recognizing emotions from the user's voice and facial expressions, means for capturing camera images, identifying real-world objects, and acquiring position and orientation information, means for selecting optimal replacement data based on the voice instructions and the recognized emotions, means for generating the selected replacement data using the object's position and orientation information, and means for displaying the generated data overlaid on the camera image, thereby enabling adaptive AR display according to the user's emotional state in real time.

[1135] A "user" is a person who uses the system to issue voice commands and replace real-world objects with virtual characters and interiors.

[1136] "Voice instructions" are voice requests or commands given by the user to the system.

[1137] "Text data" is character information converted from voice instructions using voice recognition technology.

[1138] "Emotion recognition means" is a technology that analyzes the user's voice and facial expressions to identify the user's emotional state (joy, sadness, anger, etc.).

[1139] "Camera footage" refers to real-world video data captured in real time by a camera mounted on a device.

[1140] "Object" refers to a concrete object that exists in the real world (e.g., a passerby, a vehicle, furniture, etc.).

[1141] "Position and orientation information" refers to data on the physical position and orientation of an object identified in a camera image.

[1142] "Replacement data" is digital data used to replace an object with a virtual character or virtual interior.

[1143] "Rendering information" is data that includes the representation of lighting, shadows, and other elements required for AR display.

[1144] "Real-time processing" refers to the process of analyzing, generating, and displaying data instantly, without delay.

[1145] MODE FOR CARRYING OUT THE INVENTION

[1146] This invention is a system that uses augmented reality (AR) glasses worn by the user to display real-world objects replaced with virtual characters and interiors. This system incorporates an emotion engine that recognizes the user's emotions and includes a function to adaptively change the display according to the user's emotional state.

[1147] System Configuration

[1148] Hardware and software used

[1149] Device: AR glasses

[1150] Camera: A camera for capturing images of the real world.

[1151] Microphone: A microphone for capturing user voice commands.

[1152] Display: A display for overlaying virtual objects onto real-world images.

[1153] Server: Responsible for calculation processing

[1154] Natural language processing technology: For example, Google Cloud Natural Language API

[1155] Speech recognition software: for example, Google Cloud Speech-to-Text

[1156] Emotion recognition software: Examples include IBM Watson Tone Analyzer and Microsoft Azure Face API

[1157] Image recognition technology: For example, Google Cloud Vision API

[1158] AR data generation platform: Examples include Unity and Unreal Engine

[1159] Program processing

[1160] Acquiring and analyzing voice instructions

[1161] The user issues voice commands to the AR glasses, such as "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street to characters on my commute." The device captures these commands and converts them into text using voice recognition software. The converted text is then sent to a server where it is analyzed using natural language processing technology.

[1162] emotion recognition

[1163] The device's built-in emotion engine recognizes emotions such as joy, sadness, anger, and surprise from the user's voice and facial expressions. Speech and facial recognition software analyzes this data to identify the user's emotional state.

[1164] Video Capture and Object Recognition

[1165] The device's camera captures images from the user's point of view in real time and transmits the image data to a server, which uses image recognition technology to identify real-world objects and determine their position and orientation.

[1166] Selecting and generating replacement data

[1167] The server selects the most suitable virtual character and interior design data based on the analyzed voice commands and emotion recognition data. This data is selected from a library stored in advance, and lighting and shadow information required for real-time rendering is also synchronized.

[1168] Viewing Data

[1169] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image. For example, if the user is feeling sad, an uplifting character or interior design will be displayed overlaid on the real world.

[1170] Specific examples

[1171] Example 1: Change the room interior to Swiss lodge style and adjust it according to the emotion.

[1172] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[1173] 2. The device converts the voice into text data and sends it to the server.

[1174] 3. The device's emotion engine recognizes the user's emotion as "I want to relax."

[1175] 4. The server analyzes the data and identifies instructions for interior modifications.

[1176] 5. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[1177] 6. The server selects interior design data that is more relaxing, such as a Swiss lodge.

[1178] 7. The server combines the interior data with the room's location information to generate AR data.

[1179] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. The user feels as if the interior of the room has been transformed into a Swiss lodge, creating an even more relaxing atmosphere.

[1180] Example 2: Transforming pedestrians into characters during your commute and adjusting their emotions

[1181] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[1182] 2. The device converts the voice into text data and sends it to the server.

[1183] 3. The device's emotion engine recognizes the user's emotion as "feeling stressed."

[1184] 4. The server analyzes the data and identifies the passerby's replacement instructions.

[1185] 5. The device's camera captures images of the city, and the server recognizes passersby.

[1186] 6. The server selects character data that has the effect of reducing stress based on the emotional data.

[1187] 7. The server combines the character data with the location information of passersby to generate AR data.

[1188] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. Passersby on the user's commute are transformed into characters that help reduce stress, making the commute more enjoyable.

[1189] In this way, the system of the present invention provides a variety of entertainment experiences that correspond to the user's emotions in their daily lives, contributing to improving the quality of life.

[1190] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1191] Step 1:

[1192] The user issues voice commands to the AR glasses. For example, they might say, "Decorate the room in a Swiss lodge style." The device uses a microphone to capture the user's voice. This voice data becomes the input.

[1193] Step 2:

[1194] The device converts the captured voice data into text data using voice recognition software (e.g., voice recognition API). The input is voice data and the output is text data. The converted text data is sent to the server.

[1195] Step 3:

[1196] The server analyzes the received text data using natural language processing technology (e.g., natural language processing API). The input is text data, and the output is the analysis result. Here, the content of the voice instruction is understood.

[1197] Step 4:

[1198] The device uses an emotion engine to recognize emotions from the user's voice and facial expressions. The input is the user's voice data and camera footage, and the output is emotion data. Emotion recognition software (e.g., emotion analysis API) is used to identify emotions from voice tone and facial muscle movements.

[1199] Step 5:

[1200] The device's camera captures the image from the user's point of view in real time. The input is the real-world camera image, and the output is the captured image data. The image data is sent to the server where it is processed.

[1201] Step 6:

[1202] The server analyzes the captured video data using image recognition technology (e.g., image analysis API). The input is the video data, and the output is the position and posture information of the identified objects. This information is stored in a database.

[1203] Step 7:

[1204] The server selects the optimal replacement character and interior data based on the analyzed voice instructions and emotion recognition data. The input is voice instruction data and emotion data, and the output is the selected replacement data. It also obtains appropriate 3D models and animation data from the library.

[1205] Step 8:

[1206] The server generates AR data based on the object's position and orientation information to display the selected character and interior in real-world footage. The input is the object's position and the selected data, and the output is the generated AR data. Lighting and shadow information is also incorporated.

[1207] Step 9:

[1208] The server sends the generated AR data to the device. The device displays the received AR data overlaid on the camera image. The input is AR data, and the output is an AR display overlaid on the image of the real world. Through this, users can experience virtual characters and interiors as if they were in the real world.

[1209] The system utilizes generative AI models and prompts to provide a rich and adaptive entertainment experience that responds to the user's emotions, while also achieving natural AR display in real time, improving quality of life.

[1210] (Application example 2)

[1211] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1212] In recent years, advances in augmented reality (AR) technology have enabled users to combine the real world with virtual elements for enjoyment. However, current systems simply replace real-world objects with virtual characters or interiors, and are unable to adaptively change the display in response to the user's emotions. Therefore, there is a need for the development of systems that can take user emotions into account and provide greater satisfaction and relaxation.

[1213] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring voice instructions from the user, means for converting the acquired voice instructions into text data and analyzing the instruction content, means including an emotion engine for recognizing the user's emotions, means for capturing camera footage, identifying real-world objects, and acquiring position and orientation information, means for selecting replacement data based on the emotion data and the specified character or interior, means for generating the selected replacement data using the object's position and orientation information, and means for overlaying and displaying the generated data on the camera footage. This enables adaptive virtual replacement of real-world objects based on the user's emotions.

[1214] "Means for acquiring voice instructions from the user" refers to devices or software that detect commands spoken by the user to a wearable device such as AR glasses and incorporate them into the system.

[1215] "Means for converting acquired voice instructions into text data and analyzing the content of the instructions" refers to technology or devices that use voice recognition technology to convert voice data into character data, and then analyze the character data to understand the content of the instructions.

[1216] The "means including an emotion engine for recognizing the user's emotions" refers to a software or hardware system for recognizing the user's current emotional state in real time through analysis of voice and facial expressions.

[1217] "Means for capturing camera images, identifying real-world objects, and acquiring position and orientation information" refers to technology and devices that use a camera to capture real-world images, identify objects in the images using image recognition technology, and determine their position and orientation.

[1218] "Means for selecting replacement data based on emotional data and specified character or interior" refers to an algorithm or device for selecting appropriate character or interior data in accordance with the user's emotional state and voice instructions.

[1219] "Means for generating selected replacement data using object position and orientation information" refers to technology or devices that generate data for appropriately placing selected character and interior data in the real world based on the position and orientation information of identified objects.

[1220] "Means for displaying generated data by overlaying it on camera images" refers to technology or devices for displaying generated virtual content by overlaying it on real-world images captured by a camera.

[1221] System Overview

[1222] This system uses AR glasses worn by the user to replace real-world objects with virtual characters and interiors. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it also includes a function to adaptively change the display according to the user's emotional state.

[1223] Hardware and software used

[1224] Hardware: AR glasses (e.g. Microsoft HoloLens), smart devices with cameras, servers

[1225] Software: Speech recognition engine (e.g., Google Cloud Speech-to-Text), emotion recognition engine (e.g., Microsoft Azure Emotion API), image recognition technology (e.g., YOLO, TensorFlow)

[1226] Program Processing Overview

[1227] Obtaining voice commands from the user

[1228] A microphone installed on the device captures voice commands from the user. For example, if a user says, "Make my kitchen look like a Parisian cafe," that voice input is captured.

[1229] Analysis of voice instructions

[1230] The captured voice data is converted into text data by a voice recognition engine, and the converted text data is sent to a server where the instruction content is analyzed by an analysis engine.

[1231] Emotion recognition

[1232] The device's built-in emotion recognition engine analyzes the user's voice and facial expressions in real time to recognize their emotions. For example, if it recognizes that the user is in a relaxed mood, appropriate data will be selected.

[1233] Camera image capture and object recognition

[1234] The device's camera captures video within the user's field of view, and the video data is sent in real time to a server, which uses image recognition technology to identify real-world objects and determine their position and orientation.

[1235] Selection of replacement data

[1236] The server selects the optimal replacement data (characters and interior design) based on the analyzed voice commands and the recognized emotions. For example, it may select a Parisian cafe-style interior design that has a relaxing effect.

[1237] AR data generation

[1238] The server generates AR data based on the acquired object position and orientation information, including lighting and shadow information, to display the selected character and interior in the real world.

[1239] Viewing Data

[1240] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image, making it appear to the user that real-world objects have been replaced with the specified characters or interior.

[1241] Specific examples

[1242] Example 1: Remodeling your kitchen to look like a Parisian cafe

[1243] 1. The user issues a voice command such as, "Make my kitchen look like a Parisian cafe."

[1244] 2. The device converts the voice into text data and sends it to the server.

[1245] 3. The emotion engine determines that the user wants to relax.

[1246] 4. The server analyzes the data and identifies instructions for interior modifications.

[1247] 5. The device's camera captures images of the kitchen, and the server recognizes the location of the furniture.

[1248] 6. The server selects Parisian cafe-style interior data that has a relaxing effect.

[1249] 7. The server combines the interior data with the kitchen location information to generate AR data.

[1250] 8. The server sends the generated AR data to the device, which then overlays it on the camera image, allowing the user to enjoy a relaxing Parisian cafe-style kitchen.

[1251] Examples of prompt statements

[1252] You are a programmer tasked with designing a system for a food delivery AR glasses application that changes the kitchen interior to a virtual theme (e.g., Parisian cafe) and automatically adjusts to the user's emotions. Write a program procedure that meets the following requirements:

[1253] Requirements:

[1254] 1. The user issues a voice command (e.g., "Make my kitchen look like a Parisian cafe").

[1255] 2. Convert the voice instructions into text data and send it to the server.

[1256] 3. Recognize user emotions and automatically select appropriate themes.

[1257] 4. The AR glasses' camera captures images of the kitchen, and the server recognizes the position of the furniture.

[1258] 5. The server selects appropriate theme data (e.g., interior data of a Parisian cafe).

[1259] 6. The server generates the AR data and sends it to the device.

[1260] 7. AR data is overlaid on the camera image on the device.

[1261] Follow this procedure to design your program.

[1262] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1263] Step 1:

[1264] The user utters the voice command "Make my kitchen look like a Parisian cafe." The input is the user's voice, and the output is the captured voice data. The device's microphone captures this voice and stores it in its internal memory.

[1265] Step 2:

[1266] To convert the voice data into text data, the device uses the Google Cloud Speech-to-Text engine. The input is the captured voice data, and the output is the converted text data, which is sent from the device to the server for the next analysis step.

[1267] Step 3:

[1268] The server parses the received text data. Using a natural language processing engine (e.g., various NLP libraries), the input is the transformed text data and the output is the parsed results, which include instructions for the user to transform their kitchen into a Parisian cafe style.

[1269] Step 4:

[1270] The device's built-in emotion recognition engine recognizes the user's emotions in real time. The input is the user's voice tone and facial expression data, and the output is the user's emotional data. This emotional data is sent to the server and processed together with the analysis results.

[1271] Step 5:

[1272] The device's camera captures images of the kitchen and transmits the image data to the server in real time. The input is the captured camera image, and the output is the transmitted image data.

[1273] Step 6:

[1274] The server analyzes the transmitted video data using image recognition technology (e.g., YOLO, TensorFlow). The input is the camera's video data, and the output is the position and orientation information of objects in the kitchen. This allows the server to identify the positions of furniture and other objects in the kitchen.

[1275] Step 7:

[1276] The server selects the optimal replacement data (e.g., interior design data of a Parisian cafe) based on the analyzed voice instructions and the recognized emotion data. The input is the analysis results of the voice instructions and the emotion data, and the output is the selected replacement data.

[1277] Step 8:

[1278] The server uses the object's position and orientation information to generate AR data for displaying the selected interior data in the real world. The input is the object's position and orientation information and the selected replacement data, and the output is the generated AR data. This data includes, for example, 3D models and lighting information.

[1279] Step 9:

[1280] The generated AR data is sent from the server to the device. The input is the generated AR data, and the output is the data sent to the device. The device processes the received AR data in real time and displays it overlaid on the camera image. This makes the kitchen appear to have been transformed into a Parisian cafe in the user's field of vision.

[1281] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1282] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1283] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1284] [Fourth embodiment]

[1285] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1286] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1287] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1288] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1289] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1290] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1291] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1292] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1293] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1294] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1295] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1296] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1297] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1298] This invention is a system that uses AR glasses worn by the user to replace real-world objects with virtual characters and interiors. This system allows users to enjoy their daily commute and their home life even more. The program processing is explained below in natural language.

[1299] Program processing explanation

[1300] 1. Recognition of voice commands

[1301] The user issues voice commands to the AR glasses. Possible commands include "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street on my commute into characters." The device captures the voice input and converts it into text data. The text data is then sent to a server for analysis.

[1302] 2. Parsing the instructions

[1303] The server analyzes the acquired text data. It uses natural language processing technology to understand the meaning of the instructions and identify which objects to replace. This analysis clarifies the "target (people, room interior, etc.)" and "changes (characters, Swiss lodge style, etc.)."

[1304] 3. Video Capture and Object Recognition

[1305] The device uses a camera to capture images within the user's field of view. This video data is sent in real time to a server, which uses image recognition technology to identify real-world objects (such as pedestrians, vehicles, and furniture) and obtain the position and orientation information of each object.

[1306] 4. Selection of replacement data

[1307] Based on the parsed instructions, the server selects replacement character and interior data, which is retrieved from a library containing the necessary 3D models and animation data.

[1308] 5. AR Data Generation

[1309] The server generates AR data based on the acquired object position and orientation information, placing the selected character and interior data in the real world at the appropriate position and orientation. This data also includes information necessary for real-time rendering, such as lighting and shadows.

[1310] 6. Displaying Data

[1311] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image. From the user's perspective, this makes it appear as if real-world objects have been replaced with the specified characters or interior.

[1312] Specific examples

[1313] Example 1: Change the room interior to Swiss lodge style

[1314] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[1315] 2. The device converts the voice into text data and sends it to the server.

[1316] 3. The server analyzes the data and identifies instructions for interior modifications.

[1317] 4. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[1318] 5. The server selects Swiss lodge-style interior data.

[1319] 6. The server combines the interior data with the room's location information to generate AR data.

[1320] 7. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. To the user, the interior of the room appears to have been transformed into a Swiss lodge.

[1321] Example 2: Change passersby in the city into characters during your commute

[1322] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[1323] 2. The device converts the voice into text data and sends it to the server.

[1324] 3. The server analyzes the data and identifies the passerby's replacement instructions.

[1325] 4. The device's camera captures images of the city, and the server recognizes passersby.

[1326] 5. The server selects the character data.

[1327] 6. The server combines the character data with the location information of passersby to generate AR data.

[1328] 7. The server sends the generated AR data to the device, which then displays it over the camera image. To the user, it appears as if passersby on their way to work have been transformed into characters.

[1329] This system allows users to enjoy a variety of entertainment experiences in their daily lives, improving their quality of life.

[1330] The processing flow will be explained below.

[1331] Step 1:

[1332] The user issues voice commands to the AR glasses, such as "Turn people on the street into characters on my commute."

[1333] Step 2:

[1334] The device recognizes voice commands and converts the voice data into digital form. An internal voice recognition engine analyzes this digital voice data and converts it into text data.

[1335] Step 3:

[1336] The terminal transmits the converted text data to the server, which includes the content of the voice instructions.

[1337] Step 4:

[1338] The server analyzes the received text data, using natural language processing technology to understand the meaning of the voice commands and extract the "target (people, room interior, etc.)" and "change content (character, Swiss lodge style, etc.)."

[1339] Step 5:

[1340] The device's camera captures images of the real world from the user's point of view, and the image data is sent to a server in real time.

[1341] Step 6:

[1342] The server analyzes the received video data and uses image recognition technology to identify objects in the video (passersby, vehicles, furniture, etc.), and identifies the position and posture information of each object and stores it in a database.

[1343] Step 7:

[1344] Based on the analysis results, the server selects the requested character and interior data, and retrieves the corresponding 3D model and animation data from the library.

[1345] Step 8:

[1346] The server uses the position and orientation information of the identified objects to generate AR data for displaying the selected characters and interiors in the real world, including the necessary rendering information.

[1347] Step 9:

[1348] The server sends the generated AR data to the device.

[1349] Step 10:

[1350] The device processes the AR data received in real time, and selected characters and interior designs are overlaid on the camera image.

[1351] Step 11:

[1352] When the user observes the real world through the AR glasses, the characters and interiors that have been changed as instructed are superimposed on the real world, allowing the user to experience passersby on their commute or the interior of their room as if they were replaced with virtual characters and designs.

[1353] Example 1

[1354] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1355] Systems that provide users with new entertainment experiences by replacing real-world objects with virtual characters and interiors require technology that can accurately understand user instructions and reflect them in real-time video. Furthermore, high accuracy and speed are required when systems integrate real-world video with virtual data, and technology is needed to provide users with a natural, seamless experience.

[1356] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1357] In this invention, the server includes means for receiving voice instructions from a user, means for converting the received voice instructions into text data and analyzing the instructions, means for capturing camera images, identifying real-world objects and obtaining position and orientation information, means for selecting replacement data based on the specified virtual character or interior, means for generating the selected replacement data using the object's position and orientation information, and means for displaying the generated data superimposed on the camera image. This makes it possible to quickly and accurately analyze the user's voice instructions and display virtual data that is highly accurately integrated with real-world images in real time.

[1358] "User" refers to an individual who uses the system to issue voice commands and enjoy the experience of replacing real-world objects with virtual characters and interiors.

[1359] "Voice instruction" refers to a command made by voice that a user issues to request some action or change to the system.

[1360] "Text data" refers to character string information that is generated by capturing voice instructions and converting them using voice recognition software.

[1361] "Camera footage" refers to real-world video data captured by a camera built into a device.

[1362] "Real-world objects" refer to physical objects that are within the user's field of view (e.g., pedestrians, vehicles, furniture, etc.).

[1363] "Position and orientation information" refers to information about the spatial position of an object in the real world, as well as its direction and orientation.

[1364] "Virtual character" refers to a computer-generated character that appears in place of a real-world object.

[1365] "Interior" refers to virtual data about the decoration and design of the space the user is in.

[1366] "Replacement Data" refers to information, including 3D models and animation data, used to replace real-world objects with virtual characters or interiors.

[1367] "Generated data" refers to AR data that is generated by the server based on the object's position and orientation information, and is used to apply virtual characters and interiors to the real world.

[1368] "Means for overlaying and displaying on camera images" refers to the method or technology by which the device receives the generated data, synthesizes it on camera images from the real world, and displays it to the user.

[1369] "Lens technology" refers to the optical techniques and devices used to integrate real-world images with generated data.

[1370] This invention is a system that uses AR glasses worn by the user to display real-world objects replaced with virtual characters and interiors. This system allows users to enjoy their daily commute and home life even more. Below, we will explain the specific program processing method and the hardware and software used.

[1371] System Overview

[1372] This system is primarily comprised of three components: a server, a device (including AR glasses), and a user. It combines various technologies to realize the process in which the user issues voice commands, real-world images are processed according to those commands, and virtual data is synthesized and displayed.

[1373] Hardware and Software

[1374] Device: AR glasses have built-in high-performance microphones and cameras, as well as a processor and necessary memory to process data in real time.

[1375] Server: A server with high processing power that uses natural language processing technology (e.g., OpenAI's GPT-4) and image recognition technology (e.g., YOLO and OpenCV).

[1376] Speech Recognition Software: Uses speech recognition technology such as Google Cloud Speech-to-Text to convert voice commands into text data.

[1377] Rendering engine: An engine such as Unity or Unreal Engine that integrates real-world images with virtual data for display.

[1378] Specific processing of the program

[1379] Recognizing voice commands

[1380] The user issues voice commands into the AR glasses. For example, they can give specific instructions such as "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street to characters on my commute." The device uses a built-in microphone to capture the voice and obtains the voice data in real time. The obtained voice data is then converted into text data using voice recognition software.

[1381] Analysis of voice instructions

[1382] The device sends the converted text data to a server, which then uses natural language processing technology to analyze the text data. From the analysis results, the server identifies the "subject" (people, room interior, etc.) and the "changes" (characters, Swiss lodge style, etc.).

[1383] Video Capture and Object Recognition

[1384] The device uses a built-in camera to capture images within the user's field of view. This video data is sent in real time to a server, which uses image recognition technology to identify real-world objects (such as pedestrians, vehicles, and furniture) and obtain the position and orientation information of each object.

[1385] Selection of replacement data

[1386] Based on the parsed instructions, the server selects replacement character and interior data, which is retrieved from a library containing the necessary 3D models and animation data.

[1387] AR data generation

[1388] The server generates AR data based on the acquired object position and orientation information, placing the selected character and interior data in the real world at the appropriate position and orientation. This data also includes information necessary for real-time rendering, such as lighting and shadows.

[1389] Viewing Data

[1390] The generated AR data is sent from the server to the device, which processes the received AR data in real time and displays it overlaid on the camera image. Through the AR glasses, the user can see that real-world objects have been replaced with the specified characters or interior.

[1391] Specific examples

[1392] Example 1: Change the room interior to Swiss lodge style

[1393] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[1394] 2. The device converts the voice into text data and sends it to the server.

[1395] 3. The server analyzes the data and identifies instructions for interior modifications.

[1396] 4. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[1397] 5. The server selects Swiss lodge-style interior data.

[1398] 6. The server combines the interior data with the room's location information to generate AR data.

[1399] 7. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. To the user, the interior of the room appears to have been transformed into a Swiss lodge.

[1400] Example 2: Change passersby in the city into characters during your commute

[1401] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[1402] 2. The device converts the voice into text data and sends it to the server.

[1403] 3. The server analyzes the data and identifies the passerby's replacement instructions.

[1404] 4. The device's camera captures images of the city, and the server recognizes passersby.

[1405] 5. The server selects the character data.

[1406] 6. The server combines the character data with the location information of passersby to generate AR data.

[1407] 7. The server sends the generated AR data to the device, which then displays it over the camera image. To the user, it appears as if passersby on their way to work have been transformed into characters.

[1408] Prompt Sentence Examples

[1409] 1. "How can I transform my room into a Swiss lodge?"

[1410] 2. "How do I direct the people on the street to become characters?"

[1411] This system allows users to enjoy a variety of entertainment experiences in their daily lives, improving their quality of life.

[1412] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1413] Step 1: Getting voice instructions

[1414] The user issues voice commands to the AR glasses. For example, they can give specific instructions such as "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street to characters on my commute." The device captures the voice using a built-in microphone. The input is the user's voice commands, and the output is voice data.

[1415] Step 2: Convert the audio data

[1416] The device converts the captured voice data into text data using voice recognition software (e.g., Google Cloud Speech-to-Text). It uses voice data as input and natural language processing technology to output text data. This conversion process is performed in real time.

[1417] Step 3: Analyzing voice commands

[1418] The device sends the converted text data to the server. The server uses natural language processing technology (for example, OpenAI's GPT-4) to analyze the text data. The input is the text data, and the output is the analysis results divided into the referent and the changes. From the analysis results, the "target" (for example, the room's interior) and the "changes" (for example, Swiss lodge style) are identified.

[1419] Step 4: Capture footage

[1420] The device uses the built-in camera to capture images within the user's field of view. The camera is activated and images are collected in real time. The input is the real-world image captured by the camera, and the output is the image data.

[1421] Step 5: Object Recognition

[1422] The device sends the captured video data to a server, which uses image recognition technology (e.g., YOLO or OpenCV) to identify real-world objects (e.g., furniture or passersby). The input is the video data, and the output is the object's position and orientation. The server recognizes the spatial coordinates and orientation of each object and collects the information.

[1423] Step 6: Select replacement data

[1424] The server selects replacement character and interior data based on the analyzed instructions, and obtains appropriate 3D models and animation data from the data library. The input is the analysis results and object position and orientation information, and the output is the selected 3D model and animation data.

[1425] Step 7: Generate AR data

[1426] The server generates AR data for placing the selected character and interior data in the real world based on the acquired object position and orientation information. This data also includes information necessary for real-time rendering, such as lighting and shadows. The input is the object position and orientation information, as well as the selected 3D model and animation data, and the output is the generated AR data.

[1427] Step 8: Send AR data

[1428] The server sends the generated AR data to the device. The data is compressed in an efficient format and transmitted with low latency. The input is the generated AR data, and the output is the transmitted AR data.

[1429] Step 9: Overlay on camera image

[1430] The device processes the received AR data in real time and overlays it on the camera image using a rendering engine (e.g., Unity or Unreal Engine). Through the AR glasses, the user can see that real-world objects have been replaced with virtual characters and interiors. The input is the transmitted AR data and real-world camera images, and the output is the combined image displayed in the user's field of view.

[1431] (Application example 1)

[1432] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1433] Systems that can quickly and effectively replace real-world interiors and objects with virtual characters and decorations are extremely important in the entertainment and marketing fields. However, existing systems have issues with not being able to change interiors or optimize product displays in real time according to seasons or events. There is also a lack of flexible systems that allow users to easily issue commands and have changes reflected instantly. This makes it difficult to provide an engaging customer experience in physical stores.

[1434] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1435] In this invention, the server includes means for acquiring voice instructions from a user, means for converting the acquired voice instructions into text data and analyzing the instruction content, means for capturing camera images, identifying real-world objects, and acquiring position and orientation information, means for selecting replacement data based on the designated character or interior, means for generating the selected replacement data using the object position and orientation information, means for displaying the generated data by overlaying it on the camera image, and means for generating and changing AR data of the interior and products of a physical store according to the season or an event. This enables physical stores to change the interior and product displays in real time to match the season or an event, which is expected to improve the customer experience.

[1436] "Voice instructions from the user" refers to a form in which the user inputs commands or requests by voice.

[1437] "Text data" is a data format in which voice instructions are converted into character string information.

[1438] The "instruction content" is a specific request or command given by the user through voice instructions.

[1439] "Camera footage" is real-world video data captured by a camera.

[1440] "Real-world objects" are actual objects such as furniture, products, and walls that appear in the camera image.

[1441] "Position and orientation information" is data on where an object is currently located and in what orientation it is facing.

[1442] "Replacement data" is virtual character and interior data that is roughly matched to real-world objects based on user instructions.

[1443] "Generated data" refers to virtual data that is processed and generated to be overlaid on camera images.

[1444] "Overlaying and displaying on camera images" means overlaying and displaying virtual data on images captured by a camera.

[1445] "According to the season or event" means changing the theme or decorations to suit the time of year or a particular event.

[1446] "Brick and mortar store interior" refers to the internal layout and decoration of an actual store.

[1447] "AR data of a product" is virtual data of a product displayed using augmented reality technology.

[1448] "Processing in real time" means processing data in immediate response to user instructions or changes in the environment.

[1449] "Lens technology" is a term that refers to the technology for integrating real-world images with virtual data.

[1450] "Integration" means bringing together different data and information.

[1451] This invention is a system that uses smart glasses or a smartphone worn by the user to display real-world objects replaced with virtual characters and interiors. This system enables real-world stores to change their interiors and product displays in real time according to the season or events, which is expected to improve the customer experience.

[1452] Specifically, it consists of the following elements:

[1453] Server: This receives voice instructions from the user and converts them into text data. A speech recognition API (e.g., Google Speech-to-Text API) is used. The acquired text data is then analyzed using natural language processing (NLP) technology to understand the instructions. An example of a natural language processing API is OpenAI's GPT-4. Based on the analyzed instructions, appropriate replacement data is selected. The server also receives camera images sent from the smart glasses or smartphone, identifies real-world objects using an image recognition API (e.g., Amazon Rekognition), and obtains information about the object's position and orientation. Based on this information, the server generates the selected replacement data and generates the augmented reality (AR) data needed to place the data in the real world at the optimal position and orientation.

[1454] Terminal (smart glasses or smartphone): The device allows the user to give voice commands, captures the voice, and sends it to the server. It also has the function of capturing camera images and sending them to the server. Furthermore, it overlays the AR data received from the server on the camera image in real time. A real-time rendering engine (e.g., Unity) is used for the display.

[1455] Examples:

[1456] Example 1: To change the interior of a store to suit the season, the user (store manager or staff member) issues a voice command to the smart glasses saying, "Decorate the interior of the store to suit the spring season." The smart glasses capture the voice and send it to the server. The server converts the voice into text data, analyzes it, and identifies it as "spring decoration." It receives camera footage, identifies objects in the store (shelves, products, walls, etc.), and acquires their position and orientation information. It selects spring decoration data, processes it in real time, and generates AR data. The smart glasses receive the generated data and display it overlaid on the camera footage from inside the store. In a similar manner, it is possible to change the interior of the store to suit the season or event.

[1457] Example prompt sentence:

[1458] "The interior of the store is decorated in the style of a cherry blossom garden."

[1459] "Change to Christmas decorations for the event."

[1460] This allows virtual changes to store interiors and merchandise to be made in real time, providing customers with an engaging experience.

[1461] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1462] Step 1:

[1463] The user issues a voice command. The user issues a voice command into the smart glasses or smartphone, such as "Decorate the interior of the store to match the spring season." The input is voice, and the output is voice data.

[1464] Step 2:

[1465] The device captures the audio and converts it to text data using a speech recognition API (e.g., Google Speech-to-Text API). The input here is audio data, and the output is text data. During the process, the audio is analyzed and converted into a string.

[1466] Step 3:

[1467] The terminal sends text data to the server. The input is text data, and the output is data sent to the server. The data is sent to the server using the terminal's communication function.

[1468] Step 4:

[1469] The server receives the text data and analyzes it using a natural language processing API (e.g., OpenAI's GPT-4). The input is text data, and the output is instructions as a result of the analysis (e.g., "Decorate the interior of the store in a spring-like style"). Natural language processing technology is used to understand the meaning of the text.

[1470] Step 5:

[1471] The server selects appropriate replacement data (spring decoration data) based on the analysis results. The input is the analysis results, and the output is the selected replacement data. Appropriate 3D models and animation data are selected from the library.

[1472] Step 6:

[1473] The device's camera captures images of the real world and sends the data to the server. The input is the camera image, and the output is the image data sent to the server. Images are captured and sent in real time.

[1474] Step 7:

[1475] The server receives the video data and uses an image recognition API (e.g., Amazon Rekognition) to identify real-world objects. The input is the video data, and the output is information about the identified objects (position and posture information). Objects are recognized through image recognition processing.

[1476] Step 8:

[1477] The server generates AR data based on the object's position and orientation information to apply the selected replacement data to the real world. The input is object information and replacement data, and the output is AR data. Rendering calculations are performed to place the data in the appropriate position.

[1478] Step 9:

[1479] The server sends the generated AR data to the device. The input is the generated AR data, and the output is the data sent to the device. The data is sent using the server's communication function.

[1480] Step 10:

[1481] The AR data received by the device is overlaid on the camera image in real time. The input is the AR data and the camera image, and the output is an augmented reality image that the user can see. The display is done using a real-time rendering engine (e.g., Unity).

[1482] The above steps make it possible to change the interior and product displays of physical stores in real time according to the season or events.

[1483] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1484] This invention is a system that uses AR glasses worn by the user to replace real-world objects with virtual characters and interiors. This invention includes a function that, by combining it with an emotion engine that recognizes the user's emotions, makes more adaptive changes to the display according to the user's emotional state. The program processing is explained below in natural language.

[1485] Program processing explanation

[1486] 1. Recognition of voice commands

[1487] The user issues voice commands to the AR glasses. Possible commands include "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street on my commute into characters." The device captures the voice input and converts it into text data. The text data is then sent to a server for analysis.

[1488] 2. Emotion Recognition by Emotion Engine

[1489] The device's built-in emotion engine recognizes emotions from the user's voice and facial expressions. For example, it analyzes the user's tone of voice and facial muscle movements to identify emotions such as joy, sadness, anger, and surprise.

[1490] 3. Instruction and emotion analysis

[1491] The server analyzes the acquired text data. It uses natural language processing technology to understand the meaning of the voice commands. It also analyzes emotion recognition data from the emotion engine to understand the user's emotional state. Based on the results of this analysis, it selects the most appropriate target (character or interior design).

[1492] 4. Video Capture and Object Recognition

[1493] The device's camera captures video from the user's point of view. The video data is sent in real time to a server. The server uses image recognition technology to identify real-world objects (passersby, vehicles, furniture, etc.). The server then determines the location and orientation of each object and stores them in a database.

[1494] 5. Selection of replacement data

[1495] The server selects the best replacement character and interior design data based on the analyzed instructions and the recognized emotion. For example, if the user is feeling sad, it selects an uplifting character and interior design. This data is retrieved from a library and includes the necessary 3D models and animation data.

[1496] 6. AR Data Generation

[1497] The server generates AR data based on the acquired object position and orientation information to display the selected character and interior in the real world. This data also includes information necessary for real-time rendering, such as lighting and shadows.

[1498] 7. Displaying Data

[1499] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image. From the user's perspective, this makes it appear as if real-world objects have been replaced with the specified characters or interior.

[1500] Specific examples

[1501] Example 1: Change the room interior to Swiss lodge style and adjust it according to the emotion.

[1502] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[1503] 2. The device converts the voice into text data and sends it to the server.

[1504] 3. The device's emotion engine recognizes the user's emotions and determines that they want to relax.

[1505] 4. The server analyzes the data and identifies instructions for interior modifications.

[1506] 5. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[1507] 6. Based on the emotional data, the server selects interior design data that is more relaxing, such as a Swiss lodge.

[1508] 7. The server combines the interior data with the room's location information to generate AR data.

[1509] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. The user feels as if the interior of the room has been transformed into a Swiss lodge, creating an even more relaxing atmosphere.

[1510] Example 2: Transforming pedestrians into characters during your commute and adjusting their emotions

[1511] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[1512] 2. The device converts the voice into text data and sends it to the server.

[1513] 3. The device's emotion engine recognizes the user's emotions and determines that they are "feeling stressed."

[1514] 4. The server analyzes the data and identifies the passerby's replacement instructions.

[1515] 5. The device's camera captures images of the city, and the server recognizes passersby.

[1516] 6. The server selects character data that has the effect of reducing stress based on the emotional data.

[1517] 7. The server combines the character data with the location information of passersby to generate AR data.

[1518] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. Passersby on the way to work are transformed into characters that help reduce stress, making the commute more enjoyable for the user.

[1519] This system will enable users to enjoy a variety of entertainment experiences tailored to their emotions in their daily lives, further improving their quality of life.

[1520] The processing flow will be explained below.

[1521] Step 1:

[1522] The user issues voice commands to the AR glasses, such as "Turn people on the street into characters on my commute."

[1523] Step 2:

[1524] The device captures voice commands and converts the voice data into digital form, which is then analyzed by an internal speech recognition engine and converted into text data.

[1525] Step 3:

[1526] The terminal transmits the converted text data to the server, which includes the content of the voice instructions.

[1527] Step 4:

[1528] The server analyzes the received text data, using natural language processing technology to understand the meaning of the voice commands and extract the "target (people, room interior, etc.)" and "change content (character, Swiss lodge style, etc.)."

[1529] Step 5:

[1530] The device's emotion engine recognizes emotions from the user's voice and facial expressions, analyzing the tone of voice and facial muscle movements to identify emotions such as joy, sadness, anger, and surprise.

[1531] Step 6:

[1532] The device transmits the recognized emotion data to the server.

[1533] Step 7:

[1534] The server analyzes the voice commands and emotional data and then selects the character and interior design that best suits the user's emotions and instructions. For example, if the user is feeling sad, it will select an uplifting character.

[1535] Step 8:

[1536] The device's camera captures images of the real world from the user's point of view, and the image data is sent to a server in real time.

[1537] Step 9:

[1538] The server analyzes the received video data and uses image recognition technology to identify objects in the video (passersby, vehicles, furniture, etc.), and determines the position and posture of each object and stores them in a database.

[1539] Step 10:

[1540] The server uses the object's position and orientation information to generate AR data for displaying the selected character and interior in the real world. This data also includes the necessary rendering information.

[1541] Step 11:

[1542] The server sends the generated AR data to the device.

[1543] Step 12:

[1544] The device processes the AR data received in real time, and selected characters and interior designs are overlaid on the camera image.

[1545] Step 13:

[1546] The app also adjusts based on emotions. For example, if the user is feeling stressed, characters and interior design that will help alleviate stress will be selected and reflected in the display.

[1547] Step 14:

[1548] When the user observes the real world through the AR glasses, characters and interiors that change according to voice commands and emotions are superimposed on the real world, allowing the user to experience passersby on their commute or the interior of their room as if they were replaced with virtual characters and designs.

[1549] Example 2

[1550] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1551] Conventional augmented reality (AR) systems can replace real-world objects with virtual characters or interiors based on user instructions, but they lack the ability to adaptively change the display according to the user's emotional state. This makes it difficult to provide an experience that is in line with the user's emotions. It is also difficult to process in real time and realize a display that includes natural lighting and shadows.

[1552] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1553] In this invention, the server includes means for receiving voice instructions from a user, means for converting the received voice instructions into text data and analyzing the instructions, means for recognizing emotions from the user's voice and facial expressions, means for capturing camera images, identifying real-world objects, and acquiring position and orientation information, means for selecting optimal replacement data based on the voice instructions and the recognized emotions, means for generating the selected replacement data using the object's position and orientation information, and means for displaying the generated data overlaid on the camera image, thereby enabling adaptive AR display according to the user's emotional state in real time.

[1554] A "user" is a person who uses the system to issue voice commands and replace real-world objects with virtual characters and interiors.

[1555] "Voice instructions" are voice requests or commands given by the user to the system.

[1556] "Text data" is character information converted from voice instructions using voice recognition technology.

[1557] "Emotion recognition means" is a technology that analyzes the user's voice and facial expressions to identify the user's emotional state (joy, sadness, anger, etc.).

[1558] "Camera footage" refers to real-world video data captured in real time by a camera mounted on a device.

[1559] "Object" refers to a concrete object that exists in the real world (e.g., a passerby, a vehicle, furniture, etc.).

[1560] "Position and orientation information" refers to data on the physical position and orientation of an object identified in a camera image.

[1561] "Replacement data" is digital data used to replace an object with a virtual character or virtual interior.

[1562] "Rendering information" is data that includes the representation of lighting, shadows, and other elements required for AR display.

[1563] "Real-time processing" refers to the process of analyzing, generating, and displaying data instantly, without delay.

[1564] MODE FOR CARRYING OUT THE INVENTION

[1565] This invention is a system that uses augmented reality (AR) glasses worn by the user to display real-world objects replaced with virtual characters and interiors. This system incorporates an emotion engine that recognizes the user's emotions and includes a function to adaptively change the display according to the user's emotional state.

[1566] System Configuration

[1567] Hardware and software used

[1568] Device: AR glasses

[1569] Camera: A camera for capturing images of the real world.

[1570] Microphone: A microphone for capturing user voice commands.

[1571] Display: A display for overlaying virtual objects onto real-world images.

[1572] Server: Responsible for calculation processing

[1573] Natural language processing technology: For example, Google Cloud Natural Language API

[1574] Speech recognition software: for example, Google Cloud Speech-to-Text

[1575] Emotion recognition software: Examples include IBM Watson Tone Analyzer and Microsoft Azure Face API

[1576] Image recognition technology: For example, Google Cloud Vision API

[1577] AR data generation platform: Examples include Unity and Unreal Engine

[1578] Program processing

[1579] Acquiring and analyzing voice instructions

[1580] The user issues voice commands to the AR glasses, such as "Change the interior of my room to look like a Swiss lodge" or "Change the people on the street to characters on my commute." The device captures these commands and converts them into text using voice recognition software. The converted text is then sent to a server where it is analyzed using natural language processing technology.

[1581] emotion recognition

[1582] The device's built-in emotion engine recognizes emotions such as joy, sadness, anger, and surprise from the user's voice and facial expressions. Speech and facial recognition software analyzes this data to identify the user's emotional state.

[1583] Video Capture and Object Recognition

[1584] The device's camera captures images from the user's point of view in real time and transmits the image data to a server, which uses image recognition technology to identify real-world objects and determine their position and orientation.

[1585] Selecting and generating replacement data

[1586] The server selects the most suitable virtual character and interior design data based on the analyzed voice commands and emotion recognition data. This data is selected from a library stored in advance, and lighting and shadow information required for real-time rendering is also synchronized.

[1587] Viewing Data

[1588] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image. For example, if the user is feeling sad, an uplifting character or interior design will be displayed overlaid on the real world.

[1589] Specific examples

[1590] Example 1: Change the room interior to Swiss lodge style and adjust it according to the emotion.

[1591] 1. The user issues a voice command such as "Decorate the room in a Swiss lodge style."

[1592] 2. The device converts the voice into text data and sends it to the server.

[1593] 3. The device's emotion engine recognizes the user's emotion as "I want to relax."

[1594] 4. The server analyzes the data and identifies instructions for interior modifications.

[1595] 5. The device's camera captures images of the room, and the server recognizes the position of furniture and walls.

[1596] 6. The server selects interior design data that is more relaxing, such as a Swiss lodge.

[1597] 7. The server combines the interior data with the room's location information to generate AR data.

[1598] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. The user feels as if the interior of the room has been transformed into a Swiss lodge, creating an even more relaxing atmosphere.

[1599] Example 2: Transforming pedestrians into characters during your commute and adjusting their emotions

[1600] 1. The user issues a voice command such as, "Turn people on the street into characters while I'm commuting."

[1601] 2. The device converts the voice into text data and sends it to the server.

[1602] 3. The device's emotion engine recognizes the user's emotion as "feeling stressed."

[1603] 4. The server analyzes the data and identifies the passerby's replacement instructions.

[1604] 5. The device's camera captures images of the city, and the server recognizes passersby.

[1605] 6. The server selects character data that has the effect of reducing stress based on the emotional data.

[1606] 7. The server combines the character data with the location information of passersby to generate AR data.

[1607] 8. The server sends the generated AR data to the device, which then displays it overlaid on the camera image. Passersby on the user's commute are transformed into characters that help reduce stress, making the commute more enjoyable.

[1608] In this way, the system of the present invention provides a variety of entertainment experiences that correspond to the user's emotions in their daily lives, contributing to improving the quality of life.

[1609] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1610] Step 1:

[1611] The user issues voice commands to the AR glasses. For example, they might say, "Decorate the room in a Swiss lodge style." The device uses a microphone to capture the user's voice. This voice data becomes the input.

[1612] Step 2:

[1613] The device converts the captured voice data into text data using voice recognition software (e.g., voice recognition API). The input is voice data and the output is text data. The converted text data is sent to the server.

[1614] Step 3:

[1615] The server analyzes the received text data using natural language processing technology (e.g., natural language processing API). The input is text data, and the output is the analysis result. Here, the content of the voice instruction is understood.

[1616] Step 4:

[1617] The device uses an emotion engine to recognize emotions from the user's voice and facial expressions. The input is the user's voice data and camera footage, and the output is emotion data. Emotion recognition software (e.g., emotion analysis API) is used to identify emotions from voice tone and facial muscle movements.

[1618] Step 5:

[1619] The device's camera captures the image from the user's point of view in real time. The input is the real-world camera image, and the output is the captured image data. The image data is sent to the server where it is processed.

[1620] Step 6:

[1621] The server analyzes the captured video data using image recognition technology (e.g., image analysis API). The input is the video data, and the output is the position and posture information of the identified objects. This information is stored in a database.

[1622] Step 7:

[1623] The server selects the optimal replacement character and interior data based on the analyzed voice instructions and emotion recognition data. The input is voice instruction data and emotion data, and the output is the selected replacement data. It also obtains appropriate 3D models and animation data from the library.

[1624] Step 8:

[1625] The server generates AR data based on the object's position and orientation information to display the selected character and interior in real-world footage. The input is the object's position and the selected data, and the output is the generated AR data. Lighting and shadow information is also incorporated.

[1626] Step 9:

[1627] The server sends the generated AR data to the device. The device displays the received AR data overlaid on the camera image. The input is AR data, and the output is an AR display overlaid on the image of the real world. Through this, users can experience virtual characters and interiors as if they were in the real world.

[1628] The system utilizes generative AI models and prompts to provide a rich and adaptive entertainment experience that responds to the user's emotions, while also achieving natural AR display in real time, improving quality of life.

[1629] (Application example 2)

[1630] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1631] In recent years, advances in augmented reality (AR) technology have enabled users to combine the real world with virtual elements for enjoyment. However, current systems simply replace real-world objects with virtual characters or interiors, and are unable to adaptively change the display in response to the user's emotions. Therefore, there is a need for the development of systems that can take user emotions into account and provide greater satisfaction and relaxation.

[1632] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring voice instructions from the user, means for converting the acquired voice instructions into text data and analyzing the instruction content, means including an emotion engine for recognizing the user's emotions, means for capturing camera footage, identifying real-world objects, and acquiring position and orientation information, means for selecting replacement data based on the emotion data and the specified character or interior, means for generating the selected replacement data using the object's position and orientation information, and means for overlaying and displaying the generated data on the camera footage. This enables adaptive virtual replacement of real-world objects based on the user's emotions.

[1633] "Means for acquiring voice instructions from the user" refers to devices or software that detect commands spoken by the user to a wearable device such as AR glasses and incorporate them into the system.

[1634] "Means for converting acquired voice instructions into text data and analyzing the content of the instructions" refers to technology or devices that use voice recognition technology to convert voice data into character data, and then analyze the character data to understand the content of the instructions.

[1635] The "means including an emotion engine for recognizing the user's emotions" refers to a software or hardware system for recognizing the user's current emotional state in real time through analysis of voice and facial expressions.

[1636] "Means for capturing camera images, identifying real-world objects, and acquiring position and orientation information" refers to technology and devices that use a camera to capture real-world images, identify objects in the images using image recognition technology, and determine their position and orientation.

[1637] "Means for selecting replacement data based on emotional data and specified character or interior" refers to an algorithm or device for selecting appropriate character or interior data in accordance with the user's emotional state and voice instructions.

[1638] "Means for generating selected replacement data using object position and orientation information" refers to technology or devices that generate data for appropriately placing selected character and interior data in the real world based on the position and orientation information of identified objects.

[1639] "Means for displaying generated data by overlaying it on camera images" refers to technology or devices for displaying generated virtual content by overlaying it on real-world images captured by a camera.

[1640] System Overview

[1641] This system uses AR glasses worn by the user to replace real-world objects with virtual characters and interiors. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it also includes a function to adaptively change the display according to the user's emotional state.

[1642] Hardware and software used

[1643] Hardware: AR glasses (e.g. Microsoft HoloLens), smart devices with cameras, servers

[1644] Software: Speech recognition engine (e.g., Google Cloud Speech-to-Text), emotion recognition engine (e.g., Microsoft Azure Emotion API), image recognition technology (e.g., YOLO, TensorFlow)

[1645] Program Processing Overview

[1646] Obtaining voice commands from the user

[1647] A microphone installed on the device captures voice commands from the user. For example, if a user says, "Make my kitchen look like a Parisian cafe," that voice input is captured.

[1648] Analysis of voice instructions

[1649] The captured voice data is converted into text data by a voice recognition engine, and the converted text data is sent to a server where the instruction content is analyzed by an analysis engine.

[1650] Emotion recognition

[1651] The device's built-in emotion recognition engine analyzes the user's voice and facial expressions in real time to recognize their emotions. For example, if it recognizes that the user is in a relaxed mood, appropriate data will be selected.

[1652] Camera image capture and object recognition

[1653] The device's camera captures video within the user's field of view, and the video data is sent in real time to a server, which uses image recognition technology to identify real-world objects and determine their position and orientation.

[1654] Selection of replacement data

[1655] The server selects the optimal replacement data (characters and interior design) based on the analyzed voice commands and the recognized emotions. For example, it may select a Parisian cafe-style interior design that has a relaxing effect.

[1656] AR data generation

[1657] The server generates AR data based on the acquired object position and orientation information, including lighting and shadow information, to display the selected character and interior in the real world.

[1658] Viewing Data

[1659] The generated AR data is sent from the server to the device, which processes the received data in real time and displays it overlaid on the camera image, making it appear to the user that real-world objects have been replaced with the specified characters or interior.

[1660] Specific examples

[1661] Example 1: Remodeling your kitchen to look like a Parisian cafe

[1662] 1. The user issues a voice command such as, "Make my kitchen look like a Parisian cafe."

[1663] 2. The device converts the voice into text data and sends it to the server.

[1664] 3. The emotion engine determines that the user wants to relax.

[1665] 4. The server analyzes the data and identifies instructions for interior modifications.

[1666] 5. The device's camera captures images of the kitchen, and the server recognizes the location of the furniture.

[1667] 6. The server selects Parisian cafe-style interior data that has a relaxing effect.

[1668] 7. The server combines the interior data with the kitchen location information to generate AR data.

[1669] 8. The server sends the generated AR data to the device, which then overlays it on the camera image, allowing the user to enjoy a relaxing Parisian cafe-style kitchen.

[1670] Examples of prompt statements

[1671] You are a programmer tasked with designing a system for a food delivery AR glasses application that changes the kitchen interior to a virtual theme (e.g., Parisian cafe) and automatically adjusts to the user's emotions. Write a program procedure that meets the following requirements:

[1672] Requirements:

[1673] 1. The user issues a voice command (e.g., "Make my kitchen look like a Parisian cafe").

[1674] 2. Convert the voice instructions into text data and send it to the server.

[1675] 3. Recognize user emotions and automatically select appropriate themes.

[1676] 4. The AR glasses' camera captures images of the kitchen, and the server recognizes the position of the furniture.

[1677] 5. The server selects appropriate theme data (e.g., interior data of a Parisian cafe).

[1678] 6. The server generates the AR data and sends it to the device.

[1679] 7. AR data is overlaid on the camera image on the device.

[1680] Follow this procedure to design your program.

[1681] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1682] Step 1:

[1683] The user utters the voice command "Make my kitchen look like a Parisian cafe." The input is the user's voice, and the output is the captured voice data. The device's microphone captures this voice and stores it in its internal memory.

[1684] Step 2:

[1685] To convert the voice data into text data, the device uses the Google Cloud Speech-to-Text engine. The input is the captured voice data, and the output is the converted text data, which is sent from the device to the server for the next analysis step.

[1686] Step 3:

[1687] The server parses the received text data. Using a natural language processing engine (e.g., various NLP libraries), the input is the transformed text data and the output is the parsed results, which include instructions for the user to transform their kitchen into a Parisian cafe style.

[1688] Step 4:

[1689] The device's built-in emotion recognition engine recognizes the user's emotions in real time. The input is the user's voice tone and facial expression data, and the output is the user's emotional data. This emotional data is sent to the server and processed together with the analysis results.

[1690] Step 5:

[1691] The device's camera captures images of the kitchen and transmits the image data to the server in real time. The input is the captured camera image, and the output is the transmitted image data.

[1692] Step 6:

[1693] The server analyzes the transmitted video data using image recognition technology (e.g., YOLO, TensorFlow). The input is the camera's video data, and the output is the position and orientation information of objects in the kitchen. This allows the server to identify the positions of furniture and other objects in the kitchen.

[1694] Step 7:

[1695] The server selects the optimal replacement data (e.g., interior design data of a Parisian cafe) based on the analyzed voice instructions and the recognized emotion data. The input is the analysis results of the voice instructions and the emotion data, and the output is the selected replacement data.

[1696] Step 8:

[1697] The server uses the object's position and orientation information to generate AR data for displaying the selected interior data in the real world. The input is the object's position and orientation information and the selected replacement data, and the output is the generated AR data. This data includes, for example, 3D models and lighting information.

[1698] Step 9:

[1699] The generated AR data is sent from the server to the device. The input is the generated AR data, and the output is the data sent to the device. The device processes the received AR data in real time and displays it overlaid on the camera image. This makes the kitchen appear to have been transformed into a Parisian cafe in the user's field of vision.

[1700] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1701] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1702] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1703] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1704] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1705] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1706] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1707] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1708] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1709] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1710] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1711] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1712] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1713] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1714] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1715] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1716] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1717] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1718] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1719] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1720] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1721] The following is further disclosed regarding the above embodiment.

[1722] (Claim 1)

[1723] means for receiving voice instructions from a user;

[1724] means for converting the acquired voice instruction into text data and analyzing the instruction content;

[1725] A means for capturing camera images, identifying real-world objects, and acquiring their position and orientation information;

[1726] A means for selecting replacement data based on a designated character or interior;

[1727] a means for generating selected replacement data using the position and orientation information of the object;

[1728] a means for displaying the generated data by superimposing it on a camera image;

[1729] A system including:

[1730] (Claim 2)

[1731] 10. The system of claim 1, wherein the generated data is processed in real time and displayed overlaid on a video of the real world.

[1732] (Claim 3)

[1733] 10. The system of claim 1, wherein the system utilizes lens technology to integrate real-world images with the generated data.

[1734] "Example 1"

[1735] (Claim 1)

[1736] means for receiving voice instructions from a user;

[1737] means for converting the acquired voice instruction into text data and analyzing the instruction content;

[1738] A means for capturing camera images, identifying real-world objects, and acquiring their position and orientation information;

[1739] means for selecting replacement data based on the designated virtual character or interior;

[1740] a means for generating selected replacement data using the position and orientation information of the object;

[1741] a means for displaying the generated data by superimposing it on a camera image;

[1742] A system including:

[1743] (Claim 2)

[1744] 10. The system of claim 1, wherein the generated data is processed in real time and displayed overlaid on a video of the real world.

[1745] (Claim 3)

[1746] 10. The system of claim 1, wherein the system utilizes lens technology to integrate real-world images with the generated data.

[1747] "Application Example 1"

[1748] (Claim 1)

[1749] means for receiving voice instructions from a user;

[1750] means for converting the acquired voice instruction into text data and analyzing the instruction content;

[1751] A means for capturing camera images, identifying real-world objects, and acquiring their position and orientation information;

[1752] A means for selecting replacement data based on a designated character or interior;

[1753] a means for generating selected replacement data using the position and orientation information of the object;

[1754] a means for displaying the generated data by superimposing it on a camera image;

[1755] A means to generate and change AR data for store interiors and products according to the season or event,

[1756] A system including:

[1757] (Claim 2)

[1758] 10. The system of claim 1, wherein the generated data is processed in real time and displayed overlaid on a video of the real world.

[1759] (Claim 3)

[1760] 10. The system of claim 1, wherein the system utilizes lens technology to integrate real-world images with the generated data.

[1761] "Example 2: Combining Emotion Engines"

[1762] (Claim 1)

[1763] means for receiving voice instructions from a user;

[1764] means for converting the acquired voice instruction into text data and analyzing the instruction content;

[1765] emotion recognition means for recognizing emotions from the user's voice and facial expressions;

[1766] A means for capturing camera images, identifying real-world objects, and acquiring their position and orientation information;

[1767] means for selecting optimal replacement data based on the voice instructions and the recognized emotion;

[1768] a means for generating selected replacement data using the position and orientation information of the object;

[1769] a means for displaying the generated data by superimposing it on a camera image;

[1770] A system including:

[1771] (Claim 2)

[1772] 10. The system of claim 1, wherein the generated data is processed in real time and displayed overlaid on a video of the real world.

[1773] (Claim 3)

[1774] 10. The system of claim 1, wherein the data to be displayed includes rendering information including lighting and shadows.

[1775] "Application example 2 when combining emotion engines"

[1776] (Claim 1)

[1777] means for receiving voice instructions from a user;

[1778] means for converting the acquired voice instruction into text data and analyzing the instruction content;

[1779] means including an emotion engine for recognizing an emotion of a user;

[1780] A means for capturing camera images, identifying real-world objects, and acquiring their position and orientation information;

[1781] means for selecting replacement data based on the emotion data and the designated character or interior;

[1782] a means for generating selected replacement data using the position and orientation information of the object;

[1783] a means for displaying the generated data by superimposing it on a camera image;

[1784] A system including:

[1785] (Claim 2)

[1786] 10. The system of claim 1, wherein the generated data is processed in real time and displayed overlaid on a video of the real world.

[1787] (Claim 3)

[1788] 10. The system of claim 1, wherein the system utilizes lens technology to integrate real-world images with the generated data. [Explanation of symbols]

[1789] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving voice instructions from a user; means for converting the acquired voice instruction into text data and analyzing the instruction content; A means for capturing camera images, identifying real-world objects, and acquiring their position and orientation information; A means for selecting replacement data based on a designated character or interior; means for generating selected replacement data using the object's position and orientation information; a means for displaying the generated data by superimposing it on the camera image; A system including:

2. The system according to claim 1, wherein the generated data is processed in real time and displayed overlaid on a video of the real world.

3. 10. The system of claim 1, wherein lens technology is utilized to integrate real-world images with generated data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A