system
The system addresses the lack of virtual stimulation in daily life by capturing and augmenting real-world environments with user-customizable virtual characters, enhancing user experience through real-time interaction and emotional responsiveness.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
Existing technologies fail to seamlessly integrate virtual experiences into the real world, providing limited stimulation and customization for users in mundane environments, especially during daily activities like commuting or spending time indoors.
A system that captures real-world environments, recognizes objects, replaces them with user-selected virtual characters, and overlays them in real-time using augmented reality, allowing users to customize and interact with these characters through an interface, dynamically adjusting to user preferences and emotions.
Enhances user experience by providing an individually tailored, interactive, and emotionally resonant virtual augmentation of daily life, enriching the visual and emotional engagement in everyday scenarios.
Smart Images

Figure 2026069148000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In a real environment, daily life is likely to feel monotonous and boring, and there is a need for a method to increase the interests and pleasures of users. In particular, when commuting, during daily movements, or when spending a long time indoors, there are few changes in scenery and environment, making it difficult to provide stimulation to users. In view of such a situation, it is necessary to develop a technology that seamlessly integrates a virtual non-daily experience according to the preferences of users into the real world.
Means for Solving the Problems
[0005] This invention includes an imaging means for capturing a real-world environment and a recognition means for analyzing images acquired by the imaging means to detect objects. It also includes a conversion means for replacing real-world objects with virtual characters pre-set by the user based on the object information detected by the recognition means. Furthermore, it uses an output means to overlay and display the virtual characters generated by the conversion means onto the real-world environment. This allows users to enjoy an individually customized, extraordinary visual experience in their daily lives. In addition, by including an interface means that allows users to select and change virtual characters and background themes based on user input, and a tracking means that estimates the pose of images acquired from the imaging means and dynamically adjusts the display of the virtual characters, more advanced customization and real-time display adjustments become possible, increasing the sense of realism.
[0006] "Imaging means" refers to devices and technologies for capturing the real environment as image data.
[0007] "Recognition means" refers to technologies for detecting objects from acquired image data and analyzing their type and location.
[0008] "Object information" refers to data about an object detected by a recognition means, and includes attribute information such as position, size, and orientation.
[0009] "Conversion means" refers to technology for replacing real-world objects with virtual characters set by the user based on object information.
[0010] A "virtual character" refers to a digital object selected by the user, which is displayed in place of a real-world object.
[0011] "Output means" refers to the technology for displaying the virtual character generated by the conversion means overlaid on the real environment.
[0012] "Interface means" refers to the operation screen or method that allows the user to select or change settings for the system.
[0013] "Tracking means" refers to a technology that tracks the movement of a virtual character in real time based on the object position and posture information of acquired images, and adjusts the display accordingly. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] A sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] The present invention is a system for virtually extending information about the real environment, and is specifically implemented as follows.
[0036] System Configuration
[0037] First, the user's device is equipped with a camera as a means of capturing images. The device captures images of the real world within the user's field of view in real time. At this time, some kind of compression or pre-processing is performed so that the device can process the video data seamlessly.
[0038] Object recognition and data transmission
[0039] The terminal is equipped with recognition capabilities to detect objects from captured video footage, using machine learning algorithms to identify the type and location of objects. This information is compiled as object information and transmitted to a server via the network.
[0040] Character replacement process
[0041] The server performs a conversion process based on the received object information to replace it with a virtual character selected by the user. Here, appropriate character information is selected based on the user's pre-configured preferences and themes. Necessary movement and posture information is added to the converted character data.
[0042] AR display and interaction
[0043] Character data transmitted from the server is returned to the terminal and displayed in the user's field of view, superimposed on the real world, via an output device. This allows the user to visually experience an extraordinary scene where a virtual character adorns everyday landscapes. The user can freely change the character and background theme used using the terminal's interface.
[0044] Example: Usage scenario during commuting
[0045] When a user captures their surroundings through their device's camera during their commute, the device recognizes pedestrians and vehicles and sends that information to a server. The server then replaces these with anime or fantasy characters chosen by the user, which are then displayed in their field of view. Users can customize the characters' costumes and movements according to their settings, allowing them to enjoy their commute with a different theme each day.
[0046] In this way, the present invention combines reality and virtuality to provide users with new experiences.
[0047] The following describes the processing flow.
[0048] Step 1:
[0049] The device uses its built-in camera to capture the real-world environment within the user's field of view in real time. The captured video data undergoes pre-processing such as noise reduction and contrast adjustment.
[0050] Step 2:
[0051] The device analyzes pre-processed video data and uses an object recognition algorithm to detect people, vehicles, furniture, and other objects within the video. The detection results include the type, location, and size of the objects.
[0052] Step 3:
[0053] The terminal transmits object information obtained as a result of object recognition to the server. This information is structured as a dataset including object type and location information, and is transmitted rapidly over the network.
[0054] Step 4:
[0055] Based on the received object information, the server identifies the virtual character selected by the user. The server then decides which character to replace the selected character with, according to the user's settings, and prepares the necessary data for that replacement.
[0056] Step 5:
[0057] The server transforms the prepared virtual character data to correspond to the object's position and orientation. This transformation includes parameters to adjust the character's movement and orientation.
[0058] Step 6:
[0059] The server sends the converted virtual character data to the terminal. The transmitted information includes the virtual character's model and all the data necessary for display.
[0060] Step 7:
[0061] The device uses the received virtual character data to overlay the virtual character onto the real-world image. This display occurs in real time within the user's field of view.
[0062] Step 8:
[0063] Users can change and customize the displayed virtual character and background theme through the device's interface. Interactions are reflected immediately in response to user actions.
[0064] (Example 1)
[0065] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0066] When providing users with a new visual experience by virtually augmenting and displaying information from the real world, it is necessary to seamlessly and in real time merge the environment and virtual elements without requiring advanced processing power. Furthermore, flexibility is required, allowing users to freely customize virtual elements and background settings according to their individual preferences.
[0067] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0068] In this invention, the server includes a means for capturing environmental information, a means for recognizing objects by analyzing the images obtained from the capturing means, and a means for converting the identified object information into virtual elements. This makes it possible to merge images of the real world and virtual elements in real time and provide a customized visual experience based on the user's preferences.
[0069] "Mechanism of filming" refers to devices or functions that acquire information about the real world environment as images.
[0070] "Recognition means" refers to algorithms and processes for analyzing video data obtained from a shooting means and identifying specific objects.
[0071] "Conversion means" refers to the processes and functions used to replace recognized object information with virtual elements.
[0072] "Display means" refers to a means of presenting virtual elements to the user by overlaying them with images of the real environment.
[0073] "Transmission means" refers to the function of sending and receiving object information between servers and other devices via communication.
[0074] "Operational means" refers to interface functions that allow users to select or change virtual elements and background settings.
[0075] "Tracking means" refers to a function that recognizes the position of the image obtained from the shooting means and dynamically adjusts the display of virtual elements.
[0076] This invention is a system for virtually extending real-world environmental information and enriching the user's visual experience. This system consists of three components: a terminal, a server, and a user.
[0077] The device is equipped with a camera and other means of capturing images, which acquires real-time video footage of the user's surroundings. To efficiently process the video data, the device performs pre-processing such as compression and filtering. A smart device with a high-performance processor is suitable for this process.
[0078] As a means of recognition, the device is equipped with software capable of executing machine learning algorithms to identify objects from captured video footage. Lightweight models such as TENSORFLOW® Lite are often used in this process.
[0079] The server receives object information transmitted from the terminal via communication and uses a generative AI model as a means of conversion to replace it with a virtual element. Here, the generative AI model selects an appropriate virtual character based on themes and preferences pre-set by the user and adjusts its movements and posture. The generated virtual element is processed using cloud server resources as needed.
[0080] The completed character is returned to the device and displayed as a composite image with real-world footage. AR technology is used in this process, seamlessly overlaying the character onto the user's field of view, providing an integrated image where real and virtual elements merge. Users can select and customize virtual elements through various controls, enabling a personalized and interactive experience.
[0081] As a concrete example, consider a scenario where a user takes photos of their surroundings using their device's camera while commuting. The device recognizes pedestrians and vehicles and sends that information to a server. The server replaces these with fantasy characters and animations, allowing the user to enjoy a different theme each day according to their settings.
[0082] An example of a prompt would be, "Add fantasy characters to the scenery you see during your commute." This prompt prompts the generative AI model to prepare to provide relevant content.
[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0084] Step 1:
[0085] The device uses a camera to acquire real-time video of the environment. It receives real-world video light data as input and generates compressed video frames as output. Specifically, a high-performance processor compresses the video data and performs filtering to remove noise.
[0086] Step 2:
[0087] The device executes a machine learning algorithm to recognize objects in compressed video frames. It takes compressed video frames as input and generates a dataset containing object types and location information as output. Specifically, a model such as TensorFlow Lite analyzes the data and extracts feature vectors for the identified objects.
[0088] Step 3:
[0089] The device transmits information about recognized objects to the server via communication. It receives a dataset containing object type and location information as input and generates formatted data for transfer to the server as output. The data is structured in JSON format and transmitted rapidly over the network.
[0090] Step 4:
[0091] Based on the object information received by the server, it uses a generative AI model to convert it into virtual elements. It receives the object information as input and generates a dataset of virtual elements as output. The server then runs the generative AI model, selects a virtual character according to the user's settings, and adds the necessary behavioral information.
[0092] Step 5:
[0093] The server generates virtual element data and sends it back to the terminal. It receives a dataset of virtual elements as input and generates optimized data sent to the terminal as output. The data contains information necessary for AR display and is efficiently transferred via high-speed streaming.
[0094] Step 6:
[0095] The device receives virtual elements, combines them with real-world video, and displays them to the user. It receives virtual element data and real-time video as input and generates a composite image presented to the user as output. In practice, AR technology is used to seamlessly overlay virtual characters onto the real-world environment.
[0096] Step 7:
[0097] The user selects and modifies virtual elements and backgrounds through the interface. It receives user settings as input and generates data reflecting the updated virtual element settings as output. Using various controls, users can intuitively customize the system through touch operations and other means.
[0098] (Application Example 1)
[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0100] There is a need to expand the realm of information presentation to users through the identification of objects in the real world and the combination of these objects with virtual display elements. However, existing systems have difficulty providing a consistent experience that fuses real and virtual information, and there is a lack of efficient methods for presenting information that aligns with the user's intentions, especially when multiple choices exist. This invention aims to solve these problems.
[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0102] In this invention, the server includes a video acquisition means for capturing the real environment, an analysis means for analyzing visual information obtained from the video acquisition means and recognizing objects, and an ambiguity removal means for identifying products when a user acquires video. This makes it possible to appropriately identify objects in the real world and efficiently provide virtual display elements according to the user's preferences.
[0103] "Image acquisition means" refers to devices or software that have the function of capturing the real environment.
[0104] An "analysis means" is an algorithm or system that receives visual information obtained from an image acquisition means, analyzes it, and recognizes objects.
[0105] "Conversion means" refers to a process or apparatus for converting entity information identified by analysis means into virtual display elements selected by the user.
[0106] "Output means" refers to devices or display technologies that present information to users by overlaying virtual display elements generated by conversion means onto the real environment.
[0107] "Ambiguousness elimination methods" are algorithms and techniques used by users to clearly identify identifiable products and objects when acquiring video footage.
[0108] The "presentation adjustment means" is a function for controlling and displaying story information corresponding to products identified using the ambiguity exclusion means, in accordance with the user's requests and context.
[0109] A "tracking mechanism" is a system that estimates the angle and position of visual information obtained from a video acquisition mechanism and adjusts the appearance of virtual display elements in real time.
[0110] The system for realizing this invention is implemented through a combination of many devices and software. The terminal used by the user to acquire video footage utilizes a camera and smart glasses to capture the real environment. The resulting video data is processed in real time within the terminal, undergoes necessary preprocessing, and is then sent to a server. The server receives the video data and uses machine learning algorithms (e.g., TensorFlow or PyTorch) to identify objects. The recognized object information is cross-referenced with information in a database and converted into a virtual character according to the request.
[0111] The converted character information is returned to the user's device and displayed overlaid on the real world using AR technology (e.g., ARKit / ARCore). Users can customize virtual display elements and background themes through the device's interface. For example, in a shopping mall, if a user picks up a piece of clothing and scans it with their camera, a virtual character can appear on the screen and try on the clothing. To enhance this user experience, a generative AI model is used with prompts such as, "In the next scene, the virtual character should try on the apparel item selected by the customer and visually represent the features of the item." These prompts function as input to enable more precise information presentation and support the object recognition and information generation processes.
[0112] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0113] Step 1:
[0114] The user acquires video using their device. The device's camera captures the real environment, acquiring raw video data. This data is compressed for efficient processing and undergoes pre-processing such as color correction and resolution adjustment. The input is the captured raw video data, and the output is the processed video data.
[0115] Step 2:
[0116] The device analyzes the processed video data and recognizes objects. It uses machine learning algorithms to analyze the contours and features of objects in the video and identify them. In this step, a trained model (e.g., TensorFlow) is used to assign an object category to each pixel. The input is the processed video data from the previous step, and the output is a list of recognized objects and their locations.
[0117] Step 3:
[0118] The terminal sends information about recognized objects to the server. The server receives this information, consults its database, and prepares to convert it into a virtual display element set by the user. The input is a list of objects and their locations, and the output is candidate information for the virtual character needed for the conversion.
[0119] Step 4:
[0120] The server selects an appropriate virtual character based on candidate virtual character information, according to the user's chosen theme and preferences. The selected character is then configured with the necessary movement patterns and postures. The input is candidate virtual character information, and the output is a set of specific virtual characters.
[0121] Step 5:
[0122] The server sends character data to the device, which receives it. The device then uses AR technology (e.g., ARKit / ARCore) to overlay the virtual character onto the real environment. This allows the user to enjoy interacting with the virtual character. The input is a specific set of virtual characters, and the output is a composite image displayed in the user's field of view.
[0123] Step 6:
[0124] Users can change the displayed virtual elements and background themes through the device interface. This allows for a customized experience tailored to the user's preferences. Input is user action information, and output is updated theme and virtual character information.
[0125] Step 7:
[0126] The server uses a generative AI model to generate story information for the virtual character. This process uses prompts to add more detailed elements. An example of a prompt in this step is, "In the next scene, the virtual character should try on the apparel item selected by the customer and visually represent the features of that item." The input is the updated virtual character information, and the output is the added story information.
[0127] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0128] The system of the present invention virtually augments information from the real environment and provides an interactive experience that responds to the user's emotions. Specific embodiments are described below.
[0129] System Configuration
[0130] This system uses a camera-equipped terminal to capture the real environment and incorporates recognition means for analyzing that data. This allows it to detect objects (people, vehicles, furniture, etc.) within the user's field of view and collect that information.
[0131] Introducing an emotional engine
[0132] The system uses an emotion engine that infers emotions by analyzing the user's facial expressions and voice. It evaluates data collected by the device's built-in microphone and camera to identify emotional states such as joy, anger, sadness, and happiness. Based on this, a virtual character or theme that matches the user's emotions is selected.
[0133] Character generation and tracking
[0134] The server generates a selected virtual character based on object information and user emotion data. Here, the character's movements and facial expressions are adjusted to match the user's emotions. The terminal uses tracking mechanisms to ensure the character is always seamlessly integrated with the real environment.
[0135] User interaction
[0136] Users can select and change the displayed characters and background themes through the interface. If the user's emotions change, the emotion engine detects this and dynamically updates the characters and background themes.
[0137] Learning and customization
[0138] The system learns the user's past emotional patterns and selection history, and based on that, provides a more comfortable and personalized experience. This feature allows the system to learn the user's preferences, making the suggested virtual characters and themes more personally suited to them.
[0139] Example: Personalized home entertainment
[0140] For example, when a user is relaxing at home, if the emotion engine detects the user's calm emotions, the system will select a character with a calm and relaxed theme and adjust the room's interior background accordingly. This allows the user to experience a virtual environment that makes their relaxation time even more pleasant.
[0141] These features enable the present invention to enhance the user's daily experience, making it more personalized and emotionally resonant.
[0142] The following describes the processing flow.
[0143] Step 1:
[0144] The device uses its built-in camera to capture the real-world environment surrounding the user in real time. The camera footage undergoes initial processing and is converted into a format that is easy to analyze.
[0145] Step 2:
[0146] The device performs object recognition on the captured data, identifying people and objects within the video. The recognized object information (location, type, size, etc.) is then organized for processing.
[0147] Step 3:
[0148] The device's emotion engine uses the camera and microphone to analyze the user's emotions from their face and voice. It identifies the emotional state (joy, anger, sadness, etc.) and outputs this information.
[0149] Step 4:
[0150] The device sends data about objects and emotional states to the server. This ensures that the information is collected in a secure manner.
[0151] Step 5:
[0152] The server selects a virtual character that suits the user based on the received object information and emotional state. Past usage history and emotional patterns are also taken into consideration during the selection process.
[0153] Step 6:
[0154] The server adjusts the selected virtual character's movements and facial expressions to match its emotional state. This adjusted character information is then sent to the terminal.
[0155] Step 7:
[0156] The device uses the received virtual character information to overlay the character onto the real-world image. This display is precisely integrated into the user's field of view through the device's tracking mechanism.
[0157] Step 8:
[0158] Users can manually change character types and background themes through the interface. Furthermore, the display automatically updates when changes in the user's emotions are detected.
[0159] Step 9:
[0160] The server learns from user choices and sentiment data, updating its algorithms to provide a more optimized experience for subsequent uses. This learning enhances the system's personalization capabilities.
[0161] (Example 2)
[0162] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0163] There is a growing demand for seamlessly blending real and virtual environments to provide interactive experiences that respond to user emotions. However, conventional technologies struggle to adapt to dynamic changes in the real environment and evolving user emotions, resulting in limited experiences. This invention aims to address this issue by flexibly responding to changes in user emotions and the environment, thereby providing a more personalized experience.
[0164] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0165] In this invention, the server includes a device for recording the real environment, means for identifying objects using a processing device, means for integrating virtual entities into the real environment using a generation device, and means for reflecting the characteristics of the virtual entities using an emotion analysis device. This makes it possible to provide a seamless and personalized interactive experience that responds to the user's emotions and changes in the real environment.
[0166] A "device for recording the real environment" is a hardware device that uses cameras, microphones, etc., to acquire information about the physical surroundings as digital data.
[0167] "Means for identifying objects using a processing device" refers to algorithms and hardware configurations for analyzing acquired digital data and identifying specific objects or elements.
[0168] "Means for integrating virtual entities into the real environment using a generation device" refers to technology for generating virtual entities and displaying them by overlaying them onto real-world video and audio information.
[0169] "A means of using an emotion analysis device to reflect the characteristics of a virtual entity" refers to a system that analyzes the user's emotional state and uses the results to influence the behavior and appearance of a virtual entity.
[0170] Modes for carrying out the invention
[0171] The system of the present invention integrates multiple devices and technologies to provide users with an interactive and personalized experience. The system acquires and analyzes information from the real environment, understands the user's emotions, and generates and displays virtual entities accordingly. Specific embodiments are described below.
[0172] Hardware and software configuration
[0173] 1. Recording the real environment
[0174] The device uses a camera and microphone to record the user's surroundings and audio in real time. This video and audio information is analyzed using libraries such as OpenCV.
[0175] 2. Object recognition
[0176] The device identifies people and objects based on the recorded data. This process utilizes machine learning techniques, modeled using TensorFlow and Keras.
[0177] 3. Analysis of emotions
[0178] The server uses deep learning algorithms to infer emotions from the user's facial expressions and voice. Emotion analysis is performed by an emotion engine built on the user's historical data.
[0179] 4. Creating and displaying virtual entities
[0180] The server generates virtual entities using the Unity engine based on the collected information. The device integrates these entities into the real environment using AR technology and displays them to the user.
[0181] 5. Individualization and Learning
[0182] The system uses tools like Scikit-learn to learn from the user's past emotions and choices, and then suggests suitable virtual entities and environments for the next time.
[0183] Specific example
[0184] For example, when a user wants to relax, the emotion analysis engine detects the user's calm emotions. In this case, the system selects a virtual entity with a calm and relaxing theme and adjusts the room's interior to create a virtually relaxing atmosphere. This allows the user to enjoy their relaxation time more comfortably.
[0185] Example of a prompt
[0186] "Please suggest a virtual character suitable for when the user is relaxing. Characteristics: calm, peaceful, surrounded by nature."
[0187] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0188] Step 1:
[0189] The device activates its camera and microphone to record the user's surroundings and audio. It uses real-world video and audio as input and outputs video and audio in digital data format. Specifically, the device captures video and audio at regular intervals and saves them as a data stream.
[0190] Step 2:
[0191] The device analyzes the acquired video data to detect objects. The input is the video data obtained in step 1, and an object recognition algorithm (for example, using OpenCV) is applied to generate object information, including the position and type of each object, as output. Specifically, the device processes the video frames sequentially to identify people, furniture, and other objects.
[0192] Step 3:
[0193] The device infers the user's emotions based on the acquired audio and facial video. The input is the audio and facial video data obtained in step 1, and an emotion analysis model (for example, a deep learning model using TensorFlow) is used to identify the user's emotional state as output. Specifically, the device profiles the tone of the voice and facial feature points.
[0194] Step 4:
[0195] The server generates virtual entities using object information and emotion data. The input is the data obtained in steps 2 and 3, and the server uses the Unity engine to design the virtual entities, generating 3D models and animation data as output. Specifically, the server selects an entity that matches the user's current emotion and performs individual customization.
[0196] Step 5:
[0197] The device integrates and displays the generated virtual entities in the real environment. The 3D model and animation data obtained in step 4 are used as input, and the virtual entities are rendered within the user's visual field via AR technology as output. Specifically, the device continuously updates its display according to location information and the user's movement.
[0198] Step 6:
[0199] Users interact with virtual entities and backgrounds through a provided interface. The system incorporates user selections and actions as input, dynamically updating entities and backgrounds based on this information and adjusting the display as feedback. Specifically, the interface instantly reflects the user's actions on the screen.
[0200] Step 7:
[0201] The server learns from past user sentiment data and selection history to provide appropriate suggestions for the next time. The input is historical data, and machine learning algorithms are used to generate personalized virtual entities and theme suggestions as output. Specifically, the system periodically analyzes the data and optimizes content to suit the user.
[0202] (Application Example 2)
[0203] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0204] Providing flexible and personalized customer service that responds to the diverse emotions and needs of visitors in physical stores is difficult with traditional methods. Furthermore, responding quickly to changes in visitors' emotions and suggesting the most suitable services and products is a challenge.
[0205] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0206] In this invention, the server includes an image acquisition means for acquiring the real environment, an emotion analysis means for analyzing the emotional state of visitors, and a service proposal means for selecting a customer service style based on the emotion analysis means. This enables immediate and appropriate service proposals that respond to the emotions and needs of visitors.
[0207] "Image acquisition means" refers to a device or technology for collecting visual information from the real environment.
[0208] A "judgment tool" is a function that analyzes acquired visual information and identifies a specific object.
[0209] A "conversion method" is a process that converts identified target information into a virtual target set by the user.
[0210] "Output means" refers to a mechanism that overlays a virtual object onto the real environment and performs integrated visual display.
[0211] "Emotional analysis techniques" are technologies used to analyze and infer the emotional state of visitors.
[0212] A "service proposal method" is a method for presenting the optimal customer service style and services based on information obtained through emotion analysis.
[0213] A "tracking method" is a technology for estimating the location of visual information and dynamically adjusting the display of a virtual object.
[0214] The present invention describes embodiments for carrying out the invention. The present invention realizes a system that provides personalized customer service in a physical store that responds to the emotions of the visitors.
[0215] The server uses an image acquisition device to capture the real environment. This image acquisition device is a general-purpose camera device that captures visitors' facial expressions and actions. The captured image data is analyzed by decision software to identify specific objects and features. The decision software includes image processing algorithms using OpenCV and analysis models using TensorFlow.
[0216] Furthermore, the server runs emotion analysis software to infer the visitor's emotions from the collected image and audio data. Audio data acquired by the microphone is also used in the emotion analysis. This makes it possible to identify the visitor's emotional state, such as whether they are relaxed or in a hurry.
[0217] Users can observe and utilize dynamically integrated virtual objects through smart devices. For example, a virtual character superimposed onto the real environment appears through smart glasses, supporting comfortable customer service. The display position and movement of the virtual character are dynamically adjusted by tracking software, enabling natural interaction.
[0218] As a concrete example, when a visitor in a music store wants to relax, the system analyzes the visitor's facial expressions and tone of voice to suggest a demo of relaxing music and then provides the most suitable music and service. The prompt, "Please suggest music and in-store presentation methods that are suitable for a customer who wants to relax," is input into the AI model, enabling the system to provide optimal content suggestions.
[0219] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0220] Step 1:
[0221] The server uses a camera device to capture the visitor's facial expressions and movements. This provides real-world image data as input. The server receives this image data, performs necessary preprocessing to reduce noise, and converts it into a format suitable for analysis.
[0222] Step 2:
[0223] The server inputs the pre-processed image data into the decision-making software. Using a machine learning model powered by TensorFlow, it identifies important objects and features within the image. This data processing outputs the location and attribute information of the objects (e.g., emotion estimation data from face recognition).
[0224] Step 3:
[0225] The server analyzes the identified facial recognition data using emotion analysis software. Input includes facial feature data and audio data. The software performs emotional calculations and provides output that identifies the visitor's emotional state (e.g., happy, calm, anxious).
[0226] Step 4:
[0227] The terminal generates a virtual display based on output data to overlay a virtual character onto the smart device. Graphic rendering technology is used for data processing, dynamically generating a character that expresses the identified emotion. This enables interaction tailored to the visitor's mood.
[0228] Step 5:
[0229] The user inputs a prompt into the generating AI model. Specifically, the prompt used is, "Please suggest music and store presentation methods that would be suitable when a customer wants to relax." Based on the input prompt, the generating AI model outputs the most suitable suggestions.
[0230] Step 6:
[0231] The terminal displays information to visitors and staff on their smart devices based on the outputted suggestions. These suggestions include specific customer service styles and product descriptions, providing visual and auditory feedback to enhance the user experience.
[0232] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0233] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0234] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0235] [Second Embodiment]
[0236] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0237] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0238] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0239] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0240] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0241] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0242] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0243] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0244] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0245] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0246] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0247] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0248] The present invention is a system for virtually extending information about the real environment, and is specifically implemented as follows.
[0249] System Configuration
[0250] First, the user's device is equipped with a camera as a means of capturing images. The device captures images of the real world within the user's field of view in real time. At this time, some kind of compression or pre-processing is performed so that the device can process the video data seamlessly.
[0251] Object recognition and data transmission
[0252] The terminal is equipped with recognition capabilities to detect objects from captured video footage, using machine learning algorithms to identify the type and location of objects. This information is compiled as object information and transmitted to a server via the network.
[0253] Character replacement process
[0254] The server performs a conversion process based on the received object information to replace it with a virtual character selected by the user. Here, appropriate character information is selected based on the user's pre-configured preferences and themes. Necessary movement and posture information is added to the converted character data.
[0255] AR display and interaction
[0256] Character data transmitted from the server is returned to the terminal and displayed in the user's field of view, superimposed on the real world, via an output device. This allows the user to visually experience an extraordinary scene where a virtual character adorns everyday landscapes. The user can freely change the character and background theme used using the terminal's interface.
[0257] Example: Usage scenario during commuting
[0258] When a user captures their surroundings through their device's camera during their commute, the device recognizes pedestrians and vehicles and sends that information to a server. The server then replaces these with anime or fantasy characters chosen by the user, which are then displayed in their field of view. Users can customize the characters' costumes and movements according to their settings, allowing them to enjoy their commute with a different theme each day.
[0259] In this way, the present invention combines reality and virtuality to provide users with new experiences.
[0260] The following describes the processing flow.
[0261] Step 1:
[0262] The device uses its built-in camera to capture the real-world environment within the user's field of view in real time. The captured video data undergoes pre-processing such as noise reduction and contrast adjustment.
[0263] Step 2:
[0264] The device analyzes pre-processed video data and uses an object recognition algorithm to detect people, vehicles, furniture, and other objects within the video. The detection results include the type, location, and size of the objects.
[0265] Step 3:
[0266] The terminal transmits object information obtained as a result of object recognition to the server. This information is structured as a dataset including object type and location information, and is transmitted rapidly over the network.
[0267] Step 4:
[0268] Based on the received object information, the server identifies the virtual character selected by the user. The server then decides which character to replace the selected character with, according to the user's settings, and prepares the necessary data for that replacement.
[0269] Step 5:
[0270] The server transforms the prepared virtual character data to correspond to the object's position and orientation. This transformation includes parameters to adjust the character's movement and orientation.
[0271] Step 6:
[0272] The server sends the converted virtual character data to the terminal. The transmitted information includes the virtual character's model and all the data necessary for display.
[0273] Step 7:
[0274] The device uses the received virtual character data to overlay the virtual character onto the real-world image. This display occurs in real time within the user's field of view.
[0275] Step 8:
[0276] Users can change and customize the displayed virtual character and background theme through the device's interface. Interactions are reflected immediately in response to user actions.
[0277] (Example 1)
[0278] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0279] When providing users with a new visual experience by virtually augmenting and displaying information from the real world, it is necessary to seamlessly and in real time merge the environment and virtual elements without requiring advanced processing power. Furthermore, flexibility is required, allowing users to freely customize virtual elements and background settings according to their individual preferences.
[0280] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0281] In this invention, the server includes a photographing means for acquiring environmental information, a recognition means for analyzing the video obtained from the photographing means to identify an object, and a conversion means for replacing the identified object information with virtual elements. Thereby, it becomes possible to fuse the video of the real world and virtual elements in real time and provide a customized visual experience based on the user's preferences.
[0282] The "photographing means" refers to a device or function for acquiring environmental information of the real world as a video.
[0283] The "recognition means" refers to an algorithm or process for analyzing the video data obtained from the photographing means and identifying a specific object.
[0284] The "conversion means" refers to a process or function for replacing the recognized object information with virtual elements.
[0285] The "display means" refers to a means for presenting virtual elements by superimposing them on the video of the real environment to the user.
[0286] The "transmission means" refers to a function for transmitting and receiving object information between the server and other devices via communication.
[0287] The "operation means" refers to an interface function for the user to select or change virtual elements and background settings.
[0288] The "tracking means" refers to a function for recognizing the position of the video obtained from the photographing means and dynamically adjusting the display of virtual elements.
[0289] The present invention is a system for virtually expanding environmental information of the real world and enriching the user's visual experience. This system is composed of three entities: a terminal, a server, and a user.
[0290] The device is equipped with a camera and other means of capturing images, which acquires real-time video footage of the user's surroundings. To efficiently process the video data, the device performs pre-processing such as compression and filtering. A smart device with a high-performance processor is suitable for this process.
[0291] As a means of recognition, the device is equipped with software capable of executing machine learning algorithms to identify objects from captured video footage. Lightweight models such as TensorFlow Lite are often used in this process.
[0292] The server receives object information transmitted from the terminal via communication and uses a generative AI model as a means of conversion to replace it with a virtual element. Here, the generative AI model selects an appropriate virtual character based on themes and preferences pre-set by the user and adjusts its movements and posture. The generated virtual element is processed using cloud server resources as needed.
[0293] The completed character is returned to the device and displayed as a composite image with real-world footage. AR technology is used in this process, seamlessly overlaying the character onto the user's field of view, providing an integrated image where real and virtual elements merge. Users can select and customize virtual elements through various controls, enabling a personalized and interactive experience.
[0294] As a concrete example, consider a scenario where a user takes photos of their surroundings using their device's camera while commuting. The device recognizes pedestrians and vehicles and sends that information to a server. The server replaces these with fantasy characters and animations, allowing the user to enjoy a different theme each day according to their settings.
[0295] An example of a prompt would be, "Add fantasy characters to the scenery you see during your commute." This prompt prompts the generative AI model to prepare to provide relevant content.
[0296] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0297] Step 1:
[0298] The device uses its camera to acquire real-time video of the environment. It receives real-world video light data as input and generates compressed video frames as output. Specifically, a high-performance processor compresses the video data and performs filtering to remove noise.
[0299] Step 2:
[0300] The device executes a machine learning algorithm to recognize objects in compressed video frames. It takes compressed video frames as input and generates a dataset containing object types and location information as output. Specifically, a model such as TensorFlow Lite analyzes the data and extracts feature vectors for the identified objects.
[0301] Step 3:
[0302] The device transmits information about recognized objects to the server via communication. It receives a dataset containing object type and location information as input and generates formatted data for transfer to the server as output. The data is structured in JSON format and transmitted rapidly over the network.
[0303] Step 4:
[0304] Based on the object information received by the server, it uses a generative AI model to convert it into virtual elements. It receives the object information as input and generates a dataset of virtual elements as output. The server then runs the generative AI model, selects a virtual character according to the user's settings, and adds the necessary behavioral information.
[0305] Step 5:
[0306] The server returns the virtual element data it generated to the terminal. It receives a dataset of virtual elements as input and generates the optimized data transmitted to the terminal as output. The data contains the information necessary for AR display and is efficiently transferred by high-speed streaming.
[0307] Step 6:
[0308] The terminal synthesizes the virtual elements received with the real video and displays it to the user. It receives the virtual element data and real-time video as input and generates the synthesized video presented to the user as output. In a specific operation, AR technology is used, and virtual characters are seamlessly overlapped and displayed in the real environment.
[0309] Step 7:
[0310] The user selects and changes virtual elements and backgrounds through the interface. It receives the user's setting input as input and generates data reflecting the updated virtual element settings as output. Using the operation means, the user can intuitively customize by touch operations and the like.
[0311] (Application Example 1)
[0312] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0313] There is a need to expand the area of new information presentation to users through the identification of objects in the real world and the combination with virtual display elements based on them. However, in existing systems, it is difficult to provide a consistent experience that fuses real and virtual information, and especially when there are multiple options, there is a lack of a method to efficiently present information according to the user's intention. The purpose of the present invention is to solve these problems.
[0314] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0315] In this invention, the server includes a video acquisition means for capturing the real environment, an analysis means for analyzing visual information obtained from the video acquisition means and recognizing objects, and an ambiguity removal means for identifying products when a user acquires video. This makes it possible to appropriately identify objects in the real world and efficiently provide virtual display elements according to the user's preferences.
[0316] "Image acquisition means" refers to devices or software that have the function of capturing the real environment.
[0317] An "analysis means" is an algorithm or system that receives visual information obtained from an image acquisition means, analyzes it, and recognizes objects.
[0318] "Conversion means" refers to a process or apparatus for converting entity information identified by analysis means into virtual display elements selected by the user.
[0319] "Output means" refers to devices or display technologies that present information to users by overlaying virtual display elements generated by conversion means onto the real environment.
[0320] "Ambiguousness elimination methods" are algorithms and techniques used by users to clearly identify identifiable products and objects when acquiring video footage.
[0321] The "presentation adjustment means" is a function for controlling and displaying story information corresponding to products identified using the ambiguity exclusion means, in accordance with the user's requests and context.
[0322] A "tracking mechanism" is a system that estimates the angle and position of visual information obtained from a video acquisition mechanism and adjusts the appearance of virtual display elements in real time.
[0323] The system for realizing this invention is implemented through a combination of many devices and software. The terminal used by the user to acquire video footage utilizes a camera and smart glasses to capture the real environment. The resulting video data is processed in real time within the terminal, undergoes necessary preprocessing, and is then sent to a server. The server receives the video data and uses machine learning algorithms (e.g., TensorFlow or PyTorch) to identify objects. The recognized object information is cross-referenced with information in a database and converted into a virtual character according to the request.
[0324] The converted character information is returned to the user's device and displayed overlaid on the real world using AR technology (e.g., ARKit / ARCore). Users can customize virtual display elements and background themes through the device's interface. For example, in a shopping mall, if a user picks up a piece of clothing and scans it with their camera, a virtual character can appear on the screen and try on the clothing. To enhance this user experience, a generative AI model is used with prompts such as, "In the next scene, the virtual character should try on the apparel item selected by the customer and visually represent the features of the item." These prompts function as input to enable more precise information presentation and support the object recognition and information generation processes.
[0325] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0326] Step 1:
[0327] The user acquires video using their device. The device's camera captures the real environment, acquiring raw video data. This data is compressed for efficient processing and undergoes pre-processing such as color correction and resolution adjustment. The input is the captured raw video data, and the output is the processed video data.
[0328] Step 2:
[0329] The device analyzes the processed video data and recognizes objects. It uses machine learning algorithms to analyze the contours and features of objects in the video and identify them. In this step, a trained model (e.g., TensorFlow) is used to assign an object category to each pixel. The input is the processed video data from the previous step, and the output is a list of recognized objects and their locations.
[0330] Step 3:
[0331] The terminal sends information about recognized objects to the server. The server receives this information, consults its database, and prepares to convert it into a virtual display element set by the user. The input is a list of objects and their locations, and the output is candidate information for the virtual character needed for the conversion.
[0332] Step 4:
[0333] The server selects an appropriate virtual character based on candidate virtual character information, according to the user's chosen theme and preferences. The selected character is then configured with the necessary movement patterns and postures. The input is candidate virtual character information, and the output is a set of specific virtual characters.
[0334] Step 5:
[0335] The server sends character data to the device, which receives it. The device then uses AR technology (e.g., ARKit / ARCore) to overlay the virtual character onto the real environment. This allows the user to enjoy interacting with the virtual character. The input is a specific set of virtual characters, and the output is a composite image displayed in the user's field of view.
[0336] Step 6:
[0337] Users can change the displayed virtual elements and background themes through the device interface. This allows for a customized experience tailored to the user's preferences. Input is user action information, and output is updated theme and virtual character information.
[0338] Step 7:
[0339] The server uses a generative AI model to generate story information for the virtual character. This process uses prompts to add more detailed elements. An example of a prompt in this step is, "In the next scene, the virtual character should try on the apparel item selected by the customer and visually represent the features of that item." The input is the updated virtual character information, and the output is the added story information.
[0340] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0341] The system of the present invention virtually augments information from the real environment and provides an interactive experience that responds to the user's emotions. Specific embodiments are described below.
[0342] System Configuration
[0343] This system uses a camera-equipped terminal to capture the real environment and incorporates recognition means for analyzing that data. This allows it to detect objects (people, vehicles, furniture, etc.) within the user's field of view and collect that information.
[0344] Introducing an emotional engine
[0345] The system uses an emotion engine that infers emotions by analyzing the user's facial expressions and voice. It evaluates data collected by the device's built-in microphone and camera to identify emotional states such as joy, anger, sadness, and happiness. Based on this, a virtual character or theme that matches the user's emotions is selected.
[0346] Character generation and tracking
[0347] The server generates a selected virtual character based on object information and user emotion data. Here, the character's movements and facial expressions are adjusted to match the user's emotions. The terminal uses tracking mechanisms to ensure the character is always seamlessly integrated with the real environment.
[0348] User interaction
[0349] Users can select and change the displayed characters and background themes through the interface. If the user's emotions change, the emotion engine detects this and dynamically updates the characters and background themes.
[0350] Learning and customization
[0351] The system learns the user's past emotional patterns and selection history, and based on that, provides a more comfortable and personalized experience. This feature allows the system to learn the user's preferences, making the suggested virtual characters and themes more personally suited to them.
[0352] Example: Personalized home entertainment
[0353] For example, when a user is relaxing at home, if the emotion engine detects the user's calm emotions, the system will select a character with a calm and relaxed theme and adjust the room's interior background accordingly. This allows the user to experience a virtual environment that makes their relaxation time even more pleasant.
[0354] These features enable the present invention to enhance the user's daily experience, making it more personalized and emotionally resonant.
[0355] The following describes the processing flow.
[0356] Step 1:
[0357] The device uses its built-in camera to capture the real-world environment surrounding the user in real time. The camera footage undergoes initial processing and is converted into a format that is easy to analyze.
[0358] Step 2:
[0359] The device performs object recognition on the captured data, identifying people and objects within the video. The recognized object information (location, type, size, etc.) is then organized for processing.
[0360] Step 3:
[0361] The device's emotion engine uses the camera and microphone to analyze the user's emotions from their face and voice. It identifies the emotional state (joy, anger, sadness, etc.) and outputs this information.
[0362] Step 4:
[0363] The device sends data about objects and emotional states to the server. This ensures that the information is collected in a secure manner.
[0364] Step 5:
[0365] The server selects a virtual character that suits the user based on the received object information and emotional state. Past usage history and emotional patterns are also taken into consideration during the selection process.
[0366] Step 6:
[0367] The server adjusts the selected virtual character's movements and facial expressions to match its emotional state. This adjusted character information is then sent to the terminal.
[0368] Step 7:
[0369] The device uses the received virtual character information to overlay the character onto the real-world image. This display is precisely integrated into the user's field of view through the device's tracking mechanism.
[0370] Step 8:
[0371] Users can manually change character types and background themes through the interface. Furthermore, the display automatically updates when changes in the user's emotions are detected.
[0372] Step 9:
[0373] The server learns from user choices and sentiment data, updating its algorithms to provide a more optimized experience for subsequent uses. This learning enhances the system's personalization capabilities.
[0374] (Example 2)
[0375] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0376] There is a growing demand for seamlessly blending real and virtual environments to provide interactive experiences that respond to user emotions. However, conventional technologies struggle to adapt to dynamic changes in the real environment and evolving user emotions, resulting in limited experiences. This invention aims to address this issue by flexibly responding to changes in user emotions and the environment, thereby providing a more personalized experience.
[0377] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0378] In this invention, the server includes a device for recording the real environment, means for identifying objects using a processing device, means for integrating virtual entities into the real environment using a generation device, and means for reflecting the characteristics of the virtual entities using an emotion analysis device. This makes it possible to provide a seamless and personalized interactive experience that responds to the user's emotions and changes in the real environment.
[0379] A "device for recording the real environment" is a hardware device that uses cameras, microphones, etc., to acquire information about the physical surroundings as digital data.
[0380] "Means for identifying objects using a processing device" refers to algorithms and hardware configurations for analyzing acquired digital data and identifying specific objects or elements.
[0381] "Means for integrating virtual entities into the real environment using a generation device" refers to technology for generating virtual entities and displaying them by overlaying them onto real-world video and audio information.
[0382] "A means of using an emotion analysis device to reflect the characteristics of a virtual entity" refers to a system that analyzes the user's emotional state and uses the results to influence the behavior and appearance of a virtual entity.
[0383] Modes for carrying out the invention
[0384] The system of the present invention integrates multiple devices and technologies to provide users with an interactive and personalized experience. The system acquires and analyzes information from the real environment, understands the user's emotions, and generates and displays virtual entities accordingly. Specific embodiments are described below.
[0385] Hardware and software configuration
[0386] 1. Recording the real environment
[0387] The device uses a camera and microphone to record the user's surroundings and audio in real time. This video and audio information is analyzed using libraries such as OpenCV.
[0388] 2. Object recognition
[0389] The device identifies people and objects based on the recorded data. This process utilizes machine learning techniques, modeled using TensorFlow and Keras.
[0390] 3. Analysis of emotions
[0391] The server uses deep learning algorithms to infer emotions from the user's facial expressions and voice. Emotion analysis is performed by an emotion engine built on the user's historical data.
[0392] 4. Creating and displaying virtual entities
[0393] The server generates virtual entities using the Unity engine based on the collected information. The device integrates these entities into the real environment using AR technology and displays them to the user.
[0394] 5. Individualization and Learning
[0395] The system uses tools like Scikit-learn to learn from the user's past emotions and choices, and then suggests suitable virtual entities and environments for the next time.
[0396] Specific example
[0397] For example, when a user wants to relax, the emotion analysis engine detects the user's calm emotions. In this case, the system selects a virtual entity with a calm and relaxing theme and adjusts the room's interior to create a virtually relaxing atmosphere. This allows the user to enjoy their relaxation time more comfortably.
[0398] Example of a prompt
[0399] "Please suggest a virtual character suitable for when the user is relaxing. Characteristics: calm, peaceful, surrounded by nature."
[0400] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0401] Step 1:
[0402] The device activates its camera and microphone to record the user's surroundings and audio. It uses real-world video and audio as input and outputs video and audio in digital data format. Specifically, the device captures video and audio at regular intervals and saves them as a data stream.
[0403] Step 2:
[0404] The device analyzes the acquired video data to detect objects. The input is the video data obtained in step 1, and an object recognition algorithm (for example, using OpenCV) is applied to generate object information, including the position and type of each object, as output. Specifically, the device processes the video frames sequentially to identify people, furniture, and other objects.
[0405] Step 3:
[0406] The device infers the user's emotions based on the acquired audio and facial video. The input is the audio and facial video data obtained in step 1, and an emotion analysis model (for example, a deep learning model using TensorFlow) is used to identify the user's emotional state as output. Specifically, the device profiles the tone of the voice and facial feature points.
[0407] Step 4:
[0408] The server generates virtual entities using object information and emotion data. The input is the data obtained in steps 2 and 3, and the server uses the Unity engine to design the virtual entities, generating 3D models and animation data as output. Specifically, the server selects an entity that matches the user's current emotion and performs individual customization.
[0409] Step 5:
[0410] The device integrates and displays the generated virtual entities in the real environment. The 3D model and animation data obtained in step 4 are used as input, and the virtual entities are rendered within the user's visual field via AR technology as output. Specifically, the device continuously updates its display according to location information and the user's movement.
[0411] Step 6:
[0412] Users interact with virtual entities and backgrounds through a provided interface. The system incorporates user selections and actions as input, dynamically updating entities and backgrounds based on this information and adjusting the display as feedback. Specifically, the interface instantly reflects the user's actions on the screen.
[0413] Step 7:
[0414] The server learns from past user sentiment data and selection history to provide appropriate suggestions for the next time. The input is historical data, and machine learning algorithms are used to generate personalized virtual entities and theme suggestions as output. Specifically, the system periodically analyzes the data and optimizes content to suit the user.
[0415] (Application Example 2)
[0416] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0417] Providing flexible and personalized customer service that responds to the diverse emotions and needs of visitors in physical stores is difficult with traditional methods. Furthermore, responding quickly to changes in visitors' emotions and suggesting the most suitable services and products is a challenge.
[0418] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0419] In this invention, the server includes an image acquisition means for acquiring the real environment, an emotion analysis means for analyzing the emotional state of visitors, and a service proposal means for selecting a customer service style based on the emotion analysis means. This enables immediate and appropriate service proposals that respond to the emotions and needs of visitors.
[0420] "Image acquisition means" refers to a device or technology for collecting visual information from the real environment.
[0421] A "judgment tool" is a function that analyzes acquired visual information and identifies a specific object.
[0422] A "conversion method" is a process that converts identified target information into a virtual target set by the user.
[0423] "Output means" refers to a mechanism that overlays a virtual object onto the real environment and performs integrated visual display.
[0424] "Emotional analysis techniques" are technologies used to analyze and infer the emotional state of visitors.
[0425] A "service proposal method" is a method for presenting the optimal customer service style and services based on information obtained through emotion analysis.
[0426] A "tracking method" is a technology for estimating the location of visual information and dynamically adjusting the display of a virtual object.
[0427] The present invention describes embodiments for carrying out the invention. The present invention realizes a system that provides personalized customer service in a physical store that responds to the emotions of the visitors.
[0428] The server uses an image acquisition device to capture the real environment. This image acquisition device is a general-purpose camera device that captures visitors' facial expressions and actions. The captured image data is analyzed by decision software to identify specific objects and features. The decision software includes image processing algorithms using OpenCV and analysis models using TensorFlow.
[0429] Furthermore, the server runs emotion analysis software to infer the visitor's emotions from the collected image and audio data. Audio data acquired by the microphone is also used in the emotion analysis. This makes it possible to identify the visitor's emotional state, such as whether they are relaxed or in a hurry.
[0430] Users can observe and utilize dynamically integrated virtual objects through smart devices. For example, a virtual character superimposed onto the real environment appears through smart glasses, supporting comfortable customer service. The display position and movement of the virtual character are dynamically adjusted by tracking software, enabling natural interaction.
[0431] As a concrete example, when a visitor in a music store wants to relax, the system analyzes the visitor's facial expressions and tone of voice to suggest a demo of relaxing music and then provides the most suitable music and service. The prompt, "Please suggest music and in-store presentation methods that are suitable for a customer who wants to relax," is input into the AI model, enabling the system to provide optimal content suggestions.
[0432] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0433] Step 1:
[0434] The server uses a camera device to capture the visitor's facial expressions and movements. This provides real-world image data as input. The server receives this image data, performs necessary preprocessing to reduce noise, and converts it into a format suitable for analysis.
[0435] Step 2:
[0436] The server inputs the pre-processed image data into the decision-making software. Using a machine learning model powered by TensorFlow, it identifies important objects and features within the image. This data processing outputs the location and attribute information of the objects (e.g., emotion estimation data from face recognition).
[0437] Step 3:
[0438] The server analyzes the identified facial recognition data using emotion analysis software. Input includes facial feature data and audio data. The software performs emotional calculations and provides output that identifies the visitor's emotional state (e.g., happy, calm, anxious).
[0439] Step 4:
[0440] The terminal generates a virtual display based on output data to overlay a virtual character onto the smart device. Graphic rendering technology is used for data processing, dynamically generating a character that expresses the identified emotion. This enables interaction tailored to the visitor's mood.
[0441] Step 5:
[0442] The user inputs a prompt into the generating AI model. Specifically, the prompt used is, "Please suggest music and store presentation methods that would be suitable when a customer wants to relax." Based on the input prompt, the generating AI model outputs the most suitable suggestions.
[0443] Step 6:
[0444] The terminal displays information to visitors and staff on their smart devices based on the outputted suggestions. These suggestions include specific customer service styles and product descriptions, providing visual and auditory feedback to enhance the user experience.
[0445] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0446] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0447] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0448] [Third Embodiment]
[0449] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0450] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0451] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0452] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0453] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0455] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0456] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0457] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0458] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0459] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0460] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0461] The present invention is a system for virtually extending information about the real environment, and is specifically implemented as follows.
[0462] System Configuration
[0463] First, the user's device is equipped with a camera as a means of capturing images. The device captures images of the real world within the user's field of view in real time. At this time, some kind of compression or pre-processing is performed so that the device can process the video data seamlessly.
[0464] Object recognition and data transmission
[0465] The terminal is equipped with recognition capabilities to detect objects from captured video footage, using machine learning algorithms to identify the type and location of objects. This information is compiled as object information and transmitted to a server via the network.
[0466] Character replacement process
[0467] The server performs a conversion process based on the received object information to replace it with a virtual character selected by the user. Here, appropriate character information is selected based on the user's pre-configured preferences and themes. Necessary movement and posture information is added to the converted character data.
[0468] AR display and interaction
[0469] Character data transmitted from the server is returned to the terminal and displayed in the user's field of view, superimposed on the real world, via an output device. This allows the user to visually experience an extraordinary scene where a virtual character adorns everyday landscapes. The user can freely change the character and background theme used using the terminal's interface.
[0470] Example: Usage scenario during commuting
[0471] When a user captures their surroundings through their device's camera during their commute, the device recognizes pedestrians and vehicles and sends that information to a server. The server then replaces these with anime or fantasy characters chosen by the user, which are then displayed in their field of view. Users can customize the characters' costumes and movements according to their settings, allowing them to enjoy their commute with a different theme each day.
[0472] In this way, the present invention combines reality and virtuality to provide users with new experiences.
[0473] The following describes the processing flow.
[0474] Step 1:
[0475] The device uses its built-in camera to capture the real-world environment within the user's field of view in real time. The captured video data undergoes pre-processing such as noise reduction and contrast adjustment.
[0476] Step 2:
[0477] The device analyzes pre-processed video data and uses an object recognition algorithm to detect people, vehicles, furniture, and other objects within the video. The detection results include the type, location, and size of the objects.
[0478] Step 3:
[0479] The terminal transmits object information obtained as a result of object recognition to the server. This information is structured as a dataset including object type and location information, and is transmitted rapidly over the network.
[0480] Step 4:
[0481] Based on the received object information, the server identifies the virtual character selected by the user. The server then decides which character to replace the selected character with, according to the user's settings, and prepares the necessary data for that replacement.
[0482] Step 5:
[0483] The server transforms the prepared virtual character data to correspond to the object's position and orientation. This transformation includes parameters to adjust the character's movement and orientation.
[0484] Step 6:
[0485] The server sends the converted virtual character data to the terminal. The transmitted information includes the virtual character's model and all the data necessary for display.
[0486] Step 7:
[0487] The device uses the received virtual character data to overlay the virtual character onto the real-world image. This display occurs in real time within the user's field of view.
[0488] Step 8:
[0489] Users can change and customize the displayed virtual character and background theme through the device's interface. Interactions are reflected immediately in response to user actions.
[0490] (Example 1)
[0491] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0492] When providing users with a new visual experience by virtually augmenting and displaying information from the real world, it is necessary to seamlessly and in real time merge the environment and virtual elements without requiring advanced processing power. Furthermore, flexibility is required, allowing users to freely customize virtual elements and background settings according to their individual preferences.
[0493] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0494] In this invention, the server includes a means for capturing environmental information, a means for recognizing objects by analyzing the images obtained from the capturing means, and a means for converting the identified object information into virtual elements. This makes it possible to merge images of the real world and virtual elements in real time and provide a customized visual experience based on the user's preferences.
[0495] "Mechanism of filming" refers to devices or functions that acquire information about the real world environment as images.
[0496] "Recognition means" refers to algorithms and processes for analyzing video data obtained from a shooting means and identifying specific objects.
[0497] "Conversion means" refers to the processes and functions used to replace recognized object information with virtual elements.
[0498] "Display means" refers to a means of presenting virtual elements to the user by overlaying them with images of the real environment.
[0499] "Transmission means" refers to the function of sending and receiving object information between servers and other devices via communication.
[0500] "Operational means" refers to interface functions that allow users to select or change virtual elements and background settings.
[0501] "Tracking means" refers to a function that recognizes the position of the image obtained from the shooting means and dynamically adjusts the display of virtual elements.
[0502] This invention is a system for virtually extending real-world environmental information and enriching the user's visual experience. This system consists of three entities: a terminal, a server, and a user.
[0503] The device is equipped with a camera and other means of capturing images, which acquires real-time video footage of the user's surroundings. To efficiently process the video data, the device performs pre-processing such as compression and filtering. A smart device with a high-performance processor is suitable for this process.
[0504] As a means of recognition, the device is equipped with software capable of executing machine learning algorithms to identify objects from captured video footage. Lightweight models such as TensorFlow Lite are often used in this process.
[0505] The server receives object information transmitted from the terminal via communication and uses a generative AI model as a means of conversion to replace it with a virtual element. Here, the generative AI model selects an appropriate virtual character based on themes and preferences pre-set by the user and adjusts its movements and posture. The generated virtual element is processed using cloud server resources as needed.
[0506] The completed character is returned to the device and displayed as a composite image with real-world footage. AR technology is used in this process, seamlessly overlaying the character onto the user's field of view, providing an integrated image where real and virtual elements merge. Users can select and customize virtual elements through various controls, enabling a personalized and interactive experience.
[0507] As a concrete example, consider a scenario where a user takes photos of their surroundings using their device's camera while commuting. The device recognizes pedestrians and vehicles and sends that information to a server. The server replaces these with fantasy characters and animations, allowing the user to enjoy a different theme each day according to their settings.
[0508] An example of a prompt would be, "Add fantasy characters to the scenery you see during your commute." This prompt prompts the generative AI model to prepare to provide relevant content.
[0509] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0510] Step 1:
[0511] The device uses a camera to acquire real-time video of the environment. It receives real-world video light data as input and generates compressed video frames as output. Specifically, a high-performance processor compresses the video data and performs filtering to remove noise.
[0512] Step 2:
[0513] The device executes a machine learning algorithm to recognize objects in compressed video frames. It takes compressed video frames as input and generates a dataset containing object types and location information as output. Specifically, a model such as TensorFlow Lite analyzes the data and extracts feature vectors for the identified objects.
[0514] Step 3:
[0515] The device transmits information about recognized objects to the server via communication. It receives a dataset containing object type and location information as input and generates formatted data for transfer to the server as output. The data is structured in JSON format and transmitted rapidly over the network.
[0516] Step 4:
[0517] Based on the object information received by the server, it uses a generative AI model to convert it into virtual elements. It receives the object information as input and generates a dataset of virtual elements as output. The server then runs the generative AI model, selects a virtual character according to the user's settings, and adds the necessary behavioral information.
[0518] Step 5:
[0519] The server generates virtual element data and sends it back to the terminal. It receives a dataset of virtual elements as input and generates optimized data sent to the terminal as output. The data contains information necessary for AR display and is efficiently transferred via high-speed streaming.
[0520] Step 6:
[0521] The device receives virtual elements, combines them with real-world video, and displays them to the user. It receives virtual element data and real-time video as input and generates a composite image presented to the user as output. In practice, AR technology is used to seamlessly overlay virtual characters onto the real-world environment.
[0522] Step 7:
[0523] The user selects and modifies virtual elements and backgrounds through the interface. It receives user settings as input and generates data reflecting the updated virtual element settings as output. Using various controls, users can intuitively customize the system through touch operations and other means.
[0524] (Application Example 1)
[0525] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0526] There is a need to expand the realm of information presentation to users through the identification of objects in the real world and the combination of these objects with virtual display elements. However, existing systems have difficulty providing a consistent experience that fuses real and virtual information, and there is a lack of efficient methods for presenting information that aligns with the user's intentions, especially when multiple choices exist. This invention aims to solve these problems.
[0527] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0528] In this invention, the server includes a video acquisition means for capturing the real environment, an analysis means for analyzing visual information obtained from the video acquisition means and recognizing objects, and an ambiguity removal means for identifying products when a user acquires video. This makes it possible to appropriately identify objects in the real world and efficiently provide virtual display elements according to the user's preferences.
[0529] "Image acquisition means" refers to devices or software that have the function of capturing the real environment.
[0530] An "analysis means" is an algorithm or system that receives visual information obtained from an image acquisition means, analyzes it, and recognizes objects.
[0531] "Conversion means" refers to a process or apparatus for converting entity information identified by analysis means into virtual display elements selected by the user.
[0532] "Output means" refers to devices or display technologies that present information to users by overlaying virtual display elements generated by conversion means onto the real environment.
[0533] "Ambiguousness elimination methods" are algorithms and techniques used by users to clearly identify identifiable products and objects when acquiring video footage.
[0534] The "presentation adjustment means" is a function for controlling and displaying story information corresponding to products identified using the ambiguity exclusion means, in accordance with the user's requests and context.
[0535] A "tracking mechanism" is a system that estimates the angle and position of visual information obtained from a video acquisition mechanism and adjusts the appearance of virtual display elements in real time.
[0536] The system for realizing this invention is implemented through a combination of many devices and software. The terminal used by the user to acquire video footage utilizes a camera and smart glasses to capture the real environment. The resulting video data is processed in real time within the terminal, undergoes necessary preprocessing, and is then sent to a server. The server receives the video data and uses machine learning algorithms (e.g., TensorFlow or PyTorch) to identify objects. The recognized object information is cross-referenced with information in a database and converted into a virtual character according to the request.
[0537] The converted character information is returned to the user's device and displayed overlaid on the real world using AR technology (e.g., ARKit / ARCore). Users can customize virtual display elements and background themes through the device's interface. For example, in a shopping mall, if a user picks up a piece of clothing and scans it with their camera, a virtual character can appear on the screen and try on the clothing. To enhance this user experience, a generative AI model is used with prompts such as, "In the next scene, the virtual character should try on the apparel item selected by the customer and visually represent the features of the item." These prompts function as input to enable more precise information presentation and support the object recognition and information generation processes.
[0538] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0539] Step 1:
[0540] The user acquires video using their device. The device's camera captures the real environment, acquiring raw video data. This data is compressed for efficient processing and undergoes pre-processing such as color correction and resolution adjustment. The input is the captured raw video data, and the output is the processed video data.
[0541] Step 2:
[0542] The device analyzes the processed video data and recognizes objects. It uses machine learning algorithms to analyze the contours and features of objects in the video and identify them. In this step, a trained model (e.g., TensorFlow) is used to assign an object category to each pixel. The input is the processed video data from the previous step, and the output is a list of recognized objects and their locations.
[0543] Step 3:
[0544] The terminal sends information about recognized objects to the server. The server receives this information, consults its database, and prepares to convert it into a virtual display element set by the user. The input is a list of objects and their locations, and the output is candidate information for the virtual character needed for the conversion.
[0545] Step 4:
[0546] The server selects an appropriate virtual character based on candidate virtual character information, according to the user's chosen theme and preferences. The selected character is then configured with the necessary movement patterns and postures. The input is candidate virtual character information, and the output is a set of specific virtual characters.
[0547] Step 5:
[0548] The server sends character data to the device, which receives it. The device then uses AR technology (e.g., ARKit / ARCore) to overlay the virtual character onto the real environment. This allows the user to enjoy interacting with the virtual character. The input is a specific set of virtual characters, and the output is a composite image displayed in the user's field of view.
[0549] Step 6:
[0550] Users can change the displayed virtual elements and background themes through the device interface. This allows for a customized experience tailored to the user's preferences. Input is user action information, and output is updated theme and virtual character information.
[0551] Step 7:
[0552] The server uses a generative AI model to generate story information for the virtual character. This process uses prompts to add more detailed elements. An example of a prompt in this step is, "In the next scene, the virtual character should try on the apparel item selected by the customer and visually represent the features of that item." The input is the updated virtual character information, and the output is the added story information.
[0553] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0554] The system of the present invention virtually augments information from the real environment and provides an interactive experience that responds to the user's emotions. Specific embodiments are described below.
[0555] System Configuration
[0556] This system uses a camera-equipped terminal to capture the real environment and incorporates recognition means for analyzing that data. This allows it to detect objects (people, vehicles, furniture, etc.) within the user's field of view and collect that information.
[0557] Introducing an emotional engine
[0558] The system uses an emotion engine that infers emotions by analyzing the user's facial expressions and voice. It evaluates data collected by the device's built-in microphone and camera to identify emotional states such as joy, anger, sadness, and happiness. Based on this, a virtual character or theme that matches the user's emotions is selected.
[0559] Character generation and tracking
[0560] The server generates a selected virtual character based on object information and user emotion data. Here, the character's movements and facial expressions are adjusted to match the user's emotions. The terminal uses tracking mechanisms to ensure the character is always seamlessly integrated with the real environment.
[0561] User interaction
[0562] Users can select and change the displayed characters and background themes through the interface. If the user's emotions change, the emotion engine detects this and dynamically updates the characters and background themes.
[0563] Learning and customization
[0564] The system learns the user's past emotional patterns and selection history, and based on that, provides a more comfortable and personalized experience. This feature allows the system to learn the user's preferences, making the suggested virtual characters and themes more personally suited to them.
[0565] Example: Personalized home entertainment
[0566] For example, when a user is relaxing at home, if the emotion engine detects the user's calm emotions, the system will select a character with a calm and relaxed theme and adjust the room's interior background accordingly. This allows the user to experience a virtual environment that makes their relaxation time even more pleasant.
[0567] These features enable the present invention to enhance the user's daily experience, making it more personalized and emotionally resonant.
[0568] The following describes the processing flow.
[0569] Step 1:
[0570] The device uses its built-in camera to capture the real-world environment surrounding the user in real time. The camera footage undergoes initial processing and is converted into a format that is easy to analyze.
[0571] Step 2:
[0572] The device performs object recognition on the captured data, identifying people and objects within the video. The recognized object information (location, type, size, etc.) is then organized for processing.
[0573] Step 3:
[0574] The device's emotion engine uses the camera and microphone to analyze the user's emotions from their face and voice. It identifies the emotional state (joy, anger, sadness, etc.) and outputs this information.
[0575] Step 4:
[0576] The device sends data about objects and emotional states to the server. This ensures that the information is collected in a secure manner.
[0577] Step 5:
[0578] The server selects a virtual character that suits the user based on the received object information and emotional state. Past usage history and emotional patterns are also taken into consideration during the selection process.
[0579] Step 6:
[0580] The server adjusts the selected virtual character's movements and facial expressions to match its emotional state. This adjusted character information is then sent to the terminal.
[0581] Step 7:
[0582] The device uses the received virtual character information to overlay the character onto the real-world image. This display is precisely integrated into the user's field of view through the device's tracking mechanism.
[0583] Step 8:
[0584] Users can manually change character types and background themes through the interface. Furthermore, the display automatically updates when changes in the user's emotions are detected.
[0585] Step 9:
[0586] The server learns from user choices and sentiment data, updating its algorithms to provide a more optimized experience for subsequent uses. This learning enhances the system's personalization capabilities.
[0587] (Example 2)
[0588] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0589] There is a growing demand for seamlessly blending real and virtual environments to provide interactive experiences that respond to user emotions. However, conventional technologies struggle to adapt to dynamic changes in the real environment and evolving user emotions, resulting in limited experiences. This invention aims to address this issue by flexibly responding to changes in user emotions and the environment, thereby providing a more personalized experience.
[0590] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0591] In this invention, the server includes a device for recording the real environment, means for identifying objects using a processing device, means for integrating virtual entities into the real environment using a generation device, and means for reflecting the characteristics of the virtual entities using an emotion analysis device. This makes it possible to provide a seamless and personalized interactive experience that responds to the user's emotions and changes in the real environment.
[0592] A "device for recording the real environment" is a hardware device that uses cameras, microphones, etc., to acquire information about the physical surroundings as digital data.
[0593] "Means for identifying objects using a processing device" refers to algorithms and hardware configurations for analyzing acquired digital data and identifying specific objects or elements.
[0594] "Means for integrating virtual entities into the real environment using a generation device" refers to technology for generating virtual entities and displaying them by overlaying them onto real-world video and audio information.
[0595] "A means of using an emotion analysis device to reflect the characteristics of a virtual entity" refers to a system that analyzes the user's emotional state and uses the results to influence the behavior and appearance of a virtual entity.
[0596] Modes for carrying out the invention
[0597] The system of the present invention integrates multiple devices and technologies to provide users with an interactive and personalized experience. The system acquires and analyzes information from the real environment, understands the user's emotions, and generates and displays virtual entities accordingly. Specific embodiments are described below.
[0598] Hardware and software configuration
[0599] 1. Recording the real environment
[0600] The device uses a camera and microphone to record the user's surroundings and audio in real time. This video and audio information is analyzed using libraries such as OpenCV.
[0601] 2. Object recognition
[0602] The device identifies people and objects based on the recorded data. This process utilizes machine learning techniques, modeled using TensorFlow and Keras.
[0603] 3. Analysis of emotions
[0604] The server uses deep learning algorithms to infer emotions from the user's facial expressions and voice. Emotion analysis is performed by an emotion engine built on the user's historical data.
[0605] 4. Creating and displaying virtual entities
[0606] The server generates virtual entities using the Unity engine based on the collected information. The device integrates these entities into the real environment using AR technology and displays them to the user.
[0607] 5. Individualization and Learning
[0608] The system uses tools like Scikit-learn to learn from the user's past emotions and choices, and then suggests suitable virtual entities and environments for the next time.
[0609] Specific example
[0610] For example, when a user wants to relax, the emotion analysis engine detects the user's calm emotions. In this case, the system selects a virtual entity with a calm and relaxing theme and adjusts the room's interior to create a virtually relaxing atmosphere. This allows the user to enjoy their relaxation time more comfortably.
[0611] Example of a prompt
[0612] "Please suggest a virtual character suitable for when the user is relaxing. Characteristics: calm, peaceful, surrounded by nature."
[0613] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0614] Step 1:
[0615] The device activates its camera and microphone to record the user's surroundings and audio. It uses real-world video and audio as input and outputs video and audio in digital data format. Specifically, the device captures video and audio at regular intervals and saves them as a data stream.
[0616] Step 2:
[0617] The device analyzes the acquired video data to detect objects. The input is the video data obtained in step 1, and an object recognition algorithm (for example, using OpenCV) is applied to generate object information, including the position and type of each object, as output. Specifically, the device processes the video frames sequentially to identify people, furniture, and other objects.
[0618] Step 3:
[0619] The device infers the user's emotions based on the acquired audio and facial video. The input is the audio and facial video data obtained in step 1, and an emotion analysis model (for example, a deep learning model using TensorFlow) is used to identify the user's emotional state as output. Specifically, the device profiles the tone of the voice and facial feature points.
[0620] Step 4:
[0621] The server generates virtual entities using object information and emotion data. The input is the data obtained in steps 2 and 3, and the server uses the Unity engine to design the virtual entities, generating 3D models and animation data as output. Specifically, the server selects an entity that matches the user's current emotion and performs individual customization.
[0622] Step 5:
[0623] The device integrates and displays the generated virtual entities in the real environment. The 3D model and animation data obtained in step 4 are used as input, and the virtual entities are rendered within the user's visual field via AR technology as output. Specifically, the device continuously updates its display according to location information and the user's movement.
[0624] Step 6:
[0625] Users interact with virtual entities and backgrounds through a provided interface. The system incorporates user selections and actions as input, dynamically updating entities and backgrounds based on this information and adjusting the display as feedback. Specifically, the interface instantly reflects the user's actions on the screen.
[0626] Step 7:
[0627] The server learns from past user sentiment data and selection history to provide appropriate suggestions for the next time. The input is historical data, and machine learning algorithms are used to generate personalized virtual entities and theme suggestions as output. Specifically, the system periodically analyzes the data and optimizes content to suit the user.
[0628] (Application Example 2)
[0629] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0630] Providing flexible and personalized customer service that responds to the diverse emotions and needs of visitors in physical stores is difficult with traditional methods. Furthermore, responding quickly to changes in visitors' emotions and suggesting the most suitable services and products is a challenge.
[0631] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0632] In this invention, the server includes an image acquisition means for acquiring the real environment, an emotion analysis means for analyzing the emotional state of visitors, and a service proposal means for selecting a customer service style based on the emotion analysis means. This enables immediate and appropriate service proposals that respond to the emotions and needs of visitors.
[0633] "Image acquisition means" refers to a device or technology for collecting visual information from the real environment.
[0634] A "judgment tool" is a function that analyzes acquired visual information and identifies a specific object.
[0635] A "conversion method" is a process that converts identified target information into a virtual target set by the user.
[0636] "Output means" refers to a mechanism that overlays a virtual object onto the real environment and performs integrated visual display.
[0637] "Emotional analysis techniques" are technologies used to analyze and infer the emotional state of visitors.
[0638] A "service proposal method" is a method for presenting the optimal customer service style and services based on information obtained through emotion analysis.
[0639] A "tracking method" is a technology for estimating the location of visual information and dynamically adjusting the display of a virtual object.
[0640] The present invention describes embodiments for carrying out the invention. The present invention realizes a system that provides personalized customer service in a physical store that responds to the emotions of the visitors.
[0641] The server uses an image acquisition device to capture the real environment. This image acquisition device is a general-purpose camera device that captures visitors' facial expressions and actions. The captured image data is analyzed by decision software to identify specific objects and features. The decision software includes image processing algorithms using OpenCV and analysis models using TensorFlow.
[0642] Furthermore, the server runs emotion analysis software to infer the visitor's emotions from the collected image and audio data. Audio data acquired by the microphone is also used in the emotion analysis. This makes it possible to identify the visitor's emotional state, such as whether they are relaxed or in a hurry.
[0643] Users can observe and utilize dynamically integrated virtual objects through smart devices. For example, a virtual character superimposed onto the real environment appears through smart glasses, supporting comfortable customer service. The display position and movement of the virtual character are dynamically adjusted by tracking software, enabling natural interaction.
[0644] As a concrete example, when a visitor in a music store wants to relax, the system analyzes the visitor's facial expressions and tone of voice to suggest a demo of relaxing music and then provides the most suitable music and service. The prompt, "Please suggest music and in-store presentation methods that are suitable for a customer who wants to relax," is input into the AI model, enabling the system to provide optimal content suggestions.
[0645] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0646] Step 1:
[0647] The server uses a camera device to capture the visitor's facial expressions and movements. This provides real-world image data as input. The server receives this image data, performs necessary preprocessing to reduce noise, and converts it into a format suitable for analysis.
[0648] Step 2:
[0649] The server inputs the pre-processed image data into the decision-making software. Using a machine learning model powered by TensorFlow, it identifies important objects and features within the image. This data processing outputs the location and attribute information of the objects (e.g., emotion estimation data from face recognition).
[0650] Step 3:
[0651] The server analyzes the identified facial recognition data using emotion analysis software. Input includes facial feature data and audio data. The software performs emotional calculations and provides output that identifies the visitor's emotional state (e.g., happy, calm, anxious).
[0652] Step 4:
[0653] The terminal generates a virtual display based on output data to overlay a virtual character onto the smart device. Graphic rendering technology is used for data processing, dynamically generating a character that expresses the identified emotion. This enables interaction tailored to the visitor's mood.
[0654] Step 5:
[0655] The user inputs a prompt into the generating AI model. Specifically, the prompt used is, "Please suggest music and store presentation methods that would be suitable when a customer wants to relax." Based on the input prompt, the generating AI model outputs the most suitable suggestions.
[0656] Step 6:
[0657] The terminal displays information to visitors and staff on their smart devices based on the outputted suggestions. These suggestions include specific customer service styles and product descriptions, providing visual and auditory feedback to enhance the user experience.
[0658] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0659] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0660] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0661] [Fourth Embodiment]
[0662] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0663] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0664] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0665] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0666] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0667] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0668] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0669] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0670] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0671] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0672] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0673] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0674] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0675] The present invention is a system for virtually extending information about the real environment, and is specifically implemented as follows.
[0676] System Configuration
[0677] First, the user's device is equipped with a camera as a means of capturing images. The device captures images of the real world within the user's field of view in real time. At this time, some kind of compression or pre-processing is performed so that the device can process the video data seamlessly.
[0678] Object recognition and data transmission
[0679] The terminal is equipped with recognition capabilities to detect objects from captured video footage, using machine learning algorithms to identify the type and location of objects. This information is compiled as object information and transmitted to a server via the network.
[0680] Character replacement process
[0681] The server performs a conversion process based on the received object information to replace it with a virtual character selected by the user. Here, appropriate character information is selected based on the user's pre-configured preferences and themes. Necessary movement and posture information is added to the converted character data.
[0682] AR display and interaction
[0683] Character data transmitted from the server is returned to the terminal and displayed in the user's field of view, superimposed on the real world, via an output device. This allows the user to visually experience an extraordinary scene where a virtual character adorns everyday landscapes. The user can freely change the character and background theme used using the terminal's interface.
[0684] Example: Usage scenario during commuting
[0685] When a user captures their surroundings through their device's camera during their commute, the device recognizes pedestrians and vehicles and sends that information to a server. The server then replaces these with anime or fantasy characters chosen by the user, which are then displayed in their field of view. Users can customize the characters' costumes and movements according to their settings, allowing them to enjoy their commute with a different theme each day.
[0686] In this way, the present invention combines reality and virtuality to provide users with new experiences.
[0687] The following describes the processing flow.
[0688] Step 1:
[0689] The device uses its built-in camera to capture the real-world environment within the user's field of view in real time. The captured video data undergoes pre-processing such as noise reduction and contrast adjustment.
[0690] Step 2:
[0691] The device analyzes pre-processed video data and uses an object recognition algorithm to detect people, vehicles, furniture, and other objects within the video. The detection results include the type, location, and size of the objects.
[0692] Step 3:
[0693] The terminal transmits object information obtained as a result of object recognition to the server. This information is structured as a dataset including object type and location information, and is transmitted rapidly over the network.
[0694] Step 4:
[0695] Based on the received object information, the server identifies the virtual character selected by the user. The server then decides which character to replace the selected character with, according to the user's settings, and prepares the necessary data for that replacement.
[0696] Step 5:
[0697] The server transforms the prepared virtual character data to correspond to the object's position and orientation. This transformation includes parameters to adjust the character's movement and orientation.
[0698] Step 6:
[0699] The server sends the converted virtual character data to the terminal. The transmitted information includes the virtual character's model and all the data necessary for display.
[0700] Step 7:
[0701] The device uses the received virtual character data to overlay the virtual character onto the real-world image. This display occurs in real time within the user's field of view.
[0702] Step 8:
[0703] Users can change and customize the displayed virtual character and background theme through the device's interface. Interactions are reflected immediately in response to user actions.
[0704] (Example 1)
[0705] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0706] When providing users with a new visual experience by virtually augmenting and displaying information from the real world, it is necessary to seamlessly and in real time merge the environment and virtual elements without requiring advanced processing power. Furthermore, flexibility is required, allowing users to freely customize virtual elements and background settings according to their individual preferences.
[0707] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0708] In this invention, the server includes a means for capturing environmental information, a means for recognizing objects by analyzing the images obtained from the capturing means, and a means for converting the identified object information into virtual elements. This makes it possible to merge images of the real world and virtual elements in real time and provide a customized visual experience based on the user's preferences.
[0709] "Mechanism of filming" refers to devices or functions that acquire information about the real world environment as images.
[0710] "Recognition means" refers to algorithms and processes for analyzing video data obtained from a shooting means and identifying specific objects.
[0711] "Conversion means" refers to the processes and functions used to replace recognized object information with virtual elements.
[0712] "Display means" refers to a means of presenting virtual elements to the user by overlaying them with images of the real environment.
[0713] "Transmission means" refers to the function of sending and receiving object information between servers and other devices via communication.
[0714] "Operational means" refers to interface functions that allow users to select or change virtual elements and background settings.
[0715] "Tracking means" refers to a function that recognizes the position of the image obtained from the shooting means and dynamically adjusts the display of virtual elements.
[0716] This invention is a system for virtually extending real-world environmental information and enriching the user's visual experience. This system consists of three components: a terminal, a server, and a user.
[0717] The device is equipped with a camera and other means of capturing images, which acquires real-time video footage of the user's surroundings. To efficiently process the video data, the device performs pre-processing such as compression and filtering. A smart device with a high-performance processor is suitable for this process.
[0718] As a means of recognition, the device is equipped with software capable of executing machine learning algorithms to identify objects from captured video footage. Lightweight models such as TensorFlow Lite are often used in this process.
[0719] The server receives object information transmitted from the terminal via communication and uses a generative AI model as a means of conversion to replace it with a virtual element. Here, the generative AI model selects an appropriate virtual character based on themes and preferences pre-set by the user and adjusts its movements and posture. The generated virtual element is processed using cloud server resources as needed.
[0720] The completed character is returned to the device and displayed as a composite image with real-world footage. AR technology is used in this process, seamlessly overlaying the character onto the user's field of view, providing an integrated image where real and virtual elements merge. Users can select and customize virtual elements through various controls, enabling a personalized and interactive experience.
[0721] As a concrete example, consider a scenario where a user takes photos of their surroundings using their device's camera while commuting. The device recognizes pedestrians and vehicles and sends that information to a server. The server replaces these with fantasy characters and animations, allowing the user to enjoy a different theme each day according to their settings.
[0722] An example of a prompt would be, "Add fantasy characters to the scenery you see during your commute." This prompt prompts the generative AI model to prepare to provide relevant content.
[0723] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0724] Step 1:
[0725] The device uses a camera to acquire real-time video of the environment. It receives real-world video light data as input and generates compressed video frames as output. Specifically, a high-performance processor compresses the video data and performs filtering to remove noise.
[0726] Step 2:
[0727] The device executes a machine learning algorithm to recognize objects in compressed video frames. It takes compressed video frames as input and generates a dataset containing object types and location information as output. Specifically, a model such as TensorFlow Lite analyzes the data and extracts feature vectors for the identified objects.
[0728] Step 3:
[0729] The device transmits information about recognized objects to the server via communication. It receives a dataset containing object type and location information as input and generates formatted data for transfer to the server as output. The data is structured in JSON format and transmitted rapidly over the network.
[0730] Step 4:
[0731] Based on the object information received by the server, it uses a generative AI model to convert it into virtual elements. It receives the object information as input and generates a dataset of virtual elements as output. The server then runs the generative AI model, selects a virtual character according to the user's settings, and adds the necessary behavioral information.
[0732] Step 5:
[0733] The server generates virtual element data and sends it back to the terminal. It receives a dataset of virtual elements as input and generates optimized data sent to the terminal as output. The data contains information necessary for AR display and is efficiently transferred via high-speed streaming.
[0734] Step 6:
[0735] The device receives virtual elements, combines them with real-world video, and displays them to the user. It receives virtual element data and real-time video as input and generates a composite image presented to the user as output. In practice, AR technology is used to seamlessly overlay virtual characters onto the real-world environment.
[0736] Step 7:
[0737] The user selects and modifies virtual elements and backgrounds through the interface. It receives user settings as input and generates data reflecting the updated virtual element settings as output. Using various controls, users can intuitively customize the system through touch operations and other means.
[0738] (Application Example 1)
[0739] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0740] There is a need to expand the realm of information presentation to users through the identification of objects in the real world and the combination of these objects with virtual display elements. However, existing systems have difficulty providing a consistent experience that fuses real and virtual information, and there is a lack of efficient methods for presenting information that aligns with the user's intentions, especially when multiple choices exist. This invention aims to solve these problems.
[0741] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0742] In this invention, the server includes a video acquisition means for capturing the real environment, an analysis means for analyzing visual information obtained from the video acquisition means and recognizing objects, and an ambiguity removal means for identifying products when a user acquires video. This makes it possible to appropriately identify objects in the real world and efficiently provide virtual display elements according to the user's preferences.
[0743] "Image acquisition means" refers to devices or software that have the function of capturing the real environment.
[0744] An "analysis means" is an algorithm or system that receives visual information obtained from an image acquisition means, analyzes it, and recognizes objects.
[0745] "Conversion means" refers to a process or apparatus for converting entity information identified by analysis means into virtual display elements selected by the user.
[0746] "Output means" refers to devices or display technologies that present information to users by overlaying virtual display elements generated by conversion means onto the real environment.
[0747] "Ambiguousness elimination methods" are algorithms and techniques used by users to clearly identify identifiable products and objects when acquiring video footage.
[0748] The "presentation adjustment means" is a function for controlling and displaying story information corresponding to products identified using the ambiguity exclusion means, in accordance with the user's requests and context.
[0749] A "tracking mechanism" is a system that estimates the angle and position of visual information obtained from a video acquisition mechanism and adjusts the appearance of virtual display elements in real time.
[0750] The system for realizing this invention is implemented through a combination of many devices and software. The terminal used by the user to acquire video footage utilizes a camera and smart glasses to capture the real environment. The resulting video data is processed in real time within the terminal, undergoes necessary preprocessing, and is then sent to a server. The server receives the video data and uses machine learning algorithms (e.g., TensorFlow or PyTorch) to identify objects. The recognized object information is cross-referenced with information in a database and converted into a virtual character according to the request.
[0751] The converted character information is returned to the user's device and displayed overlaid on the real world using AR technology (e.g., ARKit / ARCore). Users can customize virtual display elements and background themes through the device's interface. For example, in a shopping mall, if a user picks up a piece of clothing and scans it with their camera, a virtual character can appear on the screen and try on the clothing. To enhance this user experience, a generative AI model is used with prompts such as, "In the next scene, the virtual character should try on the apparel item selected by the customer and visually represent the features of the item." These prompts function as input to enable more precise information presentation and support the object recognition and information generation processes.
[0752] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0753] Step 1:
[0754] The user acquires video using their device. The device's camera captures the real environment, acquiring raw video data. This data is compressed for efficient processing and undergoes pre-processing such as color correction and resolution adjustment. The input is the captured raw video data, and the output is the processed video data.
[0755] Step 2:
[0756] The device analyzes the processed video data and recognizes objects. It uses machine learning algorithms to analyze the contours and features of objects in the video and identify them. In this step, a trained model (e.g., TensorFlow) is used to assign an object category to each pixel. The input is the processed video data from the previous step, and the output is a list of recognized objects and their locations.
[0757] Step 3:
[0758] The terminal sends information about recognized objects to the server. The server receives this information, consults its database, and prepares to convert it into a virtual display element set by the user. The input is a list of objects and their locations, and the output is candidate information for the virtual character needed for the conversion.
[0759] Step 4:
[0760] The server selects an appropriate virtual character based on candidate virtual character information, according to the user's chosen theme and preferences. The selected character is then configured with the necessary movement patterns and postures. The input is candidate virtual character information, and the output is a set of specific virtual characters.
[0761] Step 5:
[0762] The server sends character data to the device, which receives it. The device then uses AR technology (e.g., ARKit / ARCore) to overlay the virtual character onto the real environment. This allows the user to enjoy interacting with the virtual character. The input is a specific set of virtual characters, and the output is a composite image displayed in the user's field of view.
[0763] Step 6:
[0764] Users can change the displayed virtual elements and background themes through the device interface. This allows for a customized experience tailored to the user's preferences. Input is user action information, and output is updated theme and virtual character information.
[0765] Step 7:
[0766] The server uses a generative AI model to generate story information for the virtual character. This process uses prompts to add more detailed elements. An example of a prompt in this step is, "In the next scene, the virtual character should try on the apparel item selected by the customer and visually represent the features of that item." The input is the updated virtual character information, and the output is the added story information.
[0767] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0768] The system of the present invention virtually augments information from the real environment and provides an interactive experience that responds to the user's emotions. Specific embodiments are described below.
[0769] System Configuration
[0770] This system uses a camera-equipped terminal to capture the real environment and incorporates recognition means for analyzing that data. This allows it to detect objects (people, vehicles, furniture, etc.) within the user's field of view and collect that information.
[0771] Introducing an emotional engine
[0772] The system uses an emotion engine that infers emotions by analyzing the user's facial expressions and voice. It evaluates data collected by the device's built-in microphone and camera to identify emotional states such as joy, anger, sadness, and happiness. Based on this, a virtual character or theme that matches the user's emotions is selected.
[0773] Character generation and tracking
[0774] The server generates a selected virtual character based on object information and user emotion data. Here, the character's movements and facial expressions are adjusted to match the user's emotions. The terminal uses tracking mechanisms to ensure the character is always seamlessly integrated with the real environment.
[0775] User interaction
[0776] Users can select and change the displayed characters and background themes through the interface. If the user's emotions change, the emotion engine detects this and dynamically updates the characters and background themes.
[0777] Learning and customization
[0778] The system learns the user's past emotional patterns and selection history, and based on that, provides a more comfortable and personalized experience. This feature allows the system to learn the user's preferences, making the suggested virtual characters and themes more personally suited to them.
[0779] Example: Personalized home entertainment
[0780] For example, when a user is relaxing at home, if the emotion engine detects the user's calm emotions, the system will select a character with a calm and relaxed theme and adjust the room's interior background accordingly. This allows the user to experience a virtual environment that makes their relaxation time even more pleasant.
[0781] These features enable the present invention to enhance the user's daily experience, making it more personalized and emotionally resonant.
[0782] The following describes the processing flow.
[0783] Step 1:
[0784] The device uses its built-in camera to capture the real-world environment surrounding the user in real time. The camera footage undergoes initial processing and is converted into a format that is easy to analyze.
[0785] Step 2:
[0786] The device performs object recognition from the captured data, identifying people and objects in the video. The recognized object information (location, type, size, etc.) is then organized for processing.
[0787] Step 3:
[0788] The device's emotion engine uses the camera and microphone to analyze the user's emotions from their face and voice. It identifies the emotional state (joy, anger, sadness, etc.) and outputs this information.
[0789] Step 4:
[0790] The device sends data about objects and emotional states to the server. This ensures that the information is collected in a secure manner.
[0791] Step 5:
[0792] The server selects a virtual character that suits the user based on the received object information and emotional state. Past usage history and emotional patterns are also taken into consideration during the selection process.
[0793] Step 6:
[0794] The server adjusts the selected virtual character's movements and facial expressions to match its emotional state. This adjusted character information is then sent to the terminal.
[0795] Step 7:
[0796] The device uses the received virtual character information to overlay the character onto the real-world image. This display is precisely integrated into the user's field of view through the device's tracking mechanism.
[0797] Step 8:
[0798] Users can manually change character types and background themes through the interface. Furthermore, the display automatically updates when changes in the user's emotions are detected.
[0799] Step 9:
[0800] The server learns from user choices and sentiment data, updating its algorithms to provide a more optimized experience for subsequent uses. This learning enhances the system's personalization capabilities.
[0801] (Example 2)
[0802] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0803] There is a growing demand for seamlessly blending real and virtual environments to provide interactive experiences that respond to user emotions. However, conventional technologies struggle to adapt to dynamic changes in the real environment and evolving user emotions, resulting in limited experiences. This invention aims to address this issue by flexibly responding to changes in user emotions and the environment, thereby providing a more personalized experience.
[0804] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0805] In this invention, the server includes a device for recording the real environment, means for identifying objects using a processing device, means for integrating virtual entities into the real environment using a generation device, and means for reflecting the characteristics of the virtual entities using an emotion analysis device. This makes it possible to provide a seamless and personalized interactive experience that responds to the user's emotions and changes in the real environment.
[0806] A "device for recording the real environment" is a hardware device that uses cameras, microphones, etc., to acquire information about the physical surroundings as digital data.
[0807] "Means for identifying objects using a processing device" refers to algorithms and hardware configurations for analyzing acquired digital data and identifying specific objects or elements.
[0808] "Means for integrating virtual entities into the real environment using a generation device" refers to technology for generating virtual entities and displaying them by overlaying them onto real-world video and audio information.
[0809] "A means of using an emotion analysis device to reflect the characteristics of a virtual entity" refers to a system that analyzes the user's emotional state and uses the results to influence the behavior and appearance of a virtual entity.
[0810] Modes for carrying out the invention
[0811] The system of the present invention integrates multiple devices and technologies to provide users with an interactive and personalized experience. The system acquires and analyzes information from the real environment, understands the user's emotions, and generates and displays virtual entities accordingly. Specific embodiments are described below.
[0812] Hardware and software configuration
[0813] 1. Recording the real environment
[0814] The device uses a camera and microphone to record the user's surroundings and audio in real time. This video and audio information is analyzed using libraries such as OpenCV.
[0815] 2. Object recognition
[0816] The device identifies people and objects based on the recorded data. This process utilizes machine learning techniques, modeled using TensorFlow and Keras.
[0817] 3. Analysis of emotions
[0818] The server uses deep learning algorithms to infer emotions from the user's facial expressions and voice. Emotion analysis is performed by an emotion engine built on the user's historical data.
[0819] 4. Creating and displaying virtual entities
[0820] The server generates virtual entities using the Unity engine based on the collected information. The device integrates these entities into the real environment using AR technology and displays them to the user.
[0821] 5. Individualization and Learning
[0822] The system uses tools like Scikit-learn to learn from the user's past emotions and choices, and then suggests suitable virtual entities and environments for the next time.
[0823] Specific example
[0824] For example, when a user wants to relax, the emotion analysis engine detects the user's calm emotions. In this case, the system selects a virtual entity with a calm and relaxing theme and adjusts the room's interior to create a virtually relaxing atmosphere. This allows the user to enjoy their relaxation time more comfortably.
[0825] Example of a prompt
[0826] "Please suggest a virtual character suitable for when the user is relaxing. Characteristics: calm, peaceful, surrounded by nature."
[0827] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0828] Step 1:
[0829] The device activates its camera and microphone to record the user's surroundings and audio. It uses real-world video and audio as input and outputs video and audio in digital data format. Specifically, the device captures video and audio at regular intervals and saves them as a data stream.
[0830] Step 2:
[0831] The device analyzes the acquired video data to detect objects. The input is the video data obtained in step 1, and an object recognition algorithm (for example, using OpenCV) is applied to generate object information, including the position and type of each object, as output. Specifically, the device processes the video frames sequentially to identify people, furniture, and other objects.
[0832] Step 3:
[0833] The device infers the user's emotions based on the acquired audio and facial video. The input is the audio and facial video data obtained in step 1, and an emotion analysis model (for example, a deep learning model using TensorFlow) is used to identify the user's emotional state as output. Specifically, the device profiles the tone of the voice and facial feature points.
[0834] Step 4:
[0835] The server generates virtual entities using object information and emotion data. The input is the data obtained in steps 2 and 3, and the server uses the Unity engine to design the virtual entities, generating 3D models and animation data as output. Specifically, the server selects an entity that matches the user's current emotion and performs individual customization.
[0836] Step 5:
[0837] The device integrates and displays the generated virtual entities in the real environment. The 3D model and animation data obtained in step 4 are used as input, and the virtual entities are rendered within the user's visual field via AR technology as output. Specifically, the device continuously updates its display according to location information and the user's movement.
[0838] Step 6:
[0839] Users interact with virtual entities and backgrounds through a provided interface. The system incorporates user selections and actions as input, dynamically updating entities and backgrounds based on this information and adjusting the display as feedback. Specifically, the interface instantly reflects the user's actions on the screen.
[0840] Step 7:
[0841] The server learns from past user sentiment data and selection history to provide appropriate suggestions for the next time. The input is historical data, and machine learning algorithms are used to generate personalized virtual entities and theme suggestions as output. Specifically, the system periodically analyzes the data and optimizes content to suit the user.
[0842] (Application Example 2)
[0843] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0844] Providing flexible and personalized customer service that responds to the diverse emotions and needs of visitors in physical stores is difficult with traditional methods. Furthermore, responding quickly to changes in visitors' emotions and suggesting the most suitable services and products is a challenge.
[0845] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0846] In this invention, the server includes an image acquisition means for acquiring the real environment, an emotion analysis means for analyzing the emotional state of visitors, and a service proposal means for selecting a customer service style based on the emotion analysis means. This enables immediate and appropriate service proposals that respond to the emotions and needs of visitors.
[0847] "Image acquisition means" refers to a device or technology for collecting visual information from the real environment.
[0848] A "judgment tool" is a function that analyzes acquired visual information and identifies a specific object.
[0849] A "conversion method" is a process that converts identified target information into a virtual target set by the user.
[0850] "Output means" refers to a mechanism that overlays a virtual object onto the real environment and performs integrated visual display.
[0851] "Emotional analysis techniques" are technologies used to analyze and infer the emotional state of visitors.
[0852] A "service proposal method" is a method for presenting the optimal customer service style and services based on information obtained through emotion analysis.
[0853] A "tracking method" is a technology for estimating the location of visual information and dynamically adjusting the display of a virtual object.
[0854] The present invention describes embodiments for carrying out the invention. The present invention realizes a system that provides personalized customer service in a physical store that responds to the emotions of the visitors.
[0855] The server uses an image acquisition device to capture the real environment. This image acquisition device is a general-purpose camera device that captures visitors' facial expressions and actions. The captured image data is analyzed by decision software to identify specific objects and features. The decision software includes image processing algorithms using OpenCV and analysis models using TensorFlow.
[0856] Furthermore, the server runs emotion analysis software to infer the visitor's emotions from the collected image and audio data. Audio data acquired by the microphone is also used in the emotion analysis. This makes it possible to identify the visitor's emotional state, such as whether they are relaxed or in a hurry.
[0857] Users can observe and utilize dynamically integrated virtual objects through smart devices. For example, a virtual character superimposed onto the real environment appears through smart glasses, supporting comfortable customer service. The display position and movement of the virtual character are dynamically adjusted by tracking software, enabling natural interaction.
[0858] As a concrete example, when a visitor in a music store wants to relax, the system analyzes the visitor's facial expressions and tone of voice to suggest a demo of relaxing music and then provides the most suitable music and service. The prompt, "Please suggest music and in-store presentation methods that are suitable for a customer who wants to relax," is input into the AI model, enabling the system to provide optimal content suggestions.
[0859] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0860] Step 1:
[0861] The server uses a camera device to capture the visitor's facial expressions and movements. This provides real-world image data as input. The server receives this image data, performs necessary preprocessing to reduce noise, and converts it into a format suitable for analysis.
[0862] Step 2:
[0863] The server inputs the pre-processed image data into the decision-making software. Using a machine learning model powered by TensorFlow, it identifies important objects and features within the image. This data processing outputs the location and attribute information of the objects (e.g., emotion estimation data from face recognition).
[0864] Step 3:
[0865] The server analyzes the identified facial recognition data using emotion analysis software. Input includes facial feature data and audio data. The software performs emotional calculations and provides output that identifies the visitor's emotional state (e.g., happy, calm, anxious).
[0866] Step 4:
[0867] The terminal generates a virtual display based on output data to overlay a virtual character onto the smart device. Graphic rendering technology is used for data processing, dynamically generating a character that expresses the identified emotion. This enables interaction tailored to the visitor's mood.
[0868] Step 5:
[0869] The user inputs a prompt into the generating AI model. Specifically, the prompt used is, "Please suggest music and store presentation methods that would be suitable when a customer wants to relax." Based on the input prompt, the generating AI model outputs the most suitable suggestions.
[0870] Step 6:
[0871] The terminal displays information to visitors and staff on their smart devices based on the outputted suggestions. These suggestions include specific customer service styles and product descriptions, providing visual and auditory feedback to enhance the user experience.
[0872] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0873] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0874] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0875] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0876] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0877] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0878] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0879] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0880] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0881] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0882] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0883] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0884] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0885] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0886] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0887] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0888] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0889] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0890] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0891] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0892] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0893] The following is further disclosed regarding the embodiments described above.
[0894] (Claim 1)
[0895] An imaging means for capturing the real environment,
[0896] A recognition means that analyzes the image acquired from the imaging means to detect an object,
[0897] A conversion means that replaces object information detected by the recognition means with a virtual character set by the user,
[0898] An output means that displays the virtual character generated by the conversion means superimposed on the real environment,
[0899] A system that includes this.
[0900] (Claim 2)
[0901] The system according to claim 1, comprising an interface means for receiving user input and allowing the user to select and change virtual characters and background themes.
[0902] (Claim 3)
[0903] The system according to claim 1, further comprising tracking means for estimating the pose of an image acquired from an imaging means and dynamically adjusting the display of a virtual character.
[0904] "Example 1"
[0905] (Claim 1)
[0906] A means of capturing environmental information,
[0907] A recognition means that analyzes images obtained from a shooting means to identify objects,
[0908] A transformation means that replaces object information identified by a recognition means with a virtual element selected by the user,
[0909] A display means that overlays virtual elements generated by a conversion means onto environmental information and displays them,
[0910] A means for transmitting and receiving object information via communication,
[0911] A system that includes this.
[0912] (Claim 2)
[0913] The system according to claim 1, comprising an operating means for receiving user input and allowing the user to select or change virtual elements and background settings.
[0914] (Claim 3)
[0915] The system according to claim 1, further comprising tracking means for estimating the position of images obtained from a shooting means and dynamically adjusting the display of virtual elements.
[0916] "Application Example 1"
[0917] (Claim 1)
[0918] A means of acquiring video to capture the real environment,
[0919] An analysis means that analyzes the visual information obtained from the image acquisition means and recognizes objects,
[0920] A conversion means that converts the entity information identified by the analysis means into a virtual display element selected by the user,
[0921] An output means that overlays the virtual display elements generated by the conversion means onto the real environment to present information,
[0922] When a user acquires a video, a means of eliminating ambiguity is used to identify the product,
[0923] A presentation adjustment means for controlling the display of story information corresponding to a product identified using the ambiguity exclusion means,
[0924] A system that includes this.
[0925] (Claim 2)
[0926] The system according to claim 1, comprising control means for receiving user input and enabling the selection and modification of virtual display elements and background themes.
[0927] (Claim 3)
[0928] The system according to claim 1, further comprising tracking means for estimating the angle of visual information obtained from video acquisition means and adjusting the position of virtual display elements in real time.
[0929] "Example 2 of combining an emotion engine"
[0930] (Claim 1)
[0931] A device for recording the real environment,
[0932] A processing device that analyzes data obtained from the device to identify an object,
[0933] A generation device that converts object information recognized by the processing device into a virtual entity associated with the user,
[0934] A display device that overlays the virtual entities generated by the generation device onto the real environment,
[0935] An emotion analysis device that analyzes the user's emotions and reflects them in the characteristics of virtual entities,
[0936] A system that includes this.
[0937] (Claim 2)
[0938] The system according to claim 1, comprising an operating device that accepts input from a user and allows the user to select and change virtual entities and background themes.
[0939] (Claim 3)
[0940] The system according to claim 1, comprising a tracking device that estimates the location of data obtained from a device and dynamically adjusts the display of a virtual entity.
[0941] "Application example 2 of combining emotional engines"
[0942] (Claim 1)
[0943] Image acquisition means for capturing the real environment,
[0944] A determination means for identifying an object by analyzing the visual information obtained from the image acquisition means,
[0945] A conversion means that converts the target information identified by the determination means into a virtual target set by the user,
[0946] An output means for integrating and displaying the virtual object generated by the conversion means in the real environment,
[0947] A means of analyzing the emotional state of visitors,
[0948] A service proposal means that selects a customer service style based on the emotion analysis means,
[0949] A system that includes this.
[0950] (Claim 2)
[0951] The system according to claim 1, comprising an interface means for receiving user input and allowing selection and modification of virtual objects and background styles.
[0952] (Claim 3)
[0953] The system according to claim 1, further comprising tracking means for estimating the position of visual information obtained from image acquisition means and dynamically adjusting the display of a virtual object. [Explanation of Symbols]
[0954] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An imaging means for capturing the real environment, A recognition means that analyzes the image acquired from the imaging means to detect an object, A conversion means that replaces object information detected by the recognition means with a virtual character set by the user, An output means that displays the virtual character generated by the conversion means superimposed on the real environment, A system that includes this.
2. The system according to claim 1, comprising an interface means for receiving user input and allowing the user to select and change virtual characters and background themes.
3. The system according to claim 1, further comprising tracking means for estimating the pose of an image acquired from an imaging means and dynamically adjusting the display of a virtual character.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A