System
The system enhances AR experiences by capturing real-time video, analyzing it with image recognition, and generating interactive content using generative AI, addressing limitations in conventional AR systems to provide detailed and personalized information.
Patent Information
- Application Number
- JP2024115269
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-29
AI Technical Summary
Conventional augmented reality (AR) systems lack the capability to provide detailed real-time information and immersive experiences due to limited information availability and inadequate analysis of video data for generating immersive content.
A system that captures real-time video using a device's camera, analyzes it with an image recognition model, generates content using generative AI, and displays it in augmented reality format, allowing for user interaction to provide detailed information and personalized experiences.
Enriches users' understanding and experience of the real world by providing rich, real-time information and immersive content based on user interactions and emotional states.
Smart Images

Figure 2026014272000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In applications using conventional augmented reality (AR) technology, the information available to users in real time is limited, making it difficult to effectively provide a detailed understanding of the real world or new experiences.Another problem is the lack of established technology for properly analyzing video data captured by device cameras and generating and displaying immersive content for users. [Means for solving the problem]
[0005] In order to solve the above-mentioned problems, the present invention provides the following means: a means for capturing video in real time using a device's camera, including a means for transmitting the captured video data to a server; a means for the server to analyze the video data using an image recognition model and generate content using a generative AI based on the analysis results; a means for transmitting the generated content to the device, and displaying the content received by the device in an augmented reality format; and a means for displaying further information in response to user interaction, thereby realizing a system that effectively provides a detailed understanding of the real world and new experiences.
[0006] "Device" is a general term for electronic devices used by users, such as computers, smartphones, tablets, and head-mounted displays.
[0007] "Camera" means a combination of hardware and software installed on a Device for capturing still or video images.
[0008] "Real-time video" refers to video data captured using a camera being processed immediately without delay.
[0009] "Server" refers to a remote computer system that receives video data sent from a device and performs analysis and content generation.
[0010] "Image recognition model" refers to an algorithm that uses machine learning or deep learning technology to identify, classify, or recognize objects, scenes, text, etc. in video data.
[0011] "Generative AI" refers to artificial intelligence technology that generates appropriate content such as text, audio, and graphics based on analysis results.
[0012] "Content" means digital data, such as text, audio, graphics, and video, that is provided to users for information or entertainment purposes.
[0013] "Augmented reality" refers to a technology that overlays digital information onto real-world images.
[0014] "User interaction" refers to all actions a user takes to operate a device, including touch operations and voice input in particular. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] This invention provides an augmented reality (AR) system that combines a device's camera, image recognition models, and generative AI to enable users to explore the real world and enjoy new experiences. Specific embodiments of the system are described below.
[0037] System Overview
[0038] The system mainly consists of three elements: the device, the server, and the user. The user captures video in real time using the device's camera, and the video data is sent to the server. The server analyzes the video data and generates appropriate content using generative AI. The generated content is sent to the device, which displays it in augmented reality format. Furthermore, it is possible to present additional information depending on the user's interactions.
[0039] Program processing overview
[0040] Video capture and transmission
[0041] When the application is launched, the device activates the camera, and the user captures real-time video with the device's camera, which is then sent to the server.
[0042] For example, if a user is viewing a painting in a museum, they can take a picture of the painting with their device's camera, and the captured image will be sent to the server in real time.
[0043] Video data analysis
[0044] The server then passes the received video data through an image recognition model, which analyzes the objects and scenes in the data. For example, it analyzes a video of a painting and identifies it as a Renaissance work.
[0045] Content Generation
[0046] The server uses generative AI to generate content based on the analysis results, such as text and audio information about Renaissance paintings and their historical background.
[0047] Sending and Displaying Content
[0048] The server then sends the generated content to the device, which then displays it in an augmented reality format. Specifically, when the user looks at a painting through the camera, historical background and explanatory text are displayed around the painting.
[0049] User Interaction
[0050] Users can manipulate the device and interact with the content. For example, touching a part of a painting can reveal more information about that part, allowing users to gain a deeper understanding.
[0051] Specific examples
[0052] A specific scenario is shown below.
[0053] Suppose a user is exploring a city's historical squares. In this case:
[0054] 1. The user takes a photo of the square using their smartphone camera.
[0055] 2. The device sends the captured video to the server.
[0056] 3. The server analyzes the video data and identifies the square as a famous tourist spot.
[0057] 4. The server generates explanatory text and historical background based on the identification results.
[0058] 5. The device displays the generated content to the user in AR format, and the user can view more information by looking at the square through their camera.
[0059] 6. When the user touches a particular building, additional information is displayed.
[0060] In this way, the present invention is a system that analyzes video data obtained in real time and provides users with realistic information, thereby enriching their real-world experience.
[0061] The processing flow will be explained below.
[0062] Step 1:
[0063] The user launches an application on the device. When the application launches, the camera automatically starts and starts capturing real-time video. The user uses the device's camera to take a picture of, for example, a painting in a museum.
[0064] Step 2:
[0065] It encodes the real-time video data captured by the device and prepares it for transmission to the server. The data is usually compressed and sent to the server over the network.
[0066] Step 3:
[0067] The server decodes the video data received from the device and passes it to the image recognition model, which analyzes the video data and identifies objects and scenes within the video.
[0068] Step 4:
[0069] The server receives the analysis results from the image recognition model and extracts relevant features, for example identifying a painting as being from the Renaissance period.
[0070] Step 5:
[0071] The server uses generative AI to generate appropriate content (text, audio, graphics, etc.) based on the results of the image recognition model analysis, such as generating a commentary about a Renaissance painting.
[0072] Step 6:
[0073] The server sends the generated content to the terminal, where it is compressed and transferred to the terminal via the network.
[0074] Step 7:
[0075] The device decodes the received content and displays it to the user in an augmented reality format, such as by overlaying explanatory text or historical context on top of the camera image.
[0076] Step 8:
[0077] Users can touch specific areas on the screen to obtain more detailed information. The device detects the user's interaction and sends the details to the server.
[0078] Step 9:
[0079] The server generates more detailed content based on the user's interactions and sends it back to the device, generating adaptive information to respond to the user's requests.
[0080] Step 10:
[0081] The device displays the additional content it receives, allowing the user to view more detailed information, such as a detailed description of the part of a painting that was touched.
[0082] This series of steps allows for seamless analysis of video data obtained in real time and the presentation of generated content, effectively providing users with a new experience.
[0083] Example 1
[0084] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0085] When exploring the real world, users have limited access to detailed context and related information, making it difficult to quickly obtain it. This hinders on-site understanding and knowledge. Therefore, there is a need for a system that can provide rich information about real-world scenery and objects in real time, allowing users to experience and understand the world more deeply.
[0086] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0087] In this invention, the server includes means for capturing video in real time using a camera of the device, means for transmitting the captured video data to the information processing device, means for analyzing the video data using an image recognition algorithm in the information processing device, means for generating content using a generative model based on the analysis results, means for transmitting the generated content to the device, means for displaying the received content in an augmented reality format in the device, and means for displaying further information in response to a user's operation, thereby providing detailed information and historical background related to real-world objects and scenes in real time, thereby enriching the user's understanding and experience.
[0088] "Device" refers to a mobile information terminal or electronic device that is directly operated by the user.
[0089] "Photography device" refers to a camera function built into a device or an externally connected camera device.
[0090] "Real-time" means that data acquisition and processing occur simultaneously, minimizing delays.
[0091] "Video data" refers to image and video data captured through a camera.
[0092] "Information processing device" refers to a server or computer that receives data sent from a device and analyzes and processes it.
[0093] "Image recognition algorithm" refers to a computational method or model for identifying objects and scenes in video data and extracting features.
[0094] A "generative model" refers to an artificial intelligence model that generates new content based on analysis results.
[0095] "Content" refers to information such as text, images, audio, and video that is displayed to users.
[0096] "Augmented reality format" refers to the technology and format that displays digital content overlaid on images of the real world.
[0097] "User operations" refers to interactions such as touch operations and voice input that users perform through the device.
[0098] "Further Information" refers to detailed or related information provided in addition to the initial content.
[0099] This invention is a system that combines a device's image capture, information processing, image recognition algorithms, and generative models to provide users exploring the real world with rich information in the form of augmented reality, allowing users to better understand and enjoy their local experiences.
[0100] The main components of the system are as follows:
[0101] device
[0102] The terminal is equipped with a camera, and the camera starts up when the user launches the application. The user captures real-time video data through the camera, which is then recorded on the terminal. The terminal then transmits this video data to an information processing device (server) via the Internet. At this time, the data is compressed before transmission, optimizing communication speed and data volume.
[0103] Information processing device (server)
[0104] The server receives the video data sent from the device and passes it to an image recognition algorithm, which identifies key objects and scenes in the video and extracts features. For example, if a video of an artwork is sent, the server can identify the period and style to which the artwork belongs.
[0105] Generative Model
[0106] The server uses a generative AI model to generate appropriate content based on the results of the image recognition algorithm. This generative AI model generates content according to a predefined prompt. For example, the prompt might say, "Generate text that describes the historical background and characteristics of this painting."
[0107] The server sends the generated content to the terminal, allowing users to obtain information in real time on-site.
[0108] Display and user interaction on the device
[0109] The device receives the content sent from the server and displays it in an augmented reality format, whereby digital information is overlaid on top of real-world scenes and objects as the user views them through the camera.
[0110] As a concrete example, consider the case where a user is viewing a painting in an art museum. When the user takes a picture of the painting with the device's camera, the video is immediately sent to the server. The server analyzes the video data and identifies the painting as a Renaissance work. Based on the analysis results, the server generates text information about the Renaissance painting and its historical background. This information is then sent to the device, and when the user looks at the painting through the camera, the historical background and commentary are displayed around the painting. Furthermore, when the user touches a specific part, additional detailed information about that part is displayed.
[0111] Prompt Sentence Examples
[0112] Below is an example of a prompt sentence to input to the generative AI model.
[0113] "Generate a historical context description of the location shown in this image."
[0114] "Generate text that explains the historical background and characteristics of this painting."
[0115] "Please provide more information about the building in this footage."
[0116] This invention is a system that provides detailed information and historical context about real-world objects and scenes in real time, enriching the user's understanding and experience.
[0117] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0118] Step 1: Capture and send video
[0119] When a user launches an application, the device activates the camera. The user uses the device's camera to capture real-time video data. The input at this point is the video captured by the user, and the output is real-time video data. The device compresses this video data and sends it to a server via the Internet. In concrete terms, for example, if a user takes a photo of a painting in an art museum, the video is immediately sent to the server.
[0120] Step 2: Receiving and analyzing video data
[0121] The server receives compressed video data sent from the device and decompresses it. The input is compressed video data, and the output is decompressed video data. The server then inputs this video data into an image recognition algorithm, which performs data analysis to identify key objects and scenes. Specifically, the server analyzes video containing a painting and identifies the period and style to which the painting belongs.
[0122] Step 3: Content generation
[0123] The server inputs a prompt to the generative AI model based on the analysis results of the image recognition algorithm. This input consists of the image recognition results (e.g., the period and style of the painting) and the prompt (e.g., "Please generate text that explains the historical background and characteristics of this painting"). Based on this, the server uses the generative AI model to generate content (e.g., explanatory text and historical background). Specifically, the generative AI model generates an "explanatory text about a Renaissance painting," and the output is specific text information.
[0124] Step 4: Submitting generated content
[0125] The server sends the generated content to the device. At this point, the input is the content obtained from the generative AI model, and the output is the data sent to the device. Specifically, the server sends the generated text and audio information to the device.
[0126] Step 5: Display content
[0127] The device receives the content sent from the server and displays it in augmented reality format. The input is generated content data from the server, and the output is AR content displayed on the user's device screen. Specifically, the device overlays generated explanatory text and historical background information on top of the image displayed through the camera.
[0128] Step 6: User Interaction
[0129] Users can operate the device and interact with the displayed content. The input for this interaction is the user's touch operation or voice input, and the output is additional information displayed on the device. Specifically, when a user touches a part of the painting, detailed information about that part pops up.
[0130] As described above, the system provides users with detailed information in real time throughout each processing step, enriching their real-world experience.
[0131] (Application example 1)
[0132] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0133] Modern autonomous vehicles require real-time assistance systems to help drivers and passengers reach their destinations more safely and efficiently. However, current navigation systems rely on static map information and do not adequately reflect real-time road conditions and traffic signs. This calls for improved visual information and navigation guidance for drivers. In addition, parking assistance and warnings of dangerous areas are also important functions for autonomous vehicles.
[0134] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0135] In this invention, the server includes means for capturing video in real time using a camera of the device, means for transmitting the captured video data to the server, means for analyzing the video data using an image recognition model in the server, means for generating content using a generative AI based on the analysis results, means for transmitting the generated content to the device, means for displaying the received content in an augmented reality format in the device, means for displaying further information in response to user interaction, means for analyzing the road condition video captured using the camera and generating information about traffic signs and road conditions using a generative AI model, means for displaying the generated traffic signs and navigation information in an augmented reality format on the vehicle's display, and means for displaying the generated navigation information and warning information in an augmented reality format on the vehicle's windshield, thereby enabling dynamic navigation guidance, parking assistance, and danger area warnings based on real-time road conditions and traffic signs.
[0136] "Device Camera" refers to a camera device used to capture video in real time.
[0137] "Video Data" means data containing real-time video information captured by a device's camera.
[0138] A "server" is a computer system that receives video data sent from a device, analyzes it, and generates content.
[0139] An "image recognition model" is a type of machine learning model used to analyze video data and identify its content.
[0140] "Generative AI" is an artificial intelligence technology that automatically generates necessary content based on analysis results.
[0141] "Augmented reality (AR)" is a technology that overlays digital data onto real-world visual information.
[0142] "User interaction" is the act of a user interacting with a system through a device.
[0143] "Traffic signs" are signs on roads that indicate traffic rules, precautions, etc.
[0144] "Navigation information" refers to information that includes guidance and instructions to a destination.
[0145] A "dangerous area" is an area where there is a danger that must be avoided when a vehicle passes through.
[0146] A "vehicle windshield" is a transparent glass portion of a vehicle that provides forward visibility.
[0147] To implement this invention, three elements are required: a device, a server, and a user. The detailed configurations and operations of these elements will be described below.
[0148] System configuration
[0149] device
[0150] The device mainly includes a camera and a display mounted on an autonomous vehicle. The camera captures the road conditions ahead of the vehicle in real time and sends the video data to a server. The display displays the generated AI content received from the server in an augmented reality (AR) format.
[0151] server
[0152] The server is the central processing center for the received video data. It uses the following software to analyze the data and generate content:
[0153] Image recognition model (TensorFlow, Keras, etc.)
[0154] Analyzing traffic signs and road conditions in video data
[0155] Generative AI models (such as Hugging Face's Transformers library)
[0156] Generate appropriate navigation and warning information based on the analysis results
[0157] user
[0158] The user is the driver or passenger of the vehicle. The user visually checks and interacts with the AR content displayed on the vehicle display.
[0159] System Operation
[0160] 1. Video capture and transmission
[0161] The vehicle's camera captures road conditions in real time and transmits the video data to a server.
[0162] 2. Analysis of video data
[0163] The server passes the received video data to an image recognition model, which analyzes traffic signs and road conditions.
[0164] 3. Content Generation
[0165] The server uses generation AI based on the analysis results to generate navigation information and danger warning information.
[0166] 4. Submitting and Displaying Content
[0167] The server sends the generated content to the device and displays it in AR format on the vehicle's display.
[0168] 5. User Interaction
[0169] Users can interact with the content on the display using touch or voice input to obtain more detailed information.
[0170] Specific examples
[0171] Some specific scenarios for autonomous vehicles operating in urban areas include:
[0172] Traffic sign recognition and display: The camera recognizes speed limit signs and displays the speed limit information in AR format on the display.
[0173] Navigate to your destination: Turn-by-turn directions based on real-time video.
[0174] Danger Area Warning: Recognizes obstacles on the road and displays warnings to the user.
[0175] Parking Assist: Recognizes suitable parking spaces and displays parking guidance.
[0176] Prompt Sentence Examples
[0177] "Explicate the traffic rules for near central possible places."
[0178] As described above, the present invention is a system that supports safe and efficient driving by analyzing real-time video data and providing the user with realistic navigation and warning information.
[0179] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0180] Step 1:
[0181] The terminal captures video in real time. This is done by a camera mounted on the car, capturing the road conditions ahead. The input is real-time road video, and the output is the captured video data. This video data is sent to the server for further processing.
[0182] Step 2:
[0183] The device transmits the captured video data to the server. This transmission is done in real time over the network. The input is the captured video data, and the output is the video data received by the server. The video data is ready to be analyzed on the server side.
[0184] Step 3:
[0185] The server inputs the received video data into an image recognition model for analysis. The software used here is TensorFlow and Keras. The input is the video data received by the server, and the output is the analysis results. This analysis result includes information on traffic signs and road conditions contained in the video.
[0186] Step 4:
[0187] The server uses a generative AI model to generate appropriate content based on the analysis results obtained from the image recognition model. This content includes navigation information and warning information. The input is the analysis results, and the output is the generated content. The software used here is the Hugging Face Transformers library.
[0188] Step 5:
[0189] The server sends the generated content to the terminal. This happens in real time over the network. The input is the generated content and the output is the content received by the terminal. The content is ready for subsequent display processing.
[0190] Step 6:
[0191] The device displays the received content in augmented reality (AR) format, overlaying navigation and warning information on the vehicle's windshield or display. The input is the received content, and the output is the information displayed in AR format, making it easier for users to visually confirm the information.
[0192] Step 7:
[0193] The user interacts with the content displayed on the display. This can be done by touch or voice input. The input is the user's action, and the output is a request to display additional information or more detailed information. The additional information is reflected on the vehicle's display.
[0194] In this way, we have explained how the data is processed and the results obtained based on the specific operations at each step, which clarifies the processing flow and details of the entire system.
[0195] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0196] This invention provides an augmented reality (AR) system that combines a device's camera, image recognition model, generative AI, and emotion engine to enable users to explore the real world and enjoy new experiences. In particular, this system has the ability to recognize a user's emotions and generate and display content accordingly. A specific embodiment of the system is described below.
[0197] System Overview
[0198] The system is primarily composed of four elements: the device, the server, the user, and the emotion engine. The user captures video in real time using the device's camera, and the video data is sent to the server. The server analyzes the video data and generates appropriate content using generative AI. The generated content is then sent to the device, which displays it in augmented reality format. The emotion engine then analyzes the user's emotions in real time, and the displayed content is automatically adjusted based on the results.
[0199] Program processing overview
[0200] Video capture and transmission
[0201] When the application is launched, the device activates the camera, and the user captures real-time video with the device's camera, which is then sent to the server.
[0202] For example, if a user is viewing a painting in a museum, they can take a picture of the painting with their device's camera, and the captured image will be sent to the server in real time.
[0203] Video data analysis
[0204] The server then passes the received video data through an image recognition model, which analyzes the objects and scenes in the data. For example, it analyzes a video of a painting and identifies it as a Renaissance work.
[0205] Content Generation
[0206] The server uses generative AI to generate content based on the analysis results, such as text and audio information about Renaissance paintings and their historical background.
[0207] Emotion analysis using an emotion engine
[0208] The device captures the user's facial expressions, voice tone, and body movements, and sends this data to the emotion engine for analysis, which analyzes the user's emotional state and recognizes emotions in real time.
[0209] Sending and Displaying Content
[0210] The server applies the analysis results of the emotion engine to the generated content and generates and adjusts the content to match the user's emotions. For example, if the user is excited, detailed explanations or interactive elements can be added.
[0211] The device decodes the received content and displays it to the user in an augmented reality format, with explanatory text and historical context overlaid on top of the camera image.
[0212] User Interaction
[0213] Users can operate the device and interact with the content. For example, by touching a part of a painting, detailed information about that part can be displayed. The emotion engine analyzes the user's emotions based on their interactions and provides appropriate content.
[0214] Specific examples
[0215] A specific scenario is shown below.
[0216] Suppose a user is exploring a city's historical squares. In this case:
[0217] 1. The user takes a photo of the square using their smartphone camera.
[0218] 2. The device sends the captured video to the server.
[0219] 3. The server analyzes the video data and identifies the square as a famous tourist spot.
[0220] 4. The server generates explanatory text and historical background based on the identification results.
[0221] 5. The device uses an emotion engine to analyze the user's facial expressions and movements and sends the analysis results to the server.
[0222] 6. The server adjusts the content based on the sentiment analysis results.
[0223] 7. The device displays the generated content to the user in AR format, and the user can view more information by looking at the square through their camera.
[0224] 8. When the user touches a particular building, additional information is displayed.
[0225] In this way, by adding an emotion engine, the present invention is a system that can provide content according to the user's emotional state, thereby providing a more personalized experience.
[0226] The processing flow will be explained below.
[0227] Step 1:
[0228] The user launches an application on the device. When the application launches, the camera automatically starts and starts capturing real-time video. The user uses the device's camera to take a picture of, for example, a painting in a museum.
[0229] Step 2:
[0230] It encodes the real-time video data captured by the device and prepares it for transmission to the server. The data is usually compressed and sent to the server over the network.
[0231] Step 3:
[0232] The server decodes the video data received from the device and passes it to the image recognition model, which analyzes the video data and identifies objects and scenes within the video.
[0233] Step 4:
[0234] The server receives the analysis results from the image recognition model and extracts relevant features, for example identifying a painting as being from the Renaissance period.
[0235] Step 5:
[0236] The server uses generative AI to generate appropriate content (text, audio, graphics, etc.) based on the results of the image recognition model analysis, such as generating a commentary about a Renaissance painting.
[0237] Step 6:
[0238] The device captures the user's facial expressions, vocal tone, and body movements, encoding this data and preparing it for analysis by sending it to the emotion engine.
[0239] Step 7:
[0240] The emotion engine analyzes data sent from the device and recognizes the user's emotional state in real time, such as excitement, joy, and surprise.
[0241] Step 8:
[0242] The server applies the analysis results of the emotion engine to the generated content and adjusts the content to match the user's emotions. For example, if the user is excited, it adds detailed explanations or interactive elements.
[0243] Step 9:
[0244] The server sends the adjusted content to the terminal, where it is compressed and transferred to the terminal over the network.
[0245] Step 10:
[0246] The device decodes the received content and displays it to the user in an augmented reality format, such as by overlaying explanatory text or historical context on top of the camera image.
[0247] Step 11:
[0248] Users can touch specific areas on the screen to obtain more detailed information. The device detects the user's interaction and sends the details to the server.
[0249] Step 12:
[0250] The server generates more detailed content based on the user's interactions and sends it back to the device, generating adaptive information to respond to the user's requests.
[0251] Step 13:
[0252] The device displays the additional content it receives, allowing the user to view more detailed information, such as a detailed description of the part of a painting that was touched.
[0253] This series of steps allows for seamless analysis of video data obtained in real time and the presentation of generated content based on the user's emotional state, providing a more personalized experience for the user.
[0254] Example 2
[0255] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0256] Conventional augmented reality (AR) systems provide uniform information without considering the user's emotional state, resulting in a lack of personalized experience for each individual user. Furthermore, they lack the ability to adapt to real-time user interactions and emotional states, resulting in a limited user experience.
[0257] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0258] In this invention, the server includes means for capturing video in real time using a camera of the device, means for transmitting the captured video data to the server, means for analyzing the video data using an image recognition model in the server, means for generating content using generative artificial intelligence, means for transmitting the generated content to the device, means for displaying the received content in an augmented reality format in the device, means for displaying further information in response to user interaction, means for analyzing user emotions using an emotion engine, and means for adjusting content based on the emotion analysis result, thereby enabling the provision of personalized content according to the user's emotional state and realizing an adaptive user experience in real time.
[0259] "Device" is a general term for electronic devices that are directly operated by the user, including cameras, displays, microphones, etc.
[0260] A "camera" is a photographing device that is installed on a device and that captures images in real time.
[0261] "Real time" refers to the responsiveness in which processing or operations are performed immediately.
[0262] "Video" is a general term for visual data captured by a camera.
[0263] "Video data" means a digital representation of captured video.
[0264] A "server" is a central processing unit for processing, storing, and analyzing data, and connects to devices via a network.
[0265] An "image recognition model" is a machine learning algorithm that analyzes video data and identifies and classifies its content.
[0266] "Analysis" is the process of processing given data to clarify its characteristics and content.
[0267] "Generative AI" refers to artificial intelligence that generates new content based on given data and prompts.
[0268] "Content" is a general term for information provided to users, including text, images, audio, etc.
[0269] Augmented reality (AR) is a technology that overlays digital information onto the real world.
[0270] "Display" is the act of visually presenting information on a device's display.
[0271] "User interaction" refers to the operations or inputs that a user makes with content through a device.
[0272] An "emotion engine" is software or algorithm that analyzes a user's facial expressions, vocal tone, and body movements to identify their emotional state.
[0273] "Sentiment analysis" is the process of identifying user emotions based on collected data.
[0274] "Adjust" means changing content or settings to suit specific conditions or criteria.
[0275] This invention provides an augmented reality (AR) system that allows users to explore the real world and enjoy new experiences by combining a device's camera, image recognition model, generative artificial intelligence, and emotion engine. In particular, it has the feature of recognizing the user's emotions and generating and displaying content accordingly.
[0276] System configuration
[0277] The system mainly consists of four elements: device, server, user, and emotion engine.
[0278] 1. Device:
[0279] Camera: The device is equipped with a camera that captures video in real time.
[0280] Display: The device has a display for displaying AR content.
[0281] Microphone: Includes a microphone to capture the user's voice input.
[0282] Applications: Applications for video capture, transmission, display, and emotional data collection are installed.
[0283] 2. Server:
[0284] Storage: Has storage for saving video data and analysis results.
[0285] Image recognition model: An image recognition model (e.g., TensorFlow or OpenCV) used to analyze video data and identify objects and scenes.
[0286] Generative AI: Generative artificial intelligence to generate the required content (e.g., OpenAI's GPT-3 or DALL-E).
[0287] 3. User:
[0288] Users operate the device, capture images with the camera, and interact with the displayed AR content.
[0289] 4. Emotion Engine:
[0290] Software or algorithms that analyze a user's facial expressions, vocal tone, or body movements to identify their emotional state (e.g., Affectiva or Microsoft's Emotion API).
[0291] System Operation
[0292] 1. Video capture and transmission:
[0293] When the application is launched, the device activates the camera, and the user captures real-time video with the device's camera, which then transmits the video data to a server over the Internet.
[0294] Example: If a user is viewing a painting in a museum, they can take a picture of the painting with their camera and send the image to the server.
[0295] 2. Video data analysis:
[0296] The server stores the received video data in storage and passes it to an image recognition model for analysis, which identifies objects and scenes within the video data.
[0297] Example: Analyzing a video of a painting and identifying it as a Renaissance work.
[0298] 3. Content Generation:
[0299] Based on the analysis results, the server generates the necessary content by giving a prompt (e.g., "Please tell me the historical background of this painting") to the generation AI. The generated content is then sent to the device.
[0300] Example: Generating detailed descriptions and historical context for Renaissance artworks.
[0301] 4. Sentiment analysis using emotion engine:
[0302] The device uses a camera and microphone to capture the user's facial expressions, voice tone, and body movements, and sends them to an emotion engine, which analyzes them to determine the user's emotional state.
[0303] Example: Analyzing when a user is excited.
[0304] 5. Submitting and Displaying Content:
[0305] The server adjusts the generated content based on the emotion analysis results and sends it to the device, where it is displayed in AR format.
[0306] Example: Text information is displayed as an overlay on top of the camera image.
[0307] 6. User Interaction:
[0308] Users interact with the displayed content via the device's touchscreen or voice commands, and this interaction data is sent back to the emotion engine to re-analyze the user's emotional state.
[0309] Example: Touching a part of a particular painting will reveal more details.
[0310] Prompt Sentence Examples
[0311] Below are some example prompts for generative AI models:
[0312] 1. "Can you explain the historical background of this square?"
[0313] 2. "Please provide more information about this painting."
[0314] 3. "Generate additional information to display when the user is excited."
[0315] As described above, this system provides personalized content according to the user's emotional state, providing a real-time adaptive user experience.
[0316] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0317] System program processing flow
[0318] Step 1: Capture and send video
[0319] When the application is launched, the device activates the camera and the user captures video in real time. The captured video data is sent to the server via the Internet. The input of this step is the real-time video captured by the device's camera, and the output is the video data sent to the server. In concrete terms, when a user uses their smartphone to take a video of a historic square, the video is sent to the server in real time.
[0320] Step 2: Analyzing the video data
[0321] The server stores the received video data in storage and then passes it to an image recognition model (e.g., TensorFlow or OpenCV) for analysis. This model identifies objects and scenes within the video data. The input to this step is the video data received by the server, and the output is the analysis results for the identified objects and scenes. Specifically, the server analyzes the video data and identifies that the square is a famous tourist spot.
[0322] Step 3: Content generation
[0323] Based on the analysis results, the server generates the required content by providing a prompt to a generation AI (e.g., OpenAI's GPT-3 or DALL-E). The input to this step is the analysis results, and the output is the generated content (e.g., explanatory text or historical background information). Specifically, the server inputs a prompt such as "Explain the historical background of this square," and the generation AI generates a detailed explanatory text about the square.
[0324] Step 4: Emotion analysis using the emotion engine
[0325] The device uses the device's camera and microphone to capture the user's facial expressions, voice tone, and body movements, and sends this data to the emotion engine. The emotion engine analyzes the data and identifies the user's emotional state. The input to this step is data related to the user's facial expressions, voice tone, and body movements, and the output is the analysis result of the user's emotional state. Specifically, the camera captures the user's facial expression when they look at the square, and the emotion engine analyzes it to identify that they are excited.
[0326] Step 5: Submit and display content
[0327] The server adjusts the generated content based on the emotion analysis results and sends it to the device. The input for this step is the emotion analysis results and the generated content, and the output is the adjusted content. The device displays the received content in AR format. The input for this step is the adjusted content, and the output is the AR content displayed on the device's display. Specifically, text information is displayed as an overlay on the camera image.
[0328] Step 6: User Interaction
[0329] The user interacts with the displayed content using the device's touchscreen or voice commands. This interaction data is sent back to the emotion engine, which re-analyzes the user's emotional state. The input of this step is the user's interaction data, and the output is adjusted content based on the re-analyzed emotional state. Specifically, when the user touches a part of a particular painting, more detailed information is displayed, which is again adjusted based on the emotion engine.
[0330] Each step works in tandem to provide the user with a personalized experience that adapts in real time.
[0331] (Application example 2)
[0332] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0333] Shopping in physical stores requires users to gather a lot of information and make choices, which is time-consuming and labor-intensive. It is also difficult to provide personalized product information in real time that reflects the user's mood and emotions. Conventional systems do not suggest content that takes the user's emotions into account, resulting in a limited user experience.
[0334] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0335] In this invention, the server includes means for analyzing video data using an image recognition model, means for generating content using a generative AI, and means for analyzing user emotions, which allows the server to adjust the generated content based on the user's emotions and provide personalized product information in real time.
[0336] A "device" is an electronic device including a camera, which is operated by a user to capture video and perform various data processing.
[0337] "Real-time" refers to the ability to process data almost instantly and provide results immediately.
[0338] "Video Data" means the digital form of visual information captured by a Device's camera.
[0339] "Server" means a computer system that receives, analyzes, and processes data sent from the Device.
[0340] An "image recognition model" is an algorithm that uses computer vision technology to identify objects and scenes in video data.
[0341] "Generative AI" is artificial intelligence that generates new content, such as text or images, based on given data and prompts.
[0342] "Content" refers to information such as text, images, audio, and video that is displayed on a device.
[0343] "Augmented reality" is a technology that displays virtual information overlaid on images of the real world.
[0344] "Emotion analysis" is the process of analyzing a user's emotional state from their facial expressions, voice, etc.
[0345] "User interaction" refers to the reactions and actions that users take toward content by operating a device.
[0346] "Adjustment" refers to appropriately changing the content and presentation of content according to the user's emotions and interactions.
[0347] This invention is an augmented reality (AR) shopping assistant system that combines real-time user emotion analysis and content generation. The system consists of a device, a server, a generative AI model, and an emotion engine.
[0348] Hardware and software used
[0349] 1. Device:
[0350] Smartphone: camera, microphone, display
[0351] Software: Camera API, ARKit or ARCore
[0352] 2. Server:
[0353] Hardware: High-performance computer server
[0354] Software: Image recognition model (TensorFlow, PyTorch), generative AI (GPT-4), sentiment analysis engine (Affectiva SDK)
[0355] System processing overview
[0356] 1. Video capture and transmission:
[0357] Users turn on their smartphone camera and take pictures of products and the interior of the store, and the captured video data is sent from the smartphone to the server.
[0358] 2. Video data analysis:
[0359] The server then passes the received video data through an image recognition model to analyze the objects and scenes in the data, for example, identifying whether an item is an electronic appliance or an item of clothing.
[0360] 3. Content Generation:
[0361] The server generates content using a generative AI based on the image recognition results. The generative AI (GPT-4) generates text information such as detailed information, reviews, and prices about the identified products.
[0362] Example prompt sentence:
[0363] Given an image of a product as input, which product is this?
[0364] Please generate a detailed description for this item.
[0365] 4. Emotion analysis:
[0366] Data such as the user's facial expressions and voice tone are captured using the smartphone's camera and microphone, and then sent to the emotion analysis engine (Affectiva SDK) for analysis, which identifies the user's emotional state (e.g., excitement, satisfaction, etc.).
[0367] 5. Content Adjustment:
[0368] The server then tailors the generated content based on the sentiment analysis results: for example, if the user is excited, it adds details about new or limited edition products.
[0369] 6. Display of Content:
[0370] The server sends the adjusted content to the device, which then displays it to the user in an augmented reality format, overlaying text and images on top of the camera image.
[0371] 7. User Interaction:
[0372] Users can get more detailed information by touching specific areas on the smartphone screen, and voice input is also possible for interaction.
[0373] Specific examples
[0374] Consider a case where a user is looking for a new smartphone in a shopping mall. The user takes a picture of a smartphone on display and sends the video data to a server. The server analyzes the video data and identifies the smartphone. Then, a generative AI generates product information, prices, and reviews, and a sentiment analysis engine analyzes the user's emotions. For example, if the sentiment analysis engine determines that the user is satisfied, the server adds information about premium products and accessories. The tailored content is sent to the device and displayed in AR.
[0375] In this way, users can enjoy a personalized shopping experience that is tailored to their emotional state in real time.
[0376] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0377] Step 1:
[0378] The user activates the smartphone camera to take pictures of products and the store interior. The input is real-time video data captured by the camera, and the output is the video data stored in the smartphone.
[0379] Step 2:
[0380] The terminal sends the captured video data to the server. The input is the video data captured in step 1, and the output is the video data sent to the server via the Internet.
[0381] Step 3:
[0382] The server passes the received video data to an image recognition model, which analyzes the objects and scenes in the data. The input is the video data sent to the server, and the output is information about the recognized objects and scenes. Specifically, an image recognition model (e.g., TensorFlow, PyTorch) is used to identify that a product belongs to a specific category.
[0383] Step 4:
[0384] The server generates content using a generative AI model based on the image recognition results. The input is the image recognition results, and the output is generated text information such as product information, reviews, and prices. Specifically, the server uses a generative AI (e.g., GPT-4) to execute the following prompts:
[0385] Given an image of a product as input, which product is this?
[0386] Please generate a detailed description for this item.
[0387] Step 5:
[0388] The device captures data such as the user's facial expressions and voice tone using a camera and microphone, and sends it to an emotion analysis engine for analysis. The input is the user's real-time facial expressions and voice data, and the output is the user's emotional state (e.g., excitement, satisfaction, etc.). Specifically, emotions are analyzed using the Affectiva SDK.
[0389] Step 6:
[0390] The server adjusts the generated content based on the emotion analysis results. The input is the emotion analysis results and the generated content, and the output is content adjusted to match the user's emotional state. Specifically, if the user is excited, it adds detailed information about new products or limited edition items.
[0391] Step 7:
[0392] The server sends the modified content to the device. The input is the modified content and the output is the content sent to the device over the Internet.
[0393] Step 8:
[0394] The device displays the received content to the user in an augmented reality format. The input is the adjusted content sent in step 7, and the output is the content overlaid on the device screen in an AR format. Specifically, text and images are displayed on top of the camera image using ARKit or ARCore.
[0395] Step 9:
[0396] Users can get more detailed information by touching specific areas on the smartphone screen. The input is the user's touch or voice input, and the output is additional information. Specifically, when a user touches a specific product, more detailed information about that product is displayed.
[0397] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0398] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0399] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0400] [Second embodiment]
[0401] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0402] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0403] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0404] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0405] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0406] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0407] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0408] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0409] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0410] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0411] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0412] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0413] This invention provides an augmented reality (AR) system that combines a device's camera, image recognition models, and generative AI to enable users to explore the real world and enjoy new experiences. Specific embodiments of the system are described below.
[0414] System Overview
[0415] The system mainly consists of three elements: the device, the server, and the user. The user captures video in real time using the device's camera, and the video data is sent to the server. The server analyzes the video data and generates appropriate content using generative AI. The generated content is sent to the device, which displays it in augmented reality format. Furthermore, it is possible to present additional information depending on the user's interactions.
[0416] Program processing overview
[0417] Video capture and transmission
[0418] When the application is launched, the device activates the camera, and the user captures real-time video with the device's camera, which is then sent to the server.
[0419] For example, if a user is viewing a painting in a museum, they can take a picture of the painting with their device's camera, and the captured image will be sent to the server in real time.
[0420] Video data analysis
[0421] The server then passes the received video data through an image recognition model, which analyzes the objects and scenes in the data. For example, it analyzes a video of a painting and identifies it as a Renaissance work.
[0422] Content Generation
[0423] The server uses generative AI to generate content based on the analysis results, such as text and audio information about Renaissance paintings and their historical background.
[0424] Sending and Displaying Content
[0425] The server then sends the generated content to the device, which then displays it in an augmented reality format. Specifically, when the user looks at a painting through the camera, historical background and explanatory text are displayed around the painting.
[0426] User Interaction
[0427] Users can manipulate the device and interact with the content. For example, touching a part of a painting can reveal more information about that part, allowing users to gain a deeper understanding.
[0428] Specific examples
[0429] A specific scenario is shown below.
[0430] Suppose a user is exploring a city's historical squares. In this case:
[0431] 1. The user takes a photo of the square using their smartphone camera.
[0432] 2. The device sends the captured video to the server.
[0433] 3. The server analyzes the video data and identifies the square as a famous tourist spot.
[0434] 4. The server generates explanatory text and historical background based on the identification results.
[0435] 5. The device displays the generated content to the user in AR format, and the user can view more information by looking at the square through their camera.
[0436] 6. When the user touches a particular building, additional information is displayed.
[0437] In this way, the present invention is a system that analyzes video data obtained in real time and provides users with realistic information, thereby enriching their real-world experience.
[0438] The processing flow will be explained below.
[0439] Step 1:
[0440] The user launches an application on the device. When the application launches, the camera automatically starts and starts capturing real-time video. The user uses the device's camera to take a picture of, for example, a painting in a museum.
[0441] Step 2:
[0442] It encodes the real-time video data captured by the device and prepares it for transmission to the server. The data is usually compressed and sent to the server over the network.
[0443] Step 3:
[0444] The server decodes the video data received from the device and passes it to the image recognition model, which analyzes the video data and identifies objects and scenes within the video.
[0445] Step 4:
[0446] The server receives the analysis results from the image recognition model and extracts relevant features, for example identifying a painting as being from the Renaissance period.
[0447] Step 5:
[0448] The server uses generative AI to generate appropriate content (text, audio, graphics, etc.) based on the results of the image recognition model analysis, such as generating a commentary about a Renaissance painting.
[0449] Step 6:
[0450] The server sends the generated content to the terminal, where it is compressed and transferred to the terminal via the network.
[0451] Step 7:
[0452] The device decodes the received content and displays it to the user in an augmented reality format, such as by overlaying explanatory text or historical context on top of the camera image.
[0453] Step 8:
[0454] Users can touch specific areas on the screen to obtain more detailed information. The device detects the user's interaction and sends the details to the server.
[0455] Step 9:
[0456] The server generates more detailed content based on the user's interactions and sends it back to the device, generating adaptive information to respond to the user's requests.
[0457] Step 10:
[0458] The device displays the additional content it receives, allowing the user to view more detailed information, such as a detailed description of the part of a painting that was touched.
[0459] This series of steps allows for seamless analysis of video data obtained in real time and the presentation of generated content, effectively providing users with a new experience.
[0460] Example 1
[0461] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0462] When exploring the real world, users have limited access to detailed context and related information, making it difficult to quickly obtain it. This hinders on-site understanding and knowledge. Therefore, there is a need for a system that can provide rich information about real-world scenery and objects in real time, allowing users to experience and understand the world more deeply.
[0463] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0464] In this invention, the server includes means for capturing video in real time using a camera of the device, means for transmitting the captured video data to the information processing device, means for analyzing the video data using an image recognition algorithm in the information processing device, means for generating content using a generative model based on the analysis results, means for transmitting the generated content to the device, means for displaying the received content in an augmented reality format in the device, and means for displaying further information in response to a user's operation, thereby providing detailed information and historical background related to real-world objects and scenes in real time, thereby enriching the user's understanding and experience.
[0465] "Device" refers to a mobile information terminal or electronic device that is directly operated by the user.
[0466] "Photography device" refers to a camera function built into a device or an externally connected camera device.
[0467] "Real-time" means that data acquisition and processing occur simultaneously, minimizing delays.
[0468] "Video data" refers to image and video data captured through a camera.
[0469] "Information processing device" refers to a server or computer that receives data sent from a device and analyzes and processes it.
[0470] "Image recognition algorithm" refers to a computational method or model for identifying objects and scenes in video data and extracting features.
[0471] A "generative model" refers to an artificial intelligence model that generates new content based on analysis results.
[0472] "Content" refers to information such as text, images, audio, and video that is displayed to users.
[0473] "Augmented reality format" refers to the technology and format that displays digital content overlaid on images of the real world.
[0474] "User operations" refers to interactions such as touch operations and voice input that users perform through the device.
[0475] "Further Information" refers to detailed or related information provided in addition to the initial content.
[0476] This invention is a system that combines a device's image capture, information processing, image recognition algorithms, and generative models to provide users exploring the real world with rich information in the form of augmented reality, allowing users to better understand and enjoy their local experiences.
[0477] The main components of the system are as follows:
[0478] device
[0479] The terminal is equipped with a camera, and the camera starts up when the user launches the application. The user captures real-time video data through the camera, which is then recorded on the terminal. The terminal then transmits this video data to an information processing device (server) via the Internet. At this time, the data is compressed before transmission, optimizing communication speed and data volume.
[0480] Information processing device (server)
[0481] The server receives the video data sent from the device and passes it to an image recognition algorithm, which identifies key objects and scenes in the video and extracts features. For example, if a video of an artwork is sent, the server can identify the period and style to which the artwork belongs.
[0482] Generative Model
[0483] The server uses a generative AI model to generate appropriate content based on the results of the image recognition algorithm. This generative AI model generates content according to a predefined prompt. For example, the prompt might say, "Generate text that describes the historical background and characteristics of this painting."
[0484] The server sends the generated content to the terminal, allowing users to obtain information in real time on-site.
[0485] Display and user interaction on the device
[0486] The device receives the content sent from the server and displays it in an augmented reality format, whereby digital information is overlaid on top of real-world scenes and objects as the user views them through the camera.
[0487] As a concrete example, consider the case where a user is viewing a painting in an art museum. When the user takes a picture of the painting with the device's camera, the video is immediately sent to the server. The server analyzes the video data and identifies the painting as a Renaissance work. Based on the analysis results, the server generates text information about the Renaissance painting and its historical background. This information is then sent to the device, and when the user looks at the painting through the camera, the historical background and commentary are displayed around the painting. Furthermore, when the user touches a specific part, additional detailed information about that part is displayed.
[0488] Prompt Sentence Examples
[0489] Below is an example of a prompt sentence to input to the generative AI model.
[0490] "Generate a historical context description of the location shown in this image."
[0491] "Generate text that explains the historical background and characteristics of this painting."
[0492] "Please provide more information about the building in this footage."
[0493] This invention is a system that provides detailed information and historical context about real-world objects and scenes in real time, enriching the user's understanding and experience.
[0494] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0495] Step 1: Capture and send video
[0496] When a user launches an application, the device activates the camera. The user uses the device's camera to capture real-time video data. The input at this point is the video captured by the user, and the output is real-time video data. The device compresses this video data and sends it to a server via the Internet. In concrete terms, for example, if a user takes a photo of a painting in an art museum, the video is immediately sent to the server.
[0497] Step 2: Receiving and analyzing video data
[0498] The server receives compressed video data sent from the device and decompresses it. The input is compressed video data, and the output is decompressed video data. The server then inputs this video data into an image recognition algorithm, which performs data analysis to identify key objects and scenes. Specifically, the server analyzes video containing a painting and identifies the period and style to which the painting belongs.
[0499] Step 3: Content generation
[0500] The server inputs a prompt to the generative AI model based on the analysis results of the image recognition algorithm. This input consists of the image recognition results (e.g., the period and style of the painting) and the prompt (e.g., "Please generate text that explains the historical background and characteristics of this painting"). Based on this, the server uses the generative AI model to generate content (e.g., explanatory text and historical background). Specifically, the generative AI model generates an "explanatory text about a Renaissance painting," and the output is specific text information.
[0501] Step 4: Submitting generated content
[0502] The server sends the generated content to the device. At this point, the input is the content obtained from the generative AI model, and the output is the data sent to the device. Specifically, the server sends the generated text and audio information to the device.
[0503] Step 5: Display content
[0504] The device receives the content sent from the server and displays it in augmented reality format. The input is generated content data from the server, and the output is AR content displayed on the user's device screen. Specifically, the device overlays generated explanatory text and historical background information on top of the image displayed through the camera.
[0505] Step 6: User Interaction
[0506] Users can operate the device and interact with the displayed content. The input for this interaction is the user's touch operation or voice input, and the output is additional information displayed on the device. Specifically, when a user touches a part of the painting, detailed information about that part pops up.
[0507] As described above, the system provides users with detailed information in real time throughout each processing step, enriching their real-world experience.
[0508] (Application example 1)
[0509] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0510] Modern autonomous vehicles require real-time assistance systems to help drivers and passengers reach their destinations more safely and efficiently. However, current navigation systems rely on static map information and do not adequately reflect real-time road conditions and traffic signs. This calls for improved visual information and navigation guidance for drivers. In addition, parking assistance and warnings of dangerous areas are also important functions for autonomous vehicles.
[0511] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0512] In this invention, the server includes means for capturing video in real time using a camera of the device, means for transmitting the captured video data to the server, means for analyzing the video data using an image recognition model in the server, means for generating content using a generative AI based on the analysis results, means for transmitting the generated content to the device, means for displaying the received content in an augmented reality format in the device, means for displaying further information in response to user interaction, means for analyzing the road condition video captured using the camera and generating information about traffic signs and road conditions using a generative AI model, means for displaying the generated traffic signs and navigation information in an augmented reality format on the vehicle's display, and means for displaying the generated navigation information and warning information in an augmented reality format on the vehicle's windshield, thereby enabling dynamic navigation guidance, parking assistance, and danger area warnings based on real-time road conditions and traffic signs.
[0513] "Device Camera" refers to a camera device used to capture video in real time.
[0514] "Video Data" means data containing real-time video information captured by a device's camera.
[0515] A "server" is a computer system that receives video data sent from a device, analyzes it, and generates content.
[0516] An "image recognition model" is a type of machine learning model used to analyze video data and identify its content.
[0517] "Generative AI" is an artificial intelligence technology that automatically generates necessary content based on analysis results.
[0518] "Augmented reality (AR)" is a technology that overlays digital data onto real-world visual information.
[0519] "User interaction" is the act of a user interacting with a system through a device.
[0520] "Traffic signs" are signs on roads that indicate traffic rules, precautions, etc.
[0521] "Navigation information" refers to information that includes guidance and instructions to a destination.
[0522] A "dangerous area" is an area where there is a danger that must be avoided when a vehicle passes through.
[0523] A "vehicle windshield" is a transparent glass portion of a vehicle that provides forward visibility.
[0524] To implement this invention, three elements are required: a device, a server, and a user. The detailed configurations and operations of these elements will be described below.
[0525] System configuration
[0526] device
[0527] The device mainly includes a camera and a display mounted on an autonomous vehicle. The camera captures the road conditions ahead of the vehicle in real time and sends the video data to a server. The display displays the generated AI content received from the server in an augmented reality (AR) format.
[0528] server
[0529] The server is the central processing center for the received video data. It uses the following software to analyze the data and generate content:
[0530] Image recognition model (TensorFlow, Keras, etc.)
[0531] Analyzing traffic signs and road conditions in video data
[0532] Generative AI models (such as Hugging Face's Transformers library)
[0533] Generate appropriate navigation and warning information based on the analysis results
[0534] user
[0535] The user is the driver or passenger of the vehicle. The user visually checks and interacts with the AR content displayed on the vehicle display.
[0536] System Operation
[0537] 1. Video capture and transmission
[0538] The vehicle's camera captures road conditions in real time and transmits the video data to a server.
[0539] 2. Analysis of video data
[0540] The server passes the received video data to an image recognition model, which analyzes traffic signs and road conditions.
[0541] 3. Content Generation
[0542] The server uses generation AI based on the analysis results to generate navigation information and danger warning information.
[0543] 4. Submitting and Displaying Content
[0544] The server sends the generated content to the device and displays it in AR format on the vehicle's display.
[0545] 5. User Interaction
[0546] Users can interact with the content on the display using touch or voice input to obtain more detailed information.
[0547] Specific examples
[0548] Some specific scenarios for autonomous vehicles operating in urban areas include:
[0549] Traffic sign recognition and display: The camera recognizes speed limit signs and displays the speed limit information in AR format on the display.
[0550] Navigate to your destination: Turn-by-turn directions based on real-time video.
[0551] Danger Area Warning: Recognizes obstacles on the road and displays warnings to the user.
[0552] Parking Assist: Recognizes suitable parking spaces and displays parking guidance.
[0553] Prompt Sentence Examples
[0554] "Explicate the traffic rules for near central possible places."
[0555] As described above, the present invention is a system that supports safe and efficient driving by analyzing real-time video data and providing the user with realistic navigation and warning information.
[0556] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0557] Step 1:
[0558] The terminal captures video in real time. This is done by a camera mounted on the car, capturing the road conditions ahead. The input is real-time road video, and the output is the captured video data. This video data is sent to the server for further processing.
[0559] Step 2:
[0560] The device transmits the captured video data to the server. This transmission is done in real time over the network. The input is the captured video data, and the output is the video data received by the server. The video data is ready to be analyzed on the server side.
[0561] Step 3:
[0562] The server inputs the received video data into an image recognition model for analysis. The software used here is TensorFlow and Keras. The input is the video data received by the server, and the output is the analysis results. This analysis result includes information on traffic signs and road conditions contained in the video.
[0563] Step 4:
[0564] The server uses a generative AI model to generate appropriate content based on the analysis results obtained from the image recognition model. This content includes navigation information and warning information. The input is the analysis results, and the output is the generated content. The software used here is the Hugging Face Transformers library.
[0565] Step 5:
[0566] The server sends the generated content to the terminal. This happens in real time over the network. The input is the generated content and the output is the content received by the terminal. The content is ready for subsequent display processing.
[0567] Step 6:
[0568] The device displays the received content in augmented reality (AR) format, overlaying navigation and warning information on the vehicle's windshield or display. The input is the received content, and the output is the information displayed in AR format, making it easier for users to visually confirm the information.
[0569] Step 7:
[0570] The user interacts with the content displayed on the display. This can be done by touch or voice input. The input is the user's action, and the output is a request to display additional information or more detailed information. The additional information is reflected on the vehicle's display.
[0571] In this way, we have explained how the data is processed and the results obtained based on the specific operations at each step, which clarifies the processing flow and details of the entire system.
[0572] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0573] This invention provides an augmented reality (AR) system that combines a device's camera, image recognition model, generative AI, and emotion engine to enable users to explore the real world and enjoy new experiences. In particular, this system has the ability to recognize a user's emotions and generate and display content accordingly. A specific embodiment of the system is described below.
[0574] System Overview
[0575] The system is primarily composed of four elements: the device, the server, the user, and the emotion engine. The user captures video in real time using the device's camera, and the video data is sent to the server. The server analyzes the video data and generates appropriate content using generative AI. The generated content is then sent to the device, which displays it in augmented reality format. The emotion engine then analyzes the user's emotions in real time, and the displayed content is automatically adjusted based on the results.
[0576] Program processing overview
[0577] Video capture and transmission
[0578] When the application is launched, the device activates the camera, and the user captures real-time video with the device's camera, which is then sent to the server.
[0579] For example, if a user is viewing a painting in a museum, they can take a picture of the painting with their device's camera, and the captured image will be sent to the server in real time.
[0580] Video data analysis
[0581] The server then passes the received video data through an image recognition model, which analyzes the objects and scenes in the data. For example, it analyzes a video of a painting and identifies it as a Renaissance work.
[0582] Content Generation
[0583] The server uses generative AI to generate content based on the analysis results, such as text and audio information about Renaissance paintings and their historical background.
[0584] Emotion analysis using an emotion engine
[0585] The device captures the user's facial expressions, voice tone, and body movements, and sends this data to the emotion engine for analysis, which analyzes the user's emotional state and recognizes emotions in real time.
[0586] Sending and Displaying Content
[0587] The server applies the analysis results of the emotion engine to the generated content and generates and adjusts the content to match the user's emotions. For example, if the user is excited, detailed explanations or interactive elements can be added.
[0588] The device decodes the received content and displays it to the user in an augmented reality format, with explanatory text and historical context overlaid on top of the camera image.
[0589] User Interaction
[0590] Users can operate the device and interact with the content. For example, by touching a part of a painting, detailed information about that part can be displayed. The emotion engine analyzes the user's emotions based on their interactions and provides appropriate content.
[0591] Specific examples
[0592] A specific scenario is shown below.
[0593] Suppose a user is exploring a city's historical squares. In this case:
[0594] 1. The user takes a photo of the square using their smartphone camera.
[0595] 2. The device sends the captured video to the server.
[0596] 3. The server analyzes the video data and identifies the square as a famous tourist spot.
[0597] 4. The server generates explanatory text and historical background based on the identification results.
[0598] 5. The device uses an emotion engine to analyze the user's facial expressions and movements and sends the analysis results to the server.
[0599] 6. The server adjusts the content based on the sentiment analysis results.
[0600] 7. The device displays the generated content to the user in AR format, and the user can view more information by looking at the square through their camera.
[0601] 8. When the user touches a particular building, additional information is displayed.
[0602] In this way, by adding an emotion engine, the present invention is a system that can provide content according to the user's emotional state, thereby providing a more personalized experience.
[0603] The processing flow will be explained below.
[0604] Step 1:
[0605] The user launches an application on the device. When the application launches, the camera automatically starts and starts capturing real-time video. The user uses the device's camera to take a picture of, for example, a painting in a museum.
[0606] Step 2:
[0607] It encodes the real-time video data captured by the device and prepares it for transmission to the server. The data is usually compressed and sent to the server over the network.
[0608] Step 3:
[0609] The server decodes the video data received from the device and passes it to the image recognition model, which analyzes the video data and identifies objects and scenes within the video.
[0610] Step 4:
[0611] The server receives the analysis results from the image recognition model and extracts relevant features, for example identifying a painting as being from the Renaissance period.
[0612] Step 5:
[0613] The server uses generative AI to generate appropriate content (text, audio, graphics, etc.) based on the results of the image recognition model analysis, such as generating a commentary about a Renaissance painting.
[0614] Step 6:
[0615] The device captures the user's facial expressions, vocal tone, and body movements, encoding this data and preparing it for analysis by sending it to the emotion engine.
[0616] Step 7:
[0617] The emotion engine analyzes data sent from the device and recognizes the user's emotional state in real time, such as excitement, joy, and surprise.
[0618] Step 8:
[0619] The server applies the analysis results of the emotion engine to the generated content and adjusts the content to match the user's emotions. For example, if the user is excited, it adds detailed explanations or interactive elements.
[0620] Step 9:
[0621] The server sends the adjusted content to the terminal, where it is compressed and transferred to the terminal over the network.
[0622] Step 10:
[0623] The device decodes the received content and displays it to the user in an augmented reality format, such as by overlaying explanatory text or historical context on top of the camera image.
[0624] Step 11:
[0625] Users can touch specific areas on the screen to obtain more detailed information. The device detects the user's interaction and sends the details to the server.
[0626] Step 12:
[0627] The server generates more detailed content based on the user's interactions and sends it back to the device, generating adaptive information to respond to the user's requests.
[0628] Step 13:
[0629] The device displays the additional content it receives, allowing the user to view more detailed information, such as a detailed description of the part of a painting that was touched.
[0630] This series of steps allows for seamless analysis of video data obtained in real time and the presentation of generated content based on the user's emotional state, providing a more personalized experience for the user.
[0631] Example 2
[0632] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0633] Conventional augmented reality (AR) systems provide uniform information without considering the user's emotional state, resulting in a lack of personalized experience for each individual user. Furthermore, they lack the ability to adapt to real-time user interactions and emotional states, resulting in a limited user experience.
[0634] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0635] In this invention, the server includes means for capturing video in real time using a camera of the device, means for transmitting the captured video data to the server, means for analyzing the video data using an image recognition model in the server, means for generating content using generative artificial intelligence, means for transmitting the generated content to the device, means for displaying the received content in an augmented reality format in the device, means for displaying further information in response to user interaction, means for analyzing user emotions using an emotion engine, and means for adjusting content based on the emotion analysis result, thereby enabling the provision of personalized content according to the user's emotional state and realizing an adaptive user experience in real time.
[0636] "Device" is a general term for electronic devices that are directly operated by the user, including cameras, displays, microphones, etc.
[0637] A "camera" is a photographing device that is installed on a device and that captures images in real time.
[0638] "Real time" refers to the responsiveness in which processing or operations are performed immediately.
[0639] "Video" is a general term for visual data captured by a camera.
[0640] "Video data" means a digital representation of captured video.
[0641] A "server" is a central processing unit for processing, storing, and analyzing data, and connects to devices via a network.
[0642] An "image recognition model" is a machine learning algorithm that analyzes video data and identifies and classifies its content.
[0643] "Analysis" is the process of processing given data to clarify its characteristics and content.
[0644] "Generative AI" refers to artificial intelligence that generates new content based on given data and prompts.
[0645] "Content" is a general term for information provided to users, including text, images, audio, etc.
[0646] Augmented reality (AR) is a technology that overlays digital information onto the real world.
[0647] "Display" is the act of visually presenting information on a device's display.
[0648] "User interaction" refers to the operations or inputs that a user makes with content through a device.
[0649] An "emotion engine" is software or algorithm that analyzes a user's facial expressions, vocal tone, and body movements to identify their emotional state.
[0650] "Sentiment analysis" is the process of identifying user emotions based on collected data.
[0651] "Adjust" means changing content or settings to suit specific conditions or criteria.
[0652] This invention provides an augmented reality (AR) system that allows users to explore the real world and enjoy new experiences by combining a device's camera, image recognition model, generative artificial intelligence, and emotion engine. In particular, it has the feature of recognizing the user's emotions and generating and displaying content accordingly.
[0653] System configuration
[0654] The system mainly consists of four elements: device, server, user, and emotion engine.
[0655] 1. Device:
[0656] Camera: The device is equipped with a camera that captures video in real time.
[0657] Display: The device has a display for displaying AR content.
[0658] Microphone: Includes a microphone to capture the user's voice input.
[0659] Applications: Applications for video capture, transmission, display, and emotional data collection are installed.
[0660] 2. Server:
[0661] Storage: Has storage for saving video data and analysis results.
[0662] Image recognition model: An image recognition model (e.g., TensorFlow or OpenCV) used to analyze video data and identify objects and scenes.
[0663] Generative AI: Generative artificial intelligence to generate the required content (e.g., OpenAI's GPT-3 or DALL-E).
[0664] 3. User:
[0665] Users operate the device, capture images with the camera, and interact with the displayed AR content.
[0666] 4. Emotion Engine:
[0667] Software or algorithms that analyze a user's facial expressions, vocal tone, or body movements to identify their emotional state (e.g., Affectiva or Microsoft's Emotion API).
[0668] System Operation
[0669] 1. Video capture and transmission:
[0670] When the application is launched, the device activates the camera, and the user captures real-time video with the device's camera, which then transmits the video data to a server over the Internet.
[0671] Example: If a user is viewing a painting in a museum, they can take a picture of the painting with their camera and send the image to the server.
[0672] 2. Video data analysis:
[0673] The server stores the received video data in storage and passes it to an image recognition model for analysis, which identifies objects and scenes within the video data.
[0674] Example: Analyzing a video of a painting and identifying it as a Renaissance work.
[0675] 3. Content Generation:
[0676] Based on the analysis results, the server generates the necessary content by giving a prompt (e.g., "Please tell me the historical background of this painting") to the generation AI. The generated content is then sent to the device.
[0677] Example: Generating detailed descriptions and historical context for Renaissance artworks.
[0678] 4. Sentiment analysis using emotion engine:
[0679] The device uses a camera and microphone to capture the user's facial expressions, voice tone, and body movements, and sends them to an emotion engine, which analyzes them to determine the user's emotional state.
[0680] Example: Analyzing when a user is excited.
[0681] 5. Submitting and Displaying Content:
[0682] The server adjusts the generated content based on the emotion analysis results and sends it to the device, where it is displayed in AR format.
[0683] Example: Text information is displayed as an overlay on top of the camera image.
[0684] 6. User Interaction:
[0685] Users interact with the displayed content via the device's touchscreen or voice commands, and this interaction data is sent back to the emotion engine to re-analyze the user's emotional state.
[0686] Example: Touching a part of a particular painting will reveal more details.
[0687] Prompt Sentence Examples
[0688] Below are some example prompts for generative AI models:
[0689] 1. "Can you explain the historical background of this square?"
[0690] 2. "Please provide more information about this painting."
[0691] 3. "Generate additional information to display when the user is excited."
[0692] As described above, this system provides personalized content according to the user's emotional state, providing a real-time adaptive user experience.
[0693] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0694] System program processing flow
[0695] Step 1: Capture and send video
[0696] When the application is launched, the device activates the camera and the user captures video in real time. The captured video data is sent to the server via the Internet. The input of this step is the real-time video captured by the device's camera, and the output is the video data sent to the server. In concrete terms, when a user uses their smartphone to take a video of a historic square, the video is sent to the server in real time.
[0697] Step 2: Analyzing the video data
[0698] The server stores the received video data in storage and then passes it to an image recognition model (e.g., TensorFlow or OpenCV) for analysis. This model identifies objects and scenes within the video data. The input to this step is the video data received by the server, and the output is the analysis results for the identified objects and scenes. Specifically, the server analyzes the video data and identifies that the square is a famous tourist spot.
[0699] Step 3: Content generation
[0700] Based on the analysis results, the server generates the required content by providing a prompt to a generation AI (e.g., OpenAI's GPT-3 or DALL-E). The input to this step is the analysis results, and the output is the generated content (e.g., explanatory text or historical background information). Specifically, the server inputs a prompt such as "Explain the historical background of this square," and the generation AI generates a detailed explanatory text about the square.
[0701] Step 4: Emotion analysis using the emotion engine
[0702] The device uses the device's camera and microphone to capture the user's facial expressions, voice tone, and body movements, and sends this data to the emotion engine. The emotion engine analyzes the data and identifies the user's emotional state. The input to this step is data related to the user's facial expressions, voice tone, and body movements, and the output is the analysis result of the user's emotional state. Specifically, the camera captures the user's facial expression when they look at the square, and the emotion engine analyzes it to identify that they are excited.
[0703] Step 5: Submit and display content
[0704] The server adjusts the generated content based on the emotion analysis results and sends it to the device. The input for this step is the emotion analysis results and the generated content, and the output is the adjusted content. The device displays the received content in AR format. The input for this step is the adjusted content, and the output is the AR content displayed on the device's display. Specifically, text information is displayed as an overlay on the camera image.
[0705] Step 6: User Interaction
[0706] The user interacts with the displayed content using the device's touchscreen or voice commands. This interaction data is sent back to the emotion engine, which re-analyzes the user's emotional state. The input of this step is the user's interaction data, and the output is adjusted content based on the re-analyzed emotional state. Specifically, when the user touches a part of a particular painting, more detailed information is displayed, which is again adjusted based on the emotion engine.
[0707] Each step works in tandem to provide the user with a personalized experience that adapts in real time.
[0708] (Application example 2)
[0709] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0710] Shopping in physical stores requires users to gather a lot of information and make choices, which is time-consuming and labor-intensive. It is also difficult to provide personalized product information in real time that reflects the user's mood and emotions. Conventional systems do not suggest content that takes the user's emotions into account, resulting in a limited user experience.
[0711] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0712] In this invention, the server includes means for analyzing video data using an image recognition model, means for generating content using a generative AI, and means for analyzing user emotions, which allows the server to adjust the generated content based on the user's emotions and provide personalized product information in real time.
[0713] A "device" is an electronic device including a camera, which is operated by a user to capture video and perform various data processing.
[0714] "Real-time" refers to the ability to process data almost instantly and provide results immediately.
[0715] "Video Data" means the digital form of visual information captured by a Device's camera.
[0716] "Server" means a computer system that receives, analyzes, and processes data sent from the Device.
[0717] An "image recognition model" is an algorithm that uses computer vision technology to identify objects and scenes in video data.
[0718] "Generative AI" is artificial intelligence that generates new content, such as text or images, based on given data and prompts.
[0719] "Content" refers to information such as text, images, audio, and video that is displayed on a device.
[0720] "Augmented reality" is a technology that displays virtual information overlaid on images of the real world.
[0721] "Emotion analysis" is the process of analyzing a user's emotional state from their facial expressions, voice, etc.
[0722] "User interaction" refers to the reactions and actions that users take toward content by operating a device.
[0723] "Adjustment" refers to appropriately changing the content and presentation of content according to the user's emotions and interactions.
[0724] This invention is an augmented reality (AR) shopping assistant system that combines real-time user emotion analysis and content generation. The system consists of a device, a server, a generative AI model, and an emotion engine.
[0725] Hardware and software used
[0726] 1. Device:
[0727] Smartphone: camera, microphone, display
[0728] Software: Camera API, ARKit or ARCore
[0729] 2. Server:
[0730] Hardware: High-performance computer server
[0731] Software: Image recognition model (TensorFlow, PyTorch), generative AI (GPT-4), sentiment analysis engine (Affectiva SDK)
[0732] System processing overview
[0733] 1. Video capture and transmission:
[0734] Users turn on their smartphone camera and take pictures of products and the interior of the store, and the captured video data is sent from the smartphone to the server.
[0735] 2. Video data analysis:
[0736] The server then passes the received video data through an image recognition model to analyze the objects and scenes in the data, for example, identifying whether an item is an electronic appliance or an item of clothing.
[0737] 3. Content Generation:
[0738] The server generates content using a generative AI based on the image recognition results. The generative AI (GPT-4) generates text information such as detailed information, reviews, and prices about the identified products.
[0739] Example prompt sentence:
[0740] Given an image of a product as input, which product is this?
[0741] Please generate a detailed description for this item.
[0742] 4. Emotion analysis:
[0743] Data such as the user's facial expressions and voice tone are captured using the smartphone's camera and microphone, and then sent to the emotion analysis engine (Affectiva SDK) for analysis, which identifies the user's emotional state (e.g., excitement, satisfaction, etc.).
[0744] 5. Content Adjustment:
[0745] The server then tailors the generated content based on the sentiment analysis results: for example, if the user is excited, it adds details about new or limited edition products.
[0746] 6. Display of Content:
[0747] The server sends the adjusted content to the device, which then displays it to the user in an augmented reality format, overlaying text and images on top of the camera image.
[0748] 7. User Interaction:
[0749] Users can get more detailed information by touching specific areas on the smartphone screen, and voice input is also possible for interaction.
[0750] Specific examples
[0751] Consider a case where a user is looking for a new smartphone in a shopping mall. The user takes a picture of a smartphone on display and sends the video data to a server. The server analyzes the video data and identifies the smartphone. Then, a generative AI generates product information, prices, and reviews, and a sentiment analysis engine analyzes the user's emotions. For example, if the sentiment analysis engine determines that the user is satisfied, the server adds information about premium products and accessories. The tailored content is sent to the device and displayed in AR.
[0752] In this way, users can enjoy a personalized shopping experience that is tailored to their emotional state in real time.
[0753] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0754] Step 1:
[0755] The user activates the smartphone camera to take pictures of products and the store interior. The input is real-time video data captured by the camera, and the output is the video data stored in the smartphone.
[0756] Step 2:
[0757] The terminal sends the captured video data to the server. The input is the video data captured in step 1, and the output is the video data sent to the server via the Internet.
[0758] Step 3:
[0759] The server passes the received video data to an image recognition model, which analyzes the objects and scenes in the data. The input is the video data sent to the server, and the output is information about the recognized objects and scenes. Specifically, an image recognition model (e.g., TensorFlow, PyTorch) is used to identify that a product belongs to a specific category.
[0760] Step 4:
[0761] The server generates content using a generative AI model based on the image recognition results. The input is the image recognition results, and the output is generated text information such as product information, reviews, and prices. Specifically, the server uses a generative AI (e.g., GPT-4) to execute the following prompts:
[0762] Given an image of a product as input, which product is this?
[0763] Please generate a detailed description for this item.
[0764] Step 5:
[0765] The device captures data such as the user's facial expressions and voice tone using a camera and microphone, and sends it to an emotion analysis engine for analysis. The input is the user's real-time facial expressions and voice data, and the output is the user's emotional state (e.g., excitement, satisfaction, etc.). Specifically, emotions are analyzed using the Affectiva SDK.
[0766] Step 6:
[0767] The server adjusts the generated content based on the emotion analysis results. The input is the emotion analysis results and the generated content, and the output is content adjusted to match the user's emotional state. Specifically, if the user is excited, it adds detailed information about new products or limited edition items.
[0768] Step 7:
[0769] The server sends the modified content to the device. The input is the modified content and the output is the content sent to the device over the Internet.
[0770] Step 8:
[0771] The device displays the received content to the user in an augmented reality format. The input is the adjusted content sent in step 7, and the output is the content overlaid on the device screen in an AR format. Specifically, text and images are displayed on top of the camera image using ARKit or ARCore.
[0772] Step 9:
[0773] Users can get more detailed information by touching specific areas on the smartphone screen. The input is the user's touch or voice input, and the output is additional information. Specifically, when a user touches a specific product, more detailed information about that product is displayed.
[0774] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0775] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0776] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0777] [Third embodiment]
[0778] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0779] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0780] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0781] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0782] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0783] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0784] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0785] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0786] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0787] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0788] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0789] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0790] This invention provides an augmented reality (AR) system that combines a device's camera, image recognition models, and generative AI to enable users to explore the real world and enjoy new experiences. Specific embodiments of the system are described below.
[0791] System Overview
[0792] The system mainly consists of three elements: the device, the server, and the user. The user captures video in real time using the device's camera, and the video data is sent to the server. The server analyzes the video data and generates appropriate content using generative AI. The generated content is sent to the device, which displays it in augmented reality format. Furthermore, it is possible to present additional information depending on the user's interactions.
[0793] Program processing overview
[0794] Video capture and transmission
[0795] When the application is launched, the device activates the camera, and the user captures real-time video with the device's camera, which is then sent to the server.
[0796] For example, if a user is viewing a painting in a museum, they can take a picture of the painting with their device's camera, and the captured image will be sent to the server in real time.
[0797] Video data analysis
[0798] The server then passes the received video data through an image recognition model, which analyzes the objects and scenes in the data. For example, it analyzes a video of a painting and identifies it as a Renaissance work.
[0799] Content Generation
[0800] The server uses generative AI to generate content based on the analysis results, such as text and audio information about Renaissance paintings and their historical background.
[0801] Sending and Displaying Content
[0802] The server then sends the generated content to the device, which then displays it in an augmented reality format. Specifically, when the user looks at a painting through the camera, historical background and explanatory text are displayed around the painting.
[0803] User Interaction
[0804] Users can manipulate the device and interact with the content. For example, touching a part of a painting can reveal more information about that part, allowing users to gain a deeper understanding.
[0805] Specific examples
[0806] A specific scenario is shown below.
[0807] Suppose a user is exploring a city's historical squares. In this case:
[0808] 1. The user takes a photo of the square using their smartphone camera.
[0809] 2. The device sends the captured video to the server.
[0810] 3. The server analyzes the video data and identifies the square as a famous tourist spot.
[0811] 4. The server generates explanatory text and historical background based on the identification results.
[0812] 5. The device displays the generated content to the user in AR format, and the user can view more information by looking at the square through their camera.
[0813] 6. When the user touches a particular building, additional information is displayed.
[0814] In this way, the present invention is a system that analyzes video data obtained in real time and provides users with realistic information, thereby enriching their real-world experience.
[0815] The processing flow will be explained below.
[0816] Step 1:
[0817] The user launches an application on the device. When the application launches, the camera automatically starts and starts capturing real-time video. The user uses the device's camera to take a picture of, for example, a painting in a museum.
[0818] Step 2:
[0819] It encodes the real-time video data captured by the device and prepares it for transmission to the server. The data is usually compressed and sent to the server over the network.
[0820] Step 3:
[0821] The server decodes the video data received from the device and passes it to the image recognition model, which analyzes the video data and identifies objects and scenes within the video.
[0822] Step 4:
[0823] The server receives the analysis results from the image recognition model and extracts relevant features, for example identifying a painting as being from the Renaissance period.
[0824] Step 5:
[0825] The server uses generative AI to generate appropriate content (text, audio, graphics, etc.) based on the results of the image recognition model analysis, such as generating a commentary about a Renaissance painting.
[0826] Step 6:
[0827] The server sends the generated content to the terminal, where it is compressed and transferred to the terminal via the network.
[0828] Step 7:
[0829] The device decodes the received content and displays it to the user in an augmented reality format, such as by overlaying explanatory text or historical context on top of the camera image.
[0830] Step 8:
[0831] Users can touch specific areas on the screen to obtain more detailed information. The device detects the user's interaction and sends the details to the server.
[0832] Step 9:
[0833] The server generates more detailed content based on the user's interactions and sends it back to the device, generating adaptive information to respond to the user's requests.
[0834] Step 10:
[0835] The device displays the additional content it receives, allowing the user to view more detailed information, such as a detailed description of the part of a painting that was touched.
[0836] This series of steps allows for seamless analysis of video data obtained in real time and the presentation of generated content, effectively providing users with a new experience.
[0837] Example 1
[0838] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0839] When exploring the real world, users have limited access to detailed context and related information, making it difficult to quickly obtain it. This hinders on-site understanding and knowledge. Therefore, there is a need for a system that can provide rich information about real-world scenery and objects in real time, allowing users to experience and understand the world more deeply.
[0840] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0841] In this invention, the server includes means for capturing video in real time using a camera of the device, means for transmitting the captured video data to the information processing device, means for analyzing the video data using an image recognition algorithm in the information processing device, means for generating content using a generative model based on the analysis results, means for transmitting the generated content to the device, means for displaying the received content in an augmented reality format in the device, and means for displaying further information in response to a user's operation, thereby providing detailed information and historical background related to real-world objects and scenes in real time, thereby enriching the user's understanding and experience.
[0842] "Device" refers to a mobile information terminal or electronic device that is directly operated by the user.
[0843] "Photography device" refers to a camera function built into a device or an externally connected camera device.
[0844] "Real-time" means that data acquisition and processing occur simultaneously, minimizing delays.
[0845] "Video data" refers to image and video data captured through a camera.
[0846] "Information processing device" refers to a server or computer that receives data sent from a device and analyzes and processes it.
[0847] "Image recognition algorithm" refers to a computational method or model for identifying objects and scenes in video data and extracting features.
[0848] A "generative model" refers to an artificial intelligence model that generates new content based on analysis results.
[0849] "Content" refers to information such as text, images, audio, and video that is displayed to users.
[0850] "Augmented reality format" refers to the technology and format that displays digital content overlaid on images of the real world.
[0851] "User operations" refers to interactions such as touch operations and voice input that users perform through the device.
[0852] "Further Information" refers to detailed or related information provided in addition to the initial content.
[0853] This invention is a system that combines a device's image capture, information processing, image recognition algorithms, and generative models to provide users exploring the real world with rich information in the form of augmented reality, allowing users to better understand and enjoy their local experiences.
[0854] The main components of the system are as follows:
[0855] device
[0856] The terminal is equipped with a camera, and the camera starts up when the user launches the application. The user captures real-time video data through the camera, which is then recorded on the terminal. The terminal then transmits this video data to an information processing device (server) via the Internet. At this time, the data is compressed before transmission, optimizing communication speed and data volume.
[0857] Information processing device (server)
[0858] The server receives the video data sent from the device and passes it to an image recognition algorithm, which identifies key objects and scenes in the video and extracts features. For example, if a video of an artwork is sent, the server can identify the period and style to which the artwork belongs.
[0859] Generative Model
[0860] The server uses a generative AI model to generate appropriate content based on the results of the image recognition algorithm. This generative AI model generates content according to a predefined prompt. For example, the prompt might say, "Generate text that describes the historical background and characteristics of this painting."
[0861] The server sends the generated content to the terminal, allowing users to obtain information in real time on-site.
[0862] Display and user interaction on the device
[0863] The device receives the content sent from the server and displays it in an augmented reality format, whereby digital information is overlaid on top of real-world scenes and objects as the user views them through the camera.
[0864] As a concrete example, consider the case where a user is viewing a painting in an art museum. When the user takes a picture of the painting with the device's camera, the video is immediately sent to the server. The server analyzes the video data and identifies the painting as a Renaissance work. Based on the analysis results, the server generates text information about the Renaissance painting and its historical background. This information is then sent to the device, and when the user looks at the painting through the camera, the historical background and commentary are displayed around the painting. Furthermore, when the user touches a specific part, additional detailed information about that part is displayed.
[0865] Prompt Sentence Examples
[0866] Below is an example of a prompt sentence to input to the generative AI model.
[0867] "Generate a historical context description of the location shown in this image."
[0868] "Generate text that explains the historical background and characteristics of this painting."
[0869] "Please provide more information about the building in this footage."
[0870] This invention is a system that provides detailed information and historical context about real-world objects and scenes in real time, enriching the user's understanding and experience.
[0871] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0872] Step 1: Capture and send video
[0873] When a user launches an application, the device activates the camera. The user uses the device's camera to capture real-time video data. The input at this point is the video captured by the user, and the output is real-time video data. The device compresses this video data and sends it to a server via the Internet. In concrete terms, for example, if a user takes a photo of a painting in an art museum, the video is immediately sent to the server.
[0874] Step 2: Receiving and analyzing video data
[0875] The server receives compressed video data sent from the device and decompresses it. The input is compressed video data, and the output is decompressed video data. The server then inputs this video data into an image recognition algorithm, which performs data analysis to identify key objects and scenes. Specifically, the server analyzes video containing a painting and identifies the period and style to which the painting belongs.
[0876] Step 3: Content generation
[0877] The server inputs a prompt to the generative AI model based on the analysis results of the image recognition algorithm. This input consists of the image recognition results (e.g., the period and style of the painting) and the prompt (e.g., "Please generate text that explains the historical background and characteristics of this painting"). Based on this, the server uses the generative AI model to generate content (e.g., explanatory text and historical background). Specifically, the generative AI model generates an "explanatory text about a Renaissance painting," and the output is specific text information.
[0878] Step 4: Submitting generated content
[0879] The server sends the generated content to the device. At this point, the input is the content obtained from the generative AI model, and the output is the data sent to the device. Specifically, the server sends the generated text and audio information to the device.
[0880] Step 5: Display content
[0881] The device receives the content sent from the server and displays it in augmented reality format. The input is generated content data from the server, and the output is AR content displayed on the user's device screen. Specifically, the device overlays generated explanatory text and historical background information on top of the image displayed through the camera.
[0882] Step 6: User Interaction
[0883] Users can operate the device and interact with the displayed content. The input for this interaction is the user's touch operation or voice input, and the output is additional information displayed on the device. Specifically, when a user touches a part of the painting, detailed information about that part pops up.
[0884] As described above, the system provides users with detailed information in real time throughout each processing step, enriching their real-world experience.
[0885] (Application example 1)
[0886] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0887] Modern autonomous vehicles require real-time assistance systems to help drivers and passengers reach their destinations more safely and efficiently. However, current navigation systems rely on static map information and do not adequately reflect real-time road conditions and traffic signs. This calls for improved visual information and navigation guidance for drivers. In addition, parking assistance and warnings of dangerous areas are also important functions for autonomous vehicles.
[0888] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0889] In this invention, the server includes means for capturing video in real time using a camera of the device, means for transmitting the captured video data to the server, means for analyzing the video data using an image recognition model in the server, means for generating content using a generative AI based on the analysis results, means for transmitting the generated content to the device, means for displaying the received content in an augmented reality format in the device, means for displaying further information in response to user interaction, means for analyzing the road condition video captured using the camera and generating information about traffic signs and road conditions using a generative AI model, means for displaying the generated traffic signs and navigation information in an augmented reality format on the vehicle's display, and means for displaying the generated navigation information and warning information in an augmented reality format on the vehicle's windshield, thereby enabling dynamic navigation guidance, parking assistance, and danger area warnings based on real-time road conditions and traffic signs.
[0890] "Device Camera" refers to a camera device used to capture video in real time.
[0891] "Video Data" means data containing real-time video information captured by a device's camera.
[0892] A "server" is a computer system that receives video data sent from a device, analyzes it, and generates content.
[0893] An "image recognition model" is a type of machine learning model used to analyze video data and identify its content.
[0894] "Generative AI" is an artificial intelligence technology that automatically generates necessary content based on analysis results.
[0895] "Augmented reality (AR)" is a technology that overlays digital data onto real-world visual information.
[0896] "User interaction" is the act of a user interacting with a system through a device.
[0897] "Traffic signs" are signs on roads that indicate traffic rules, precautions, etc.
[0898] "Navigation information" refers to information that includes guidance and instructions to a destination.
[0899] A "dangerous area" is an area where there is a danger that must be avoided when a vehicle passes through.
[0900] A "vehicle windshield" is a transparent glass portion of a vehicle that provides forward visibility.
[0901] To implement this invention, three elements are required: a device, a server, and a user. The detailed configurations and operations of these elements will be described below.
[0902] System configuration
[0903] device
[0904] The device mainly includes a camera and a display mounted on an autonomous vehicle. The camera captures the road conditions ahead of the vehicle in real time and sends the video data to a server. The display displays the generated AI content received from the server in an augmented reality (AR) format.
[0905] server
[0906] The server is the central processing center for the received video data. It uses the following software to analyze the data and generate content:
[0907] Image recognition model (TensorFlow, Keras, etc.)
[0908] Analyzing traffic signs and road conditions in video data
[0909] Generative AI models (such as Hugging Face's Transformers library)
[0910] Generate appropriate navigation and warning information based on the analysis results
[0911] user
[0912] The user is the driver or passenger of the vehicle. The user visually checks and interacts with the AR content displayed on the vehicle display.
[0913] System Operation
[0914] 1. Video capture and transmission
[0915] The vehicle's camera captures road conditions in real time and transmits the video data to a server.
[0916] 2. Analysis of video data
[0917] The server passes the received video data to an image recognition model, which analyzes traffic signs and road conditions.
[0918] 3. Content Generation
[0919] The server uses generation AI based on the analysis results to generate navigation information and danger warning information.
[0920] 4. Submitting and Displaying Content
[0921] The server sends the generated content to the device and displays it in AR format on the vehicle's display.
[0922] 5. User Interaction
[0923] Users can interact with the content on the display using touch or voice input to obtain more detailed information.
[0924] Specific examples
[0925] Some specific scenarios for autonomous vehicles operating in urban areas include:
[0926] Traffic sign recognition and display: The camera recognizes speed limit signs and displays the speed limit information in AR format on the display.
[0927] Navigate to your destination: Turn-by-turn directions based on real-time video.
[0928] Danger Area Warning: Recognizes obstacles on the road and displays warnings to the user.
[0929] Parking Assist: Recognizes suitable parking spaces and displays parking guidance.
[0930] Prompt Sentence Examples
[0931] "Explicate the traffic rules for near central possible places."
[0932] As described above, the present invention is a system that supports safe and efficient driving by analyzing real-time video data and providing the user with realistic navigation and warning information.
[0933] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0934] Step 1:
[0935] The terminal captures video in real time. This is done by a camera mounted on the car, capturing the road conditions ahead. The input is real-time road video, and the output is the captured video data. This video data is sent to the server for further processing.
[0936] Step 2:
[0937] The device transmits the captured video data to the server. This transmission is done in real time over the network. The input is the captured video data, and the output is the video data received by the server. The video data is ready to be analyzed on the server side.
[0938] Step 3:
[0939] The server inputs the received video data into an image recognition model for analysis. The software used here is TensorFlow and Keras. The input is the video data received by the server, and the output is the analysis results. This analysis result includes information on traffic signs and road conditions contained in the video.
[0940] Step 4:
[0941] The server uses a generative AI model to generate appropriate content based on the analysis results obtained from the image recognition model. This content includes navigation information and warning information. The input is the analysis results, and the output is the generated content. The software used here is the Hugging Face Transformers library.
[0942] Step 5:
[0943] The server sends the generated content to the terminal. This happens in real time over the network. The input is the generated content and the output is the content received by the terminal. The content is ready for subsequent display processing.
[0944] Step 6:
[0945] The device displays the received content in augmented reality (AR) format, overlaying navigation and warning information on the vehicle's windshield or display. The input is the received content, and the output is the information displayed in AR format, making it easier for users to visually confirm the information.
[0946] Step 7:
[0947] The user interacts with the content displayed on the display. This can be done by touch or voice input. The input is the user's action, and the output is a request to display additional information or more detailed information. The additional information is reflected on the vehicle's display.
[0948] In this way, we have explained how the data is processed and the results obtained based on the specific operations at each step, which clarifies the processing flow and details of the entire system.
[0949] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0950] This invention provides an augmented reality (AR) system that combines a device's camera, image recognition model, generative AI, and emotion engine to enable users to explore the real world and enjoy new experiences. In particular, this system has the ability to recognize a user's emotions and generate and display content accordingly. A specific embodiment of the system is described below.
[0951] System Overview
[0952] The system is primarily composed of four elements: the device, the server, the user, and the emotion engine. The user captures video in real time using the device's camera, and the video data is sent to the server. The server analyzes the video data and generates appropriate content using generative AI. The generated content is then sent to the device, which displays it in augmented reality format. The emotion engine then analyzes the user's emotions in real time, and the displayed content is automatically adjusted based on the results.
[0953] Program processing overview
[0954] Video capture and transmission
[0955] When the application is launched, the device activates the camera, and the user captures real-time video with the device's camera, which is then sent to the server.
[0956] For example, if a user is viewing a painting in a museum, they can take a picture of the painting with their device's camera, and the captured image will be sent to the server in real time.
[0957] Video data analysis
[0958] The server then passes the received video data through an image recognition model, which analyzes the objects and scenes in the data. For example, it analyzes a video of a painting and identifies it as a Renaissance work.
[0959] Content Generation
[0960] The server uses generative AI to generate content based on the analysis results, such as text and audio information about Renaissance paintings and their historical background.
[0961] Emotion analysis using an emotion engine
[0962] The device captures the user's facial expressions, voice tone, and body movements, and sends this data to the emotion engine for analysis, which analyzes the user's emotional state and recognizes emotions in real time.
[0963] Sending and Displaying Content
[0964] The server applies the analysis results of the emotion engine to the generated content and generates and adjusts the content to match the user's emotions. For example, if the user is excited, detailed explanations or interactive elements can be added.
[0965] The device decodes the received content and displays it to the user in an augmented reality format, with explanatory text and historical context overlaid on top of the camera image.
[0966] User Interaction
[0967] Users can operate the device and interact with the content. For example, by touching a part of a painting, detailed information about that part can be displayed. The emotion engine analyzes the user's emotions based on their interactions and provides appropriate content.
[0968] Specific examples
[0969] A specific scenario is shown below.
[0970] Suppose a user is exploring a city's historical squares. In this case:
[0971] 1. The user takes a photo of the square using their smartphone camera.
[0972] 2. The device sends the captured video to the server.
[0973] 3. The server analyzes the video data and identifies the square as a famous tourist spot.
[0974] 4. The server generates explanatory text and historical background based on the identification results.
[0975] 5. The device uses an emotion engine to analyze the user's facial expressions and movements and sends the analysis results to the server.
[0976] 6. The server adjusts the content based on the sentiment analysis results.
[0977] 7. The device displays the generated content to the user in AR format, and the user can view more information by looking at the square through their camera.
[0978] 8. When the user touches a particular building, additional information is displayed.
[0979] In this way, by adding an emotion engine, the present invention is a system that can provide content according to the user's emotional state, thereby providing a more personalized experience.
[0980] The processing flow will be explained below.
[0981] Step 1:
[0982] The user launches an application on the device. When the application launches, the camera automatically starts and starts capturing real-time video. The user uses the device's camera to take a picture of, for example, a painting in a museum.
[0983] Step 2:
[0984] It encodes the real-time video data captured by the device and prepares it for transmission to the server. The data is usually compressed and sent to the server over the network.
[0985] Step 3:
[0986] The server decodes the video data received from the device and passes it to the image recognition model, which analyzes the video data and identifies objects and scenes within the video.
[0987] Step 4:
[0988] The server receives the analysis results from the image recognition model and extracts relevant features, for example identifying a painting as being from the Renaissance period.
[0989] Step 5:
[0990] The server uses generative AI to generate appropriate content (text, audio, graphics, etc.) based on the results of the image recognition model analysis, such as generating a commentary about a Renaissance painting.
[0991] Step 6:
[0992] The device captures the user's facial expressions, vocal tone, and body movements, encoding this data and preparing it for analysis by sending it to the emotion engine.
[0993] Step 7:
[0994] The emotion engine analyzes data sent from the device and recognizes the user's emotional state in real time, such as excitement, joy, and surprise.
[0995] Step 8:
[0996] The server applies the analysis results of the emotion engine to the generated content and adjusts the content to match the user's emotions. For example, if the user is excited, it adds detailed explanations or interactive elements.
[0997] Step 9:
[0998] The server sends the adjusted content to the terminal, where it is compressed and transferred to the terminal over the network.
[0999] Step 10:
[1000] The device decodes the received content and displays it to the user in an augmented reality format, such as by overlaying explanatory text or historical context on top of the camera image.
[1001] Step 11:
[1002] Users can touch specific areas on the screen to obtain more detailed information. The device detects the user's interaction and sends the details to the server.
[1003] Step 12:
[1004] The server generates more detailed content based on the user's interactions and sends it back to the device, generating adaptive information to respond to the user's requests.
[1005] Step 13:
[1006] The device displays the additional content it receives, allowing the user to view more detailed information, such as a detailed description of the part of a painting that was touched.
[1007] This series of steps allows for seamless analysis of video data obtained in real time and the presentation of generated content based on the user's emotional state, providing a more personalized experience for the user.
[1008] Example 2
[1009] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1010] Conventional augmented reality (AR) systems provide uniform information without considering the user's emotional state, resulting in a lack of personalized experience for each individual user. Furthermore, they lack the ability to adapt to real-time user interactions and emotional states, resulting in a limited user experience.
[1011] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1012] In this invention, the server includes means for capturing video in real time using a camera of the device, means for transmitting the captured video data to the server, means for analyzing the video data using an image recognition model in the server, means for generating content using generative artificial intelligence, means for transmitting the generated content to the device, means for displaying the received content in an augmented reality format in the device, means for displaying further information in response to user interaction, means for analyzing user emotions using an emotion engine, and means for adjusting content based on the emotion analysis result, thereby enabling the provision of personalized content according to the user's emotional state and realizing an adaptive user experience in real time.
[1013] "Device" is a general term for electronic devices that are directly operated by the user, including cameras, displays, microphones, etc.
[1014] A "camera" is a photographing device that is installed on a device and that captures images in real time.
[1015] "Real time" refers to the responsiveness in which processing or operations are performed immediately.
[1016] "Video" is a general term for visual data captured by a camera.
[1017] "Video data" means a digital representation of captured video.
[1018] A "server" is a central processing unit for processing, storing, and analyzing data, and connects to devices via a network.
[1019] An "image recognition model" is a machine learning algorithm that analyzes video data and identifies and classifies its content.
[1020] "Analysis" is the process of processing given data to clarify its characteristics and content.
[1021] "Generative AI" refers to artificial intelligence that generates new content based on given data and prompts.
[1022] "Content" is a general term for information provided to users, including text, images, audio, etc.
[1023] Augmented reality (AR) is a technology that overlays digital information onto the real world.
[1024] "Display" is the act of visually presenting information on a device's display.
[1025] "User interaction" refers to the operations or inputs that a user makes with content through a device.
[1026] An "emotion engine" is software or algorithm that analyzes a user's facial expressions, vocal tone, and body movements to identify their emotional state.
[1027] "Sentiment analysis" is the process of identifying user emotions based on collected data.
[1028] "Adjust" means changing content or settings to suit specific conditions or criteria.
[1029] This invention provides an augmented reality (AR) system that allows users to explore the real world and enjoy new experiences by combining a device's camera, image recognition model, generative artificial intelligence, and emotion engine. In particular, it has the feature of recognizing the user's emotions and generating and displaying content accordingly.
[1030] System configuration
[1031] The system mainly consists of four elements: device, server, user, and emotion engine.
[1032] 1. Device:
[1033] Camera: The device is equipped with a camera that captures video in real time.
[1034] Display: The device has a display for displaying AR content.
[1035] Microphone: Includes a microphone to capture the user's voice input.
[1036] Applications: Applications for video capture, transmission, display, and emotional data collection are installed.
[1037] 2. Server:
[1038] Storage: Has storage for saving video data and analysis results.
[1039] Image recognition model: An image recognition model (e.g., TensorFlow or OpenCV) used to analyze video data and identify objects and scenes.
[1040] Generative AI: Generative artificial intelligence to generate the required content (e.g., OpenAI's GPT-3 or DALL-E).
[1041] 3. User:
[1042] Users operate the device, capture images with the camera, and interact with the displayed AR content.
[1043] 4. Emotion Engine:
[1044] Software or algorithms that analyze a user's facial expressions, vocal tone, or body movements to identify their emotional state (e.g., Affectiva or Microsoft's Emotion API).
[1045] System Operation
[1046] 1. Video capture and transmission:
[1047] When the application is launched, the device activates the camera, and the user captures real-time video with the device's camera, which then transmits the video data to a server over the Internet.
[1048] Example: If a user is viewing a painting in a museum, they can take a picture of the painting with their camera and send the image to the server.
[1049] 2. Video data analysis:
[1050] The server stores the received video data in storage and passes it to an image recognition model for analysis, which identifies objects and scenes within the video data.
[1051] Example: Analyzing a video of a painting and identifying it as a Renaissance work.
[1052] 3. Content Generation:
[1053] Based on the analysis results, the server generates the necessary content by giving a prompt (e.g., "Please tell me the historical background of this painting") to the generation AI. The generated content is then sent to the device.
[1054] Example: Generating detailed descriptions and historical context for Renaissance artworks.
[1055] 4. Sentiment analysis using emotion engine:
[1056] The device uses a camera and microphone to capture the user's facial expressions, voice tone, and body movements, and sends them to an emotion engine, which analyzes them to determine the user's emotional state.
[1057] Example: Analyzing when a user is excited.
[1058] 5. Submitting and Displaying Content:
[1059] The server adjusts the generated content based on the emotion analysis results and sends it to the device, where it is displayed in AR format.
[1060] Example: Text information is displayed as an overlay on top of the camera image.
[1061] 6. User Interaction:
[1062] Users interact with the displayed content via the device's touchscreen or voice commands, and this interaction data is sent back to the emotion engine to re-analyze the user's emotional state.
[1063] Example: Touching a part of a particular painting will reveal more details.
[1064] Prompt Sentence Examples
[1065] Below are some example prompts for generative AI models:
[1066] 1. "Can you explain the historical background of this square?"
[1067] 2. "Please provide more information about this painting."
[1068] 3. "Generate additional information to display when the user is excited."
[1069] As described above, this system provides personalized content according to the user's emotional state, providing a real-time adaptive user experience.
[1070] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1071] System program processing flow
[1072] Step 1: Capture and send video
[1073] When the application is launched, the device activates the camera and the user captures video in real time. The captured video data is sent to the server via the Internet. The input of this step is the real-time video captured by the device's camera, and the output is the video data sent to the server. In concrete terms, when a user uses their smartphone to take a video of a historic square, the video is sent to the server in real time.
[1074] Step 2: Analyzing the video data
[1075] The server stores the received video data in storage and then passes it to an image recognition model (e.g., TensorFlow or OpenCV) for analysis. This model identifies objects and scenes within the video data. The input to this step is the video data received by the server, and the output is the analysis results for the identified objects and scenes. Specifically, the server analyzes the video data and identifies that the square is a famous tourist spot.
[1076] Step 3: Content generation
[1077] Based on the analysis results, the server generates the required content by providing a prompt to a generation AI (e.g., OpenAI's GPT-3 or DALL-E). The input to this step is the analysis results, and the output is the generated content (e.g., explanatory text or historical background information). Specifically, the server inputs a prompt such as "Explain the historical background of this square," and the generation AI generates a detailed explanatory text about the square.
[1078] Step 4: Emotion analysis using the emotion engine
[1079] The device uses the device's camera and microphone to capture the user's facial expressions, voice tone, and body movements, and sends this data to the emotion engine. The emotion engine analyzes the data and identifies the user's emotional state. The input to this step is data related to the user's facial expressions, voice tone, and body movements, and the output is the analysis result of the user's emotional state. Specifically, the camera captures the user's facial expression when they look at the square, and the emotion engine analyzes it to identify that they are excited.
[1080] Step 5: Submit and display content
[1081] The server adjusts the generated content based on the emotion analysis results and sends it to the device. The input for this step is the emotion analysis results and the generated content, and the output is the adjusted content. The device displays the received content in AR format. The input for this step is the adjusted content, and the output is the AR content displayed on the device's display. Specifically, text information is displayed as an overlay on the camera image.
[1082] Step 6: User Interaction
[1083] The user interacts with the displayed content using the device's touchscreen or voice commands. This interaction data is sent back to the emotion engine, which re-analyzes the user's emotional state. The input of this step is the user's interaction data, and the output is adjusted content based on the re-analyzed emotional state. Specifically, when the user touches a part of a particular painting, more detailed information is displayed, which is again adjusted based on the emotion engine.
[1084] Each step works in tandem to provide the user with a personalized experience that adapts in real time.
[1085] (Application example 2)
[1086] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1087] Shopping in physical stores requires users to gather a lot of information and make choices, which is time-consuming and labor-intensive. It is also difficult to provide personalized product information in real time that reflects the user's mood and emotions. Conventional systems do not suggest content that takes the user's emotions into account, resulting in a limited user experience.
[1088] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1089] In this invention, the server includes means for analyzing video data using an image recognition model, means for generating content using a generative AI, and means for analyzing user emotions, which allows the server to adjust the generated content based on the user's emotions and provide personalized product information in real time.
[1090] A "device" is an electronic device including a camera, which is operated by a user to capture video and perform various data processing.
[1091] "Real-time" refers to the ability to process data almost instantly and provide results immediately.
[1092] "Video Data" means the digital form of visual information captured by a Device's camera.
[1093] "Server" means a computer system that receives, analyzes, and processes data sent from the Device.
[1094] An "image recognition model" is an algorithm that uses computer vision technology to identify objects and scenes in video data.
[1095] "Generative AI" is artificial intelligence that generates new content, such as text or images, based on given data and prompts.
[1096] "Content" refers to information such as text, images, audio, and video that is displayed on a device.
[1097] "Augmented reality" is a technology that displays virtual information overlaid on images of the real world.
[1098] "Emotion analysis" is the process of analyzing a user's emotional state from their facial expressions, voice, etc.
[1099] "User interaction" refers to the reactions and actions that users take toward content by operating a device.
[1100] "Adjustment" refers to appropriately changing the content and presentation of content according to the user's emotions and interactions.
[1101] This invention is an augmented reality (AR) shopping assistant system that combines real-time user emotion analysis and content generation. The system consists of a device, a server, a generative AI model, and an emotion engine.
[1102] Hardware and software used
[1103] 1. Device:
[1104] Smartphone: camera, microphone, display
[1105] Software: Camera API, ARKit or ARCore
[1106] 2. Server:
[1107] Hardware: High-performance computer server
[1108] Software: Image recognition model (TensorFlow, PyTorch), generative AI (GPT-4), sentiment analysis engine (Affectiva SDK)
[1109] System processing overview
[1110] 1. Video capture and transmission:
[1111] Users turn on their smartphone camera and take pictures of products and the interior of the store, and the captured video data is sent from the smartphone to the server.
[1112] 2. Video data analysis:
[1113] The server then passes the received video data through an image recognition model to analyze the objects and scenes in the data, for example, identifying whether an item is an electronic appliance or an item of clothing.
[1114] 3. Content Generation:
[1115] The server generates content using a generative AI based on the image recognition results. The generative AI (GPT-4) generates text information such as detailed information, reviews, and prices about the identified products.
[1116] Example prompt sentence:
[1117] Given an image of a product as input, which product is this?
[1118] Please generate a detailed description for this item.
[1119] 4. Emotion analysis:
[1120] Data such as the user's facial expressions and voice tone are captured using the smartphone's camera and microphone, and then sent to the emotion analysis engine (Affectiva SDK) for analysis, which identifies the user's emotional state (e.g., excitement, satisfaction, etc.).
[1121] 5. Content Adjustment:
[1122] The server then tailors the generated content based on the sentiment analysis results: for example, if the user is excited, it adds details about new or limited edition products.
[1123] 6. Display of Content:
[1124] The server sends the adjusted content to the device, which then displays it to the user in an augmented reality format, overlaying text and images on top of the camera image.
[1125] 7. User Interaction:
[1126] Users can get more detailed information by touching specific areas on the smartphone screen, and voice input is also possible for interaction.
[1127] Specific examples
[1128] Consider a case where a user is looking for a new smartphone in a shopping mall. The user takes a picture of a smartphone on display and sends the video data to a server. The server analyzes the video data and identifies the smartphone. Then, a generative AI generates product information, prices, and reviews, and a sentiment analysis engine analyzes the user's emotions. For example, if the sentiment analysis engine determines that the user is satisfied, the server adds information about premium products and accessories. The tailored content is sent to the device and displayed in AR.
[1129] In this way, users can enjoy a personalized shopping experience that is tailored to their emotional state in real time.
[1130] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1131] Step 1:
[1132] The user activates the smartphone camera to take pictures of products and the store interior. The input is real-time video data captured by the camera, and the output is the video data stored in the smartphone.
[1133] Step 2:
[1134] The terminal sends the captured video data to the server. The input is the video data captured in step 1, and the output is the video data sent to the server via the Internet.
[1135] Step 3:
[1136] The server passes the received video data to an image recognition model, which analyzes the objects and scenes in the data. The input is the video data sent to the server, and the output is information about the recognized objects and scenes. Specifically, an image recognition model (e.g., TensorFlow, PyTorch) is used to identify that a product belongs to a specific category.
[1137] Step 4:
[1138] The server generates content using a generative AI model based on the image recognition results. The input is the image recognition results, and the output is generated text information such as product information, reviews, and prices. Specifically, the server uses a generative AI (e.g., GPT-4) to execute the following prompts:
[1139] Given an image of a product as input, which product is this?
[1140] Please generate a detailed description for this item.
[1141] Step 5:
[1142] The device captures data such as the user's facial expressions and voice tone using a camera and microphone, and sends it to an emotion analysis engine for analysis. The input is the user's real-time facial expressions and voice data, and the output is the user's emotional state (e.g., excitement, satisfaction, etc.). Specifically, emotions are analyzed using the Affectiva SDK.
[1143] Step 6:
[1144] The server adjusts the generated content based on the emotion analysis results. The input is the emotion analysis results and the generated content, and the output is content adjusted to match the user's emotional state. Specifically, if the user is excited, it adds detailed information about new products or limited edition items.
[1145] Step 7:
[1146] The server sends the modified content to the device. The input is the modified content and the output is the content sent to the device over the Internet.
[1147] Step 8:
[1148] The device displays the received content to the user in an augmented reality format. The input is the adjusted content sent in step 7, and the output is the content overlaid on the device screen in an AR format. Specifically, text and images are displayed on top of the camera image using ARKit or ARCore.
[1149] Step 9:
[1150] Users can get more detailed information by touching specific areas on the smartphone screen. The input is the user's touch or voice input, and the output is additional information. Specifically, when a user touches a specific product, more detailed information about that product is displayed.
[1151] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1152] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1153] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1154] [Fourth embodiment]
[1155] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1156] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1157] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1158] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1159] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1160] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1161] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1162] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1163] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1164] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1165] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1166] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1167] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1168] This invention provides an augmented reality (AR) system that combines a device's camera, image recognition models, and generative AI to enable users to explore the real world and enjoy new experiences. Specific embodiments of the system are described below.
[1169] System Overview
[1170] The system mainly consists of three elements: the device, the server, and the user. The user captures video in real time using the device's camera, and the video data is sent to the server. The server analyzes the video data and generates appropriate content using generative AI. The generated content is sent to the device, which displays it in augmented reality format. Furthermore, it is possible to present additional information depending on the user's interactions.
[1171] Program processing overview
[1172] Video capture and transmission
[1173] When the application is launched, the device activates the camera, and the user captures real-time video with the device's camera, which is then sent to the server.
[1174] For example, if a user is viewing a painting in a museum, they can take a picture of the painting with their device's camera, and the captured image will be sent to the server in real time.
[1175] Video data analysis
[1176] The server then passes the received video data through an image recognition model, which analyzes the objects and scenes in the data. For example, it analyzes a video of a painting and identifies it as a Renaissance work.
[1177] Content Generation
[1178] The server uses generative AI to generate content based on the analysis results, such as text and audio information about Renaissance paintings and their historical background.
[1179] Sending and Displaying Content
[1180] The server then sends the generated content to the device, which then displays it in an augmented reality format. Specifically, when the user looks at a painting through the camera, historical background and explanatory text are displayed around the painting.
[1181] User Interaction
[1182] Users can manipulate the device and interact with the content. For example, touching a part of a painting can reveal more information about that part, allowing users to gain a deeper understanding.
[1183] Specific examples
[1184] A specific scenario is shown below.
[1185] Suppose a user is exploring a city's historical squares. In this case:
[1186] 1. The user takes a photo of the square using their smartphone camera.
[1187] 2. The device sends the captured video to the server.
[1188] 3. The server analyzes the video data and identifies the square as a famous tourist spot.
[1189] 4. The server generates explanatory text and historical background based on the identification results.
[1190] 5. The device displays the generated content to the user in AR format, and the user can view more information by looking at the square through their camera.
[1191] 6. When the user touches a particular building, additional information is displayed.
[1192] In this way, the present invention is a system that analyzes video data obtained in real time and provides users with realistic information, thereby enriching their real-world experience.
[1193] The processing flow will be explained below.
[1194] Step 1:
[1195] The user launches an application on the device. When the application launches, the camera automatically starts and starts capturing real-time video. The user uses the device's camera to take a picture of, for example, a painting in a museum.
[1196] Step 2:
[1197] It encodes the real-time video data captured by the device and prepares it for transmission to the server. The data is usually compressed and sent to the server over the network.
[1198] Step 3:
[1199] The server decodes the video data received from the device and passes it to the image recognition model, which analyzes the video data and identifies objects and scenes within the video.
[1200] Step 4:
[1201] The server receives the analysis results from the image recognition model and extracts relevant features, for example identifying a painting as being from the Renaissance period.
[1202] Step 5:
[1203] The server uses generative AI to generate appropriate content (text, audio, graphics, etc.) based on the results of the image recognition model analysis, such as generating a commentary about a Renaissance painting.
[1204] Step 6:
[1205] The server sends the generated content to the terminal, where it is compressed and transferred to the terminal via the network.
[1206] Step 7:
[1207] The device decodes the received content and displays it to the user in an augmented reality format, such as by overlaying explanatory text or historical context on top of the camera image.
[1208] Step 8:
[1209] Users can touch specific areas on the screen to obtain more detailed information. The device detects the user's interaction and sends the details to the server.
[1210] Step 9:
[1211] The server generates more detailed content based on the user's interactions and sends it back to the device, generating adaptive information to respond to the user's requests.
[1212] Step 10:
[1213] The device displays the additional content it receives, allowing the user to view more detailed information, such as a detailed description of the part of a painting that was touched.
[1214] This series of steps allows for seamless analysis of video data obtained in real time and the presentation of generated content, effectively providing users with a new experience.
[1215] Example 1
[1216] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1217] When exploring the real world, users have limited access to detailed context and related information, making it difficult to quickly obtain it. This hinders on-site understanding and knowledge. Therefore, there is a need for a system that can provide rich information about real-world scenery and objects in real time, allowing users to experience and understand the world more deeply.
[1218] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1219] In this invention, the server includes means for capturing video in real time using a camera of the device, means for transmitting the captured video data to the information processing device, means for analyzing the video data using an image recognition algorithm in the information processing device, means for generating content using a generative model based on the analysis results, means for transmitting the generated content to the device, means for displaying the received content in an augmented reality format in the device, and means for displaying further information in response to a user's operation, thereby providing detailed information and historical background related to real-world objects and scenes in real time, thereby enriching the user's understanding and experience.
[1220] "Device" refers to a mobile information terminal or electronic device that is directly operated by the user.
[1221] "Photography device" refers to a camera function built into a device or an externally connected camera device.
[1222] "Real-time" means that data acquisition and processing occur simultaneously, minimizing delays.
[1223] "Video data" refers to image and video data captured through a camera.
[1224] "Information processing device" refers to a server or computer that receives data sent from a device and analyzes and processes it.
[1225] "Image recognition algorithm" refers to a computational method or model for identifying objects and scenes in video data and extracting features.
[1226] A "generative model" refers to an artificial intelligence model that generates new content based on analysis results.
[1227] "Content" refers to information such as text, images, audio, and video that is displayed to users.
[1228] "Augmented reality format" refers to the technology and format that displays digital content overlaid on images of the real world.
[1229] "User operations" refers to interactions such as touch operations and voice input that users perform through the device.
[1230] "Further Information" refers to detailed or related information provided in addition to the initial content.
[1231] This invention is a system that combines a device's image capture, information processing, image recognition algorithms, and generative models to provide users exploring the real world with rich information in the form of augmented reality, allowing users to better understand and enjoy their local experiences.
[1232] The main components of the system are as follows:
[1233] device
[1234] The terminal is equipped with a camera, and the camera starts up when the user launches the application. The user captures real-time video data through the camera, which is then recorded on the terminal. The terminal then transmits this video data to an information processing device (server) via the Internet. At this time, the data is compressed before transmission, optimizing communication speed and data volume.
[1235] Information processing device (server)
[1236] The server receives the video data sent from the device and passes it to an image recognition algorithm, which identifies key objects and scenes in the video and extracts features. For example, if a video of an artwork is sent, the server can identify the period and style to which the artwork belongs.
[1237] Generative Model
[1238] The server uses a generative AI model to generate appropriate content based on the results of the image recognition algorithm. This generative AI model generates content according to a predefined prompt. For example, the prompt might say, "Generate text that describes the historical background and characteristics of this painting."
[1239] The server sends the generated content to the terminal, allowing users to obtain information in real time on-site.
[1240] Display and user interaction on the device
[1241] The device receives the content sent from the server and displays it in an augmented reality format, whereby digital information is overlaid on top of real-world scenes and objects as the user views them through the camera.
[1242] As a concrete example, consider the case where a user is viewing a painting in an art museum. When the user takes a picture of the painting with the device's camera, the video is immediately sent to the server. The server analyzes the video data and identifies the painting as a Renaissance work. Based on the analysis results, the server generates text information about the Renaissance painting and its historical background. This information is then sent to the device, and when the user looks at the painting through the camera, the historical background and commentary are displayed around the painting. Furthermore, when the user touches a specific part, additional detailed information about that part is displayed.
[1243] Prompt Sentence Examples
[1244] Below is an example of a prompt sentence to input to the generative AI model.
[1245] "Generate a historical context description of the location shown in this image."
[1246] "Generate text that explains the historical background and characteristics of this painting."
[1247] "Please provide more information about the building in this footage."
[1248] This invention is a system that provides detailed information and historical context about real-world objects and scenes in real time, enriching the user's understanding and experience.
[1249] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1250] Step 1: Capture and send video
[1251] When a user launches an application, the device activates the camera. The user uses the device's camera to capture real-time video data. The input at this point is the video captured by the user, and the output is real-time video data. The device compresses this video data and sends it to a server via the Internet. In concrete terms, for example, if a user takes a photo of a painting in an art museum, the video is immediately sent to the server.
[1252] Step 2: Receiving and analyzing video data
[1253] The server receives compressed video data sent from the device and decompresses it. The input is compressed video data, and the output is decompressed video data. The server then inputs this video data into an image recognition algorithm, which performs data analysis to identify key objects and scenes. Specifically, the server analyzes video containing a painting and identifies the period and style to which the painting belongs.
[1254] Step 3: Content generation
[1255] The server inputs a prompt to the generative AI model based on the analysis results of the image recognition algorithm. This input consists of the image recognition results (e.g., the period and style of the painting) and the prompt (e.g., "Please generate text that explains the historical background and characteristics of this painting"). Based on this, the server uses the generative AI model to generate content (e.g., explanatory text and historical background). Specifically, the generative AI model generates an "explanatory text about a Renaissance painting," and the output is specific text information.
[1256] Step 4: Submitting generated content
[1257] The server sends the generated content to the device. At this point, the input is the content obtained from the generative AI model, and the output is the data sent to the device. Specifically, the server sends the generated text and audio information to the device.
[1258] Step 5: Display content
[1259] The device receives the content sent from the server and displays it in augmented reality format. The input is generated content data from the server, and the output is AR content displayed on the user's device screen. Specifically, the device overlays generated explanatory text and historical background information on top of the image displayed through the camera.
[1260] Step 6: User Interaction
[1261] Users can operate the device and interact with the displayed content. The input for this interaction is the user's touch operation or voice input, and the output is additional information displayed on the device. Specifically, when a user touches a part of the painting, detailed information about that part pops up.
[1262] As described above, the system provides users with detailed information in real time throughout each processing step, enriching their real-world experience.
[1263] (Application example 1)
[1264] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1265] Modern autonomous vehicles require real-time assistance systems to help drivers and passengers reach their destinations more safely and efficiently. However, current navigation systems rely on static map information and do not adequately reflect real-time road conditions and traffic signs. This calls for improved visual information and navigation guidance for drivers. In addition, parking assistance and warnings of dangerous areas are also important functions for autonomous vehicles.
[1266] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1267] In this invention, the server includes means for capturing video in real time using a camera of the device, means for transmitting the captured video data to the server, means for analyzing the video data using an image recognition model in the server, means for generating content using a generative AI based on the analysis results, means for transmitting the generated content to the device, means for displaying the received content in an augmented reality format in the device, means for displaying further information in response to user interaction, means for analyzing the road condition video captured using the camera and generating information about traffic signs and road conditions using a generative AI model, means for displaying the generated traffic signs and navigation information in an augmented reality format on the vehicle's display, and means for displaying the generated navigation information and warning information in an augmented reality format on the vehicle's windshield, thereby enabling dynamic navigation guidance, parking assistance, and danger area warnings based on real-time road conditions and traffic signs.
[1268] "Device Camera" refers to a camera device used to capture video in real time.
[1269] "Video Data" means data containing real-time video information captured by a device's camera.
[1270] A "server" is a computer system that receives video data sent from a device, analyzes it, and generates content.
[1271] An "image recognition model" is a type of machine learning model used to analyze video data and identify its content.
[1272] "Generative AI" is an artificial intelligence technology that automatically generates necessary content based on analysis results.
[1273] "Augmented reality (AR)" is a technology that overlays digital data onto real-world visual information.
[1274] "User interaction" is the act of a user interacting with a system through a device.
[1275] "Traffic signs" are signs on roads that indicate traffic rules, precautions, etc.
[1276] "Navigation information" refers to information that includes guidance and instructions to a destination.
[1277] A "dangerous area" is an area where there is a danger that must be avoided when a vehicle passes through.
[1278] A "vehicle windshield" is a transparent glass portion of a vehicle that provides forward visibility.
[1279] To implement this invention, three elements are required: a device, a server, and a user. The detailed configurations and operations of these elements will be described below.
[1280] System configuration
[1281] device
[1282] The device mainly includes a camera and a display mounted on an autonomous vehicle. The camera captures the road conditions ahead of the vehicle in real time and sends the video data to a server. The display displays the generated AI content received from the server in an augmented reality (AR) format.
[1283] server
[1284] The server is the central processing center for the received video data. It uses the following software to analyze the data and generate content:
[1285] Image recognition model (TensorFlow, Keras, etc.)
[1286] Analyzing traffic signs and road conditions in video data
[1287] Generative AI models (such as Hugging Face's Transformers library)
[1288] Generate appropriate navigation and warning information based on the analysis results
[1289] user
[1290] The user is the driver or passenger of the vehicle. The user visually checks and interacts with the AR content displayed on the vehicle display.
[1291] System Operation
[1292] 1. Video capture and transmission
[1293] The vehicle's camera captures road conditions in real time and transmits the video data to a server.
[1294] 2. Analysis of video data
[1295] The server passes the received video data to an image recognition model, which analyzes traffic signs and road conditions.
[1296] 3. Content Generation
[1297] The server uses generation AI based on the analysis results to generate navigation information and danger warning information.
[1298] 4. Submitting and Displaying Content
[1299] The server sends the generated content to the device and displays it in AR format on the vehicle's display.
[1300] 5. User Interaction
[1301] Users can interact with the content on the display using touch or voice input to obtain more detailed information.
[1302] Specific examples
[1303] Some specific scenarios for autonomous vehicles operating in urban areas include:
[1304] Traffic sign recognition and display: The camera recognizes speed limit signs and displays the speed limit information in AR format on the display.
[1305] Navigate to your destination: Turn-by-turn directions based on real-time video.
[1306] Danger Area Warning: Recognizes obstacles on the road and displays warnings to the user.
[1307] Parking Assist: Recognizes suitable parking spaces and displays parking guidance.
[1308] Prompt Sentence Examples
[1309] "Explicate the traffic rules for near central possible places."
[1310] As described above, the present invention is a system that supports safe and efficient driving by analyzing real-time video data and providing the user with realistic navigation and warning information.
[1311] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1312] Step 1:
[1313] The terminal captures video in real time. This is done by a camera mounted on the car, capturing the road conditions ahead. The input is real-time road video, and the output is the captured video data. This video data is sent to the server for further processing.
[1314] Step 2:
[1315] The device transmits the captured video data to the server. This transmission is done in real time over the network. The input is the captured video data, and the output is the video data received by the server. The video data is ready to be analyzed on the server side.
[1316] Step 3:
[1317] The server inputs the received video data into an image recognition model for analysis. The software used here is TensorFlow and Keras. The input is the video data received by the server, and the output is the analysis results. This analysis result includes information on traffic signs and road conditions contained in the video.
[1318] Step 4:
[1319] The server uses a generative AI model to generate appropriate content based on the analysis results obtained from the image recognition model. This content includes navigation information and warning information. The input is the analysis results, and the output is the generated content. The software used here is the Hugging Face Transformers library.
[1320] Step 5:
[1321] The server sends the generated content to the terminal. This happens in real time over the network. The input is the generated content and the output is the content received by the terminal. The content is ready for subsequent display processing.
[1322] Step 6:
[1323] The device displays the received content in augmented reality (AR) format, overlaying navigation and warning information on the vehicle's windshield or display. The input is the received content, and the output is the information displayed in AR format, making it easier for users to visually confirm the information.
[1324] Step 7:
[1325] The user interacts with the content displayed on the display. This can be done by touch or voice input. The input is the user's action, and the output is a request to display additional information or more detailed information. The additional information is reflected on the vehicle's display.
[1326] In this way, we have explained how the data is processed and the results obtained based on the specific operations at each step, which clarifies the processing flow and details of the entire system.
[1327] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1328] This invention provides an augmented reality (AR) system that combines a device's camera, image recognition model, generative AI, and emotion engine to enable users to explore the real world and enjoy new experiences. In particular, this system has the ability to recognize a user's emotions and generate and display content accordingly. A specific embodiment of the system is described below.
[1329] System Overview
[1330] The system is primarily composed of four elements: the device, the server, the user, and the emotion engine. The user captures video in real time using the device's camera, and the video data is sent to the server. The server analyzes the video data and generates appropriate content using generative AI. The generated content is then sent to the device, which displays it in augmented reality format. The emotion engine then analyzes the user's emotions in real time, and the displayed content is automatically adjusted based on the results.
[1331] Program processing overview
[1332] Video capture and transmission
[1333] When the application is launched, the device activates the camera, and the user captures real-time video with the device's camera, which is then sent to the server.
[1334] For example, if a user is viewing a painting in a museum, they can take a picture of the painting with their device's camera, and the captured image will be sent to the server in real time.
[1335] Video data analysis
[1336] The server then passes the received video data through an image recognition model, which analyzes the objects and scenes in the data. For example, it analyzes a video of a painting and identifies it as a Renaissance work.
[1337] Content Generation
[1338] The server uses generative AI to generate content based on the analysis results, such as text and audio information about Renaissance paintings and their historical background.
[1339] Emotion analysis using an emotion engine
[1340] The device captures the user's facial expressions, voice tone, and body movements, and sends this data to the emotion engine for analysis, which analyzes the user's emotional state and recognizes emotions in real time.
[1341] Sending and Displaying Content
[1342] The server applies the analysis results of the emotion engine to the generated content and generates and adjusts the content to match the user's emotions. For example, if the user is excited, detailed explanations or interactive elements can be added.
[1343] The device decodes the received content and displays it to the user in an augmented reality format, with explanatory text and historical context overlaid on top of the camera image.
[1344] User Interaction
[1345] Users can operate the device and interact with the content. For example, by touching a part of a painting, detailed information about that part can be displayed. The emotion engine analyzes the user's emotions based on their interactions and provides appropriate content.
[1346] Specific examples
[1347] A specific scenario is shown below.
[1348] Suppose a user is exploring a city's historical squares. In this case:
[1349] 1. The user takes a photo of the square using their smartphone camera.
[1350] 2. The device sends the captured video to the server.
[1351] 3. The server analyzes the video data and identifies the square as a famous tourist spot.
[1352] 4. The server generates explanatory text and historical background based on the identification results.
[1353] 5. The device uses an emotion engine to analyze the user's facial expressions and movements and sends the analysis results to the server.
[1354] 6. The server adjusts the content based on the sentiment analysis results.
[1355] 7. The device displays the generated content to the user in AR format, and the user can view more information by looking at the square through their camera.
[1356] 8. When the user touches a particular building, additional information is displayed.
[1357] In this way, by adding an emotion engine, the present invention is a system that can provide content according to the user's emotional state, thereby providing a more personalized experience.
[1358] The processing flow will be explained below.
[1359] Step 1:
[1360] The user launches an application on the device. When the application launches, the camera automatically starts and starts capturing real-time video. The user uses the device's camera to take a picture of, for example, a painting in a museum.
[1361] Step 2:
[1362] It encodes the real-time video data captured by the device and prepares it for transmission to the server. The data is usually compressed and sent to the server over the network.
[1363] Step 3:
[1364] The server decodes the video data received from the device and passes it to the image recognition model, which analyzes the video data and identifies objects and scenes within the video.
[1365] Step 4:
[1366] The server receives the analysis results from the image recognition model and extracts relevant features, for example identifying a painting as being from the Renaissance period.
[1367] Step 5:
[1368] The server uses generative AI to generate appropriate content (text, audio, graphics, etc.) based on the results of the image recognition model analysis, such as generating a commentary about a Renaissance painting.
[1369] Step 6:
[1370] The device captures the user's facial expressions, vocal tone, and body movements, encoding this data and preparing it for analysis by sending it to the emotion engine.
[1371] Step 7:
[1372] The emotion engine analyzes data sent from the device and recognizes the user's emotional state in real time, such as excitement, joy, and surprise.
[1373] Step 8:
[1374] The server applies the analysis results of the emotion engine to the generated content and adjusts the content to match the user's emotions. For example, if the user is excited, it adds detailed explanations or interactive elements.
[1375] Step 9:
[1376] The server sends the adjusted content to the terminal, where it is compressed and transferred to the terminal over the network.
[1377] Step 10:
[1378] The device decodes the received content and displays it to the user in an augmented reality format, such as by overlaying explanatory text or historical context on top of the camera image.
[1379] Step 11:
[1380] Users can touch specific areas on the screen to obtain more detailed information. The device detects the user's interaction and sends the details to the server.
[1381] Step 12:
[1382] The server generates more detailed content based on the user's interactions and sends it back to the device, generating adaptive information to respond to the user's requests.
[1383] Step 13:
[1384] The device displays the additional content it receives, allowing the user to view more detailed information, such as a detailed description of the part of a painting that was touched.
[1385] This series of steps allows for seamless analysis of video data obtained in real time and the presentation of generated content based on the user's emotional state, providing a more personalized experience for the user.
[1386] Example 2
[1387] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1388] Conventional augmented reality (AR) systems provide uniform information without considering the user's emotional state, resulting in a lack of personalized experience for each individual user. Furthermore, they lack the ability to adapt to real-time user interactions and emotional states, resulting in a limited user experience.
[1389] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1390] In this invention, the server includes means for capturing video in real time using a camera of the device, means for transmitting the captured video data to the server, means for analyzing the video data using an image recognition model in the server, means for generating content using generative artificial intelligence, means for transmitting the generated content to the device, means for displaying the received content in an augmented reality format in the device, means for displaying further information in response to user interaction, means for analyzing user emotions using an emotion engine, and means for adjusting content based on the emotion analysis result, thereby enabling the provision of personalized content according to the user's emotional state and realizing an adaptive user experience in real time.
[1391] "Device" is a general term for electronic devices that are directly operated by the user, including cameras, displays, microphones, etc.
[1392] A "camera" is a photographing device that is installed on a device and that captures images in real time.
[1393] "Real time" refers to the responsiveness in which processing or operations are performed immediately.
[1394] "Video" is a general term for visual data captured by a camera.
[1395] "Video data" means a digital representation of captured video.
[1396] A "server" is a central processing unit for processing, storing, and analyzing data, and connects to devices via a network.
[1397] An "image recognition model" is a machine learning algorithm that analyzes video data and identifies and classifies its content.
[1398] "Analysis" is the process of processing given data to clarify its characteristics and content.
[1399] "Generative AI" refers to artificial intelligence that generates new content based on given data and prompts.
[1400] "Content" is a general term for information provided to users, including text, images, audio, etc.
[1401] Augmented reality (AR) is a technology that overlays digital information onto the real world.
[1402] "Display" is the act of visually presenting information on a device's display.
[1403] "User interaction" refers to the operations or inputs that a user makes with content through a device.
[1404] An "emotion engine" is software or algorithm that analyzes a user's facial expressions, vocal tone, and body movements to identify their emotional state.
[1405] "Sentiment analysis" is the process of identifying user emotions based on collected data.
[1406] "Adjust" means changing content or settings to suit specific conditions or criteria.
[1407] This invention provides an augmented reality (AR) system that allows users to explore the real world and enjoy new experiences by combining a device's camera, image recognition model, generative artificial intelligence, and emotion engine. In particular, it has the feature of recognizing the user's emotions and generating and displaying content accordingly.
[1408] System configuration
[1409] The system mainly consists of four elements: device, server, user, and emotion engine.
[1410] 1. Device:
[1411] Camera: The device is equipped with a camera that captures video in real time.
[1412] Display: The device has a display for displaying AR content.
[1413] Microphone: Includes a microphone to capture the user's voice input.
[1414] Applications: Applications for video capture, transmission, display, and emotional data collection are installed.
[1415] 2. Server:
[1416] Storage: Has storage for saving video data and analysis results.
[1417] Image recognition model: An image recognition model (e.g., TensorFlow or OpenCV) used to analyze video data and identify objects and scenes.
[1418] Generative AI: Generative artificial intelligence to generate the required content (e.g., OpenAI's GPT-3 or DALL-E).
[1419] 3. User:
[1420] Users operate the device, capture images with the camera, and interact with the displayed AR content.
[1421] 4. Emotion Engine:
[1422] Software or algorithms that analyze a user's facial expressions, vocal tone, or body movements to identify their emotional state (e.g., Affectiva or Microsoft's Emotion API).
[1423] System Operation
[1424] 1. Video capture and transmission:
[1425] When the application is launched, the device activates the camera, and the user captures real-time video with the device's camera, which then transmits the video data to a server over the Internet.
[1426] Example: If a user is viewing a painting in a museum, they can take a picture of the painting with their camera and send the image to the server.
[1427] 2. Video data analysis:
[1428] The server stores the received video data in storage and passes it to an image recognition model for analysis, which identifies objects and scenes within the video data.
[1429] Example: Analyzing a video of a painting and identifying it as a Renaissance work.
[1430] 3. Content Generation:
[1431] Based on the analysis results, the server generates the necessary content by giving a prompt (e.g., "Please tell me the historical background of this painting") to the generation AI. The generated content is then sent to the device.
[1432] Example: Generating detailed descriptions and historical context for Renaissance artworks.
[1433] 4. Sentiment analysis using emotion engine:
[1434] The device uses a camera and microphone to capture the user's facial expressions, voice tone, and body movements, and sends them to an emotion engine, which analyzes them to determine the user's emotional state.
[1435] Example: Analyzing when a user is excited.
[1436] 5. Submitting and Displaying Content:
[1437] The server adjusts the generated content based on the emotion analysis results and sends it to the device, where it is displayed in AR format.
[1438] Example: Text information is displayed as an overlay on top of the camera image.
[1439] 6. User Interaction:
[1440] Users interact with the displayed content via the device's touchscreen or voice commands, and this interaction data is sent back to the emotion engine to re-analyze the user's emotional state.
[1441] Example: Touching a part of a particular painting will reveal more details.
[1442] Prompt Sentence Examples
[1443] Below are some example prompts for generative AI models:
[1444] 1. "Can you explain the historical background of this square?"
[1445] 2. "Please provide more information about this painting."
[1446] 3. "Generate additional information to display when the user is excited."
[1447] As described above, this system provides personalized content according to the user's emotional state, providing a real-time adaptive user experience.
[1448] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1449] System program processing flow
[1450] Step 1: Capture and send video
[1451] When the application is launched, the device activates the camera and the user captures video in real time. The captured video data is sent to the server via the Internet. The input of this step is the real-time video captured by the device's camera, and the output is the video data sent to the server. In concrete terms, when a user uses their smartphone to take a video of a historic square, the video is sent to the server in real time.
[1452] Step 2: Analyzing the video data
[1453] The server stores the received video data in storage and then passes it to an image recognition model (e.g., TensorFlow or OpenCV) for analysis. This model identifies objects and scenes within the video data. The input to this step is the video data received by the server, and the output is the analysis results for the identified objects and scenes. Specifically, the server analyzes the video data and identifies that the square is a famous tourist spot.
[1454] Step 3: Content generation
[1455] Based on the analysis results, the server generates the required content by providing a prompt to a generation AI (e.g., OpenAI's GPT-3 or DALL-E). The input to this step is the analysis results, and the output is the generated content (e.g., explanatory text or historical background information). Specifically, the server inputs a prompt such as "Explain the historical background of this square," and the generation AI generates a detailed explanatory text about the square.
[1456] Step 4: Emotion analysis using the emotion engine
[1457] The device uses the device's camera and microphone to capture the user's facial expressions, voice tone, and body movements, and sends this data to the emotion engine. The emotion engine analyzes the data and identifies the user's emotional state. The input to this step is data related to the user's facial expressions, voice tone, and body movements, and the output is the analysis result of the user's emotional state. Specifically, the camera captures the user's facial expression when they look at the square, and the emotion engine analyzes it to identify that they are excited.
[1458] Step 5: Submit and display content
[1459] The server adjusts the generated content based on the emotion analysis results and sends it to the device. The input for this step is the emotion analysis results and the generated content, and the output is the adjusted content. The device displays the received content in AR format. The input for this step is the adjusted content, and the output is the AR content displayed on the device's display. Specifically, text information is displayed as an overlay on the camera image.
[1460] Step 6: User Interaction
[1461] The user interacts with the displayed content using the device's touchscreen or voice commands. This interaction data is sent back to the emotion engine, which re-analyzes the user's emotional state. The input of this step is the user's interaction data, and the output is adjusted content based on the re-analyzed emotional state. Specifically, when the user touches a part of a particular painting, more detailed information is displayed, which is again adjusted based on the emotion engine.
[1462] Each step works in tandem to provide the user with a personalized experience that adapts in real time.
[1463] (Application example 2)
[1464] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1465] Shopping in physical stores requires users to gather a lot of information and make choices, which is time-consuming and labor-intensive. It is also difficult to provide personalized product information in real time that reflects the user's mood and emotions. Conventional systems do not suggest content that takes the user's emotions into account, resulting in a limited user experience.
[1466] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1467] In this invention, the server includes means for analyzing video data using an image recognition model, means for generating content using a generative AI, and means for analyzing user emotions, which allows the server to adjust the generated content based on the user's emotions and provide personalized product information in real time.
[1468] A "device" is an electronic device including a camera, which is operated by a user to capture video and perform various data processing.
[1469] "Real-time" refers to the ability to process data almost instantly and provide results immediately.
[1470] "Video Data" means the digital form of visual information captured by a Device's camera.
[1471] "Server" means a computer system that receives, analyzes, and processes data sent from the Device.
[1472] An "image recognition model" is an algorithm that uses computer vision technology to identify objects and scenes in video data.
[1473] "Generative AI" is artificial intelligence that generates new content, such as text or images, based on given data and prompts.
[1474] "Content" refers to information such as text, images, audio, and video that is displayed on a device.
[1475] "Augmented reality" is a technology that displays virtual information overlaid on images of the real world.
[1476] "Emotion analysis" is the process of analyzing a user's emotional state from their facial expressions, voice, etc.
[1477] "User interaction" refers to the reactions and actions that users take toward content by operating a device.
[1478] "Adjustment" refers to appropriately changing the content and presentation of content according to the user's emotions and interactions.
[1479] This invention is an augmented reality (AR) shopping assistant system that combines real-time user emotion analysis and content generation. The system consists of a device, a server, a generative AI model, and an emotion engine.
[1480] Hardware and software used
[1481] 1. Device:
[1482] Smartphone: camera, microphone, display
[1483] Software: Camera API, ARKit or ARCore
[1484] 2. Server:
[1485] Hardware: High-performance computer server
[1486] Software: Image recognition model (TensorFlow, PyTorch), generative AI (GPT-4), sentiment analysis engine (Affectiva SDK)
[1487] System processing overview
[1488] 1. Video capture and transmission:
[1489] Users turn on their smartphone camera and take pictures of products and the interior of the store, and the captured video data is sent from the smartphone to the server.
[1490] 2. Video data analysis:
[1491] The server then passes the received video data through an image recognition model to analyze the objects and scenes in the data, for example, identifying whether an item is an electronic appliance or an item of clothing.
[1492] 3. Content Generation:
[1493] The server generates content using a generative AI based on the image recognition results. The generative AI (GPT-4) generates text information such as detailed information, reviews, and prices about the identified products.
[1494] Example prompt sentence:
[1495] Given an image of a product as input, which product is this?
[1496] Please generate a detailed description for this item.
[1497] 4. Emotion analysis:
[1498] Data such as the user's facial expressions and voice tone are captured using the smartphone's camera and microphone, and then sent to the emotion analysis engine (Affectiva SDK) for analysis, which identifies the user's emotional state (e.g., excitement, satisfaction, etc.).
[1499] 5. Content Adjustment:
[1500] The server then tailors the generated content based on the sentiment analysis results: for example, if the user is excited, it adds details about new or limited edition products.
[1501] 6. Display of Content:
[1502] The server sends the adjusted content to the device, which then displays it to the user in an augmented reality format, overlaying text and images on top of the camera image.
[1503] 7. User Interaction:
[1504] Users can get more detailed information by touching specific areas on the smartphone screen, and voice input is also possible for interaction.
[1505] Specific examples
[1506] Consider a case where a user is looking for a new smartphone in a shopping mall. The user takes a picture of a smartphone on display and sends the video data to a server. The server analyzes the video data and identifies the smartphone. Then, a generative AI generates product information, prices, and reviews, and a sentiment analysis engine analyzes the user's emotions. For example, if the sentiment analysis engine determines that the user is satisfied, the server adds information about premium products and accessories. The tailored content is sent to the device and displayed in AR.
[1507] In this way, users can enjoy a personalized shopping experience that is tailored to their emotional state in real time.
[1508] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1509] Step 1:
[1510] The user activates the smartphone camera to take pictures of products and the store interior. The input is real-time video data captured by the camera, and the output is the video data stored in the smartphone.
[1511] Step 2:
[1512] The terminal sends the captured video data to the server. The input is the video data captured in step 1, and the output is the video data sent to the server via the Internet.
[1513] Step 3:
[1514] The server passes the received video data to an image recognition model, which analyzes the objects and scenes in the data. The input is the video data sent to the server, and the output is information about the recognized objects and scenes. Specifically, an image recognition model (e.g., TensorFlow, PyTorch) is used to identify that a product belongs to a specific category.
[1515] Step 4:
[1516] The server generates content using a generative AI model based on the image recognition results. The input is the image recognition results, and the output is generated text information such as product information, reviews, and prices. Specifically, the server uses a generative AI (e.g., GPT-4) to execute the following prompts:
[1517] Given an image of a product as input, which product is this?
[1518] Please generate a detailed description for this item.
[1519] Step 5:
[1520] The device captures data such as the user's facial expressions and voice tone using a camera and microphone, and sends it to an emotion analysis engine for analysis. The input is the user's real-time facial expressions and voice data, and the output is the user's emotional state (e.g., excitement, satisfaction, etc.). Specifically, emotions are analyzed using the Affectiva SDK.
[1521] Step 6:
[1522] The server adjusts the generated content based on the emotion analysis results. The input is the emotion analysis results and the generated content, and the output is content adjusted to match the user's emotional state. Specifically, if the user is excited, it adds detailed information about new products or limited edition items.
[1523] Step 7:
[1524] The server sends the modified content to the device. The input is the modified content and the output is the content sent to the device over the Internet.
[1525] Step 8:
[1526] The device displays the received content to the user in an augmented reality format. The input is the adjusted content sent in step 7, and the output is the content overlaid on the device screen in an AR format. Specifically, text and images are displayed on top of the camera image using ARKit or ARCore.
[1527] Step 9:
[1528] Users can get more detailed information by touching specific areas on the smartphone screen. The input is the user's touch or voice input, and the output is additional information. Specifically, when a user touches a specific product, more detailed information about that product is displayed.
[1529] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1530] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1531] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1532] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1533] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1534] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1535] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1536] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1537] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1538] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1539] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1540] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1541] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1542] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1543] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1544] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1545] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1546] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1547] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1548] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1549] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1550] The following is further disclosed regarding the above embodiment.
[1551] (Claim 1)
[1552] a means for capturing video in real time using a camera on the device;
[1553] means for transmitting the captured video data to a server;
[1554] means for analyzing video data using an image recognition model in a server;
[1555] A means for generating content using a generation AI based on the analysis results;
[1556] means for transmitting the generated content to a device;
[1557] means for displaying the received content in an augmented reality format at the device;
[1558] a means for displaying further information in response to user interaction;
[1559] A system including:
[1560] (Claim 2)
[1561] 10. The system of claim 1, wherein the video data captured by the device's camera is of at least one of a work of art, a natural landscape, and an urban environment.
[1562] (Claim 3)
[1563] 10. The system of claim 1, wherein a user can interact with a particular object by touch or voice input.
[1564] "Example 1"
[1565] (Claim 1)
[1566] a means for capturing video in real time using a camera of the device;
[1567] means for transmitting the acquired video data to an information processing device;
[1568] means for analyzing video data using an image recognition algorithm in an information processing device;
[1569] means for generating content using a generative model based on the analysis results;
[1570] means for transmitting the generated content to a device;
[1571] means for displaying the received content in an augmented reality format at the device;
[1572] a means for displaying further information in response to a user action;
[1573] A system including:
[1574] (Claim 2)
[1575] 10. The system of claim 1, wherein the video data captured by the device's camera is of at least one of a work of art, a natural landscape, and an urban environment.
[1576] (Claim 3)
[1577] 2. The system according to claim 1, wherein a user can operate a specific object by touch or voice input.
[1578] "Application Example 1"
[1579] (Claim 1)
[1580] a means for capturing video in real time using a camera on the device;
[1581] means for transmitting the captured video data to a server;
[1582] means for analyzing video data using an image recognition model in a server;
[1583] A means for generating content using a generation AI based on the analysis results;
[1584] means for transmitting the generated content to a device;
[1585] means for displaying the received content in an augmented reality format at the device;
[1586] a means for displaying further information in response to user interaction;
[1587] A means for analyzing road condition images captured by a camera and generating information on traffic signs and road conditions using a generation AI model;
[1588] means for displaying the generated traffic signs and navigation information in an augmented reality format on a display of the vehicle;
[1589] means for displaying the generated navigation and warning information in an augmented reality format on the windshield of the vehicle;
[1590] A system including:
[1591] (Claim 2)
[1592] 10. The system of claim 1, wherein the video data captured by the device's camera is at least one of road conditions, traffic signs, or parking spaces.
[1593] (Claim 3)
[1594] 10. The system of claim 1, wherein a user can interact with a particular object by touch or voice input.
[1595] "Example 2: Combining Emotion Engines"
[1596] (Claim 1)
[1597] a means for capturing video in real time using a camera on the device;
[1598] means for transmitting the captured video data to a server;
[1599] means for analyzing video data using an image recognition model in a server;
[1600] A means for generating content using artificial intelligence based on the analysis results;
[1601] means for transmitting the generated content to a device;
[1602] means for displaying the received content in an augmented reality format at the device;
[1603] a means for displaying further information in response to user interaction;
[1604] A means for analyzing user emotions using an emotion engine;
[1605] a means for adjusting content based on sentiment analysis results;
[1606] A system including:
[1607] (Claim 2)
[1608] 10. The system of claim 1, wherein the video data captured by the device's camera is at least one of a work of art, a natural landscape, and an urban environment.
[1609] (Claim 3)
[1610] 10. The system of claim 1, wherein a user can interact with a particular object by touch or voice input.
[1611] "Application example 2 when combining emotion engines"
[1612] (Claim 1)
[1613] a means for capturing video in real time using a camera on the device;
[1614] means for transmitting the captured video data to a server;
[1615] means for analyzing video data using an image recognition model in a server;
[1616] A means for generating content using a generation AI based on the analysis results;
[1617] means for transmitting the generated content to a device;
[1618] means for displaying the received content in an augmented reality format at the device;
[1619] A means of analyzing user emotions,
[1620] a means for adjusting the generated content based on user sentiment;
[1621] a means for displaying further information in response to user interaction;
[1622] A system including:
[1623] (Claim 2)
[1624] The system according to claim 1, wherein the video data captured by the device's camera is of at least one of products, store interiors, and exhibits.
[1625] (Claim 3)
[1626] 10. The system of claim 1, wherein a user can interact with a particular object by touch or voice input. [Explanation of symbols]
[1627] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for capturing video in real time using a camera on the device; means for transmitting the captured video data to a server; means for analyzing video data using an image recognition model in a server; A means for generating content using a generation AI based on the analysis result; means for transmitting the generated content to a device; means for displaying the received content in an augmented reality format at the device; a means for displaying further information in response to user interaction; A system including:
2. The system of claim 1 , wherein the video data captured by the device's camera is at least one of a work of art, a natural landscape, and an urban environment.
3. The system of claim 1 , wherein a user can interact with a particular object by touch or voice input.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A