system
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-11-13
- Publication Date
- 2026-05-25
AI Technical Summary
Conventional mobile phone cases lack the ability to dynamically display media information and support interactive operations based on user input, failing to provide a personalized and engaging experience.
A system that registers media information desired by the user, dynamically displays it on a display device, and allows interactive control through gesture recognition, resolution adjustment, cropping, and filter application, optimizing the media for the display area.
Enables a personalized and engaging user experience by customizing the mobile phone case to match the user's emotions and needs, providing intuitive and real-time interaction.
Smart Images

Figure 2026085708000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional mobile phone cases have had the problem that their designs are fixed and cannot support the user's individual expression. In addition, conventional cover-type devices have lacked means for dynamically displaying media information or flexibly supporting based on the user's input operations. For this reason, the user has not been able to change the display content or perform interactive operations according to the situation or context. Therefore, there has been a demand for a more attractive and practical mobile phone case for the user.
Means for Solving the Problems
[0005] The present invention solves the above problems by providing a system that registers media information desired by the user in a device and dynamically displays said media information on a display device. By storing the media information in a storage medium and managing it based on user identification information, a personalized experience can be provided to each user. Furthermore, means are provided to apply resolution adjustment, cropping, and filters in order to optimize the media information for the display area of the display device. By providing a gesture recognition function to the terminal and enabling control of the operation of media information based on the user's input actions, an interactive experience is provided to the user.
[0006] A "user" is an individual or organization that uses the system to register or display media information.
[0007] "Media information" refers to digital content such as videos and images that are registered by the user and displayed on a display device.
[0008] A "device" is an electronic device used for registering, processing, and displaying media information.
[0009] A "display device" refers to a display or screen used to visually output media information.
[0010] A "storage medium" is a digital storage device used to store and save media information and user identification information.
[0011] "User identification information" refers to data used to identify individual users and to properly manage media information.
[0012] "Resolution adjustment" is a technical process that modifies media information to match the screen size and image quality requirements of a display device.
[0013] "Trimming" is an editing process that removes unnecessary parts of media information, leaving only the necessary portion.
[0014] "Filter" refers to an image processing algorithm for adding visual effects to media information.
[0015] "Gesture recognition function" is a technology for detecting a user's body movements and instructing the corresponding operations to the system.
[0016] "Controlling an operation" is a process of changing and executing, according to an instruction, the playback of media information, pause, application of effects, etc.
Brief Description of Drawings
[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0021] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0025] [First Embodiment]
[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0038] This invention is a system that dynamically displays user-specified media information on a display device such as a mobile phone case and interactively controls the media information in response to user operations. This system mainly consists of a server and terminals, and each component functions in cooperation with the others.
[0039] The server receives media information uploaded by users from their devices and stores it linked to user identification information. This allows for efficient management of multiple user data sets. The server also analyzes the received media information, adjusting the resolution, cropping, and applying filters. Through these processes, the media information is optimized for the display device on the mobile phone case.
[0040] Meanwhile, the device receives optimized media information delivered from the server. The device also has a function to actually display this information on the display device and automatically adjust the screen quality. Furthermore, devices with gesture recognition capabilities can detect user actions such as taps and swipes. Based on this, the device sends gesture instructions to the server and receives commands to control actions such as media playback, pause, and adding effects.
[0041] As a concrete example, let's consider a scenario where a user displays their favorite landscape video on their device and effectively controls it with specific gestures. The user uploads the selected video to the server using the app. The server trims the video, applies the desired filters, and then sends it to the device. The device displays the video and, upon receiving verbal instructions to stop playback, receives instructions from the server and controls the video according to the gestures.
[0042] This invention allows users to enjoy a new experience by customizing the surface of their mobile phone to match their emotions and needs at the moment. Such a system realizes a more intuitive and engaging user experience by reflecting individual sensibilities in the device.
[0043] The following describes the processing flow.
[0044] Step 1:
[0045] The user launches the application, selects the media information they want to display, and presses the upload button. The device processes the selected media information and compresses the data appropriately before sending it to the server.
[0046] Step 2:
[0047] The server receives compressed media data sent from the terminal. Next, it uses user identification information to save the received media data to the corresponding user profile.
[0048] Step 3:
[0049] The server processes the media information into a format suitable for the display device. Specifically, it adjusts the resolution to match the screen size of the mobile phone and crops the display area. It also applies user-specified filters and effects to the media information.
[0050] Step 4:
[0051] The server begins distributing optimized media information to the terminal. It adjusts the streaming speed according to the network conditions to ensure smooth transmission.
[0052] Step 5:
[0053] The terminal interprets data received from the server and reproduces media information on the display device. During display, it automatically adjusts screen brightness and contrast to provide the best possible visual experience.
[0054] Step 6:
[0055] The user performs gestures on the phone case. The device uses its built-in sensors to detect the gestures and sends that information to the server.
[0056] Step 7:
[0057] The server analyzes gesture data and generates operation instructions based on the user's intent. These include stopping playback, pausing, and changing effects. The server then sends the generated instructions back to the terminal.
[0058] Step 8:
[0059] The terminal controls media information on the display device based on operation instructions received from the server. It instantly reflects actions such as playback, pause, and effect changes, providing the user with real-time interaction.
[0060] (Example 1)
[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0062] In recent years, there has been a growing demand for systems that dynamically display information and support user interaction. However, existing systems suffer from inefficient information optimization and management, making it difficult to achieve low-latency information display and real-time communication. Therefore, effective means are needed to provide users with an intuitive and engaging user experience.
[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0064] In this invention, the server includes means for a user to register information with the device and display it dynamically, means for storing the information and managing it based on identification information, means for optimizing the information for the display area of the display device, and means for distributing the information using real-time communication technology. This enables a system that displays information with low latency and high quality, and that is interactively operable by the user.
[0065] "Information" refers to data that a user registers on their device and is processed and displayed in digital format.
[0066] "Device" refers to the collective term for hardware and software used for registering, storing, displaying, managing, and distributing information.
[0067] A "display device" is a device used to visually display information, and includes screens and displays.
[0068] "Motion recognition functionality" refers to a function that allows a device to detect physical actions such as user gestures and taps, and to process information and perform operations in response to those actions.
[0069] "Real-time communication technology" refers to a communication method that enables the instantaneous transmission and reception of data over a network, and its immediate reflection.
[0070] "Distribution" refers to the transmission of information to another device or user, and specifically to the transfer of data over a digital network.
[0071] This invention provides a system for dynamically displaying and interactively manipulating information selected by the user. The system consists primarily of a server and a terminal, which work together in cooperation with each other.
[0072] The server receives information sent from the user's terminal, associates it with identification information, and stores it in a database. Database software or cloud storage services can be used for this process. The server further performs optimization processes such as adjusting the resolution, cropping, and applying filters to the information. Image and video editing software is recommended for this optimization.
[0073] The terminal displays optimized information received from the server onto a display device. The display device utilizes high-quality screen technology, such as an OLED display. The terminal also has a built-in motion recognition function, allowing it to detect user input. For example, the terminal uses sensor technology to detect gestures and taps and sends instruction information to the server. Based on these instructions, the server sends commands such as play, pause, and add effects.
[0074] Users can upload their favorite videos and images to the server through the application. For example, a user can select a landscape image, apply a sepia filter to it, and display it on their device. The displayed content can be interactively managed through user gestures.
[0075] An example of a prompt message is, "Upload a landscape photo, apply a sepia filter, and view it on the terminal."
[0076] This system allows users to customize information according to their individual needs and enjoy a rich digital experience through intuitive operation.
[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0078] Step 1:
[0079] The user selects information and uploads it to the server via their device. The user selects image or video files from their device via the app's interface and sends the selected data to the server by pressing the "Upload" button.
[0080] Input: User selects an image or video file
[0081] Output: Raw media data transferred to the server
[0082] Step 2:
[0083] The server stores the received information along with identification information in a database. The server saves the received files in a specific format (e.g., "User ID_File Name_Date and Time") and organizes them using database software.
[0084] Input: Raw media data sent from the device.
[0085] Output: Media data stored in the database
[0086] Step 3:
[0087] The server retrieves information from the database and performs optimization processes such as resolution adjustment, cropping, and filter application. Using Adobe software APIs, it standardizes the image resolution, crops a specified area, and applies the selected filter.
[0088] Input: Media data stored in the database
[0089] Output: Optimized media data
[0090] Step 4:
[0091] The server sends optimized information to the terminal using real-time communication technology. WebSocket is used to achieve low-latency data transfer.
[0092] Input: Optimized media data
[0093] Output: Media data sent to the terminal
[0094] Step 5:
[0095] The terminal displays the received information on the display device. The terminal uses an OLED display to show media in high quality and automatically adjusts the brightness using an ambient light sensor.
[0096] Input: Media data sent from the server
[0097] Output: Media displayed on the display device
[0098] Step 6:
[0099] Users interact with the displayed information by performing gestures on the device. The device uses sensors to detect actions such as taps and swipes and sends that data as instructions to the server.
[0100] Input: User gestures
[0101] Output: Instruction information sent to the server
[0102] Step 7:
[0103] Based on the received instructions, the server sends the necessary media control commands back to the terminal. This allows for actions such as playback, stopping, and adding effects.
[0104] Input: Instructions sent from the terminal
[0105] Output: Control commands to the terminal
[0106] (Application Example 1)
[0107] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0108] In today's commercial spaces, information provided to consumers is diversifying, but it is difficult for consumers to effectively and intuitively obtain the information they truly need. Especially for busy modern people, there is a need for personalized information delivery systems that allow quick access to necessary information. Furthermore, the sheer volume of information often overwhelms consumers, leading to confusion in their choices. A system is needed to improve this situation and realize efficient and personalized information delivery.
[0109] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0110] In this invention, the server includes means for a user to register media information with the device and dynamically display the media information, means for storing the media information on a storage medium and managing the media information based on user identification information, and means for detecting the user's gaze and actions using an eye-tracking device and displaying the media information in an individually customized manner. This makes it possible for consumers to obtain personalized information in real time based on their gaze and gestures within a store.
[0111] "User identification information" refers to information used to uniquely identify each user and manage associated media data.
[0112] "Media information" refers to media data such as images, videos, and audio registered by users.
[0113] A "display device" is a display device that dynamically displays media information registered by the user.
[0114] "Gesture recognition" is a function that detects user actions such as taps and swipes and controls media information in response to those actions.
[0115] An "eye-tracking device" is a device that includes sensors that detect the user's gaze and enable the display of information based on that gaze.
[0116] "Individually customized display using visual devices" refers to a function that displays media information tailored to individual users based on eye tracking and gesture recognition.
[0117] "Displaying with low latency" refers to a process that quickly reflects media information on the display device in real time, allowing for immediate response to user interaction.
[0118] A system implementing this invention primarily consists of three components: a server, a terminal (such as smart glasses), and a user.
[0119] The server has the function of receiving, managing, and optimizing media information registered by the user. Specifically, the server stores media information based on user identification information, adjusts the resolution to suit display on eye-tracking devices, and applies cropping and filters as needed.
[0120] The smart glasses, acting as the terminal, receive optimized media information from a server and display it in real time within the user's field of view. The smart glasses are equipped with eye-tracking sensors and gesture recognition functions, detecting the user's gaze and movements and changing the displayed content accordingly to provide personalized information.
[0121] Through this system, users can obtain real-time information and promotions for products they are interested in using their gaze and gestures within the store. When a user's gaze lingers at a certain point, detailed information related to that product is automatically displayed on the screen. For example, if a user looks at a specific product in a certain section, details and related videos for that product will be displayed.
[0122] By utilizing a generative AI model, the system enables real-time information presentation in response to a prompt question such as: How can media information be displayed and interaction be achieved when a user wearing smart glasses directs their gaze towards a specified product?
[0123] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0124] Step 1:
[0125] The user registers media information using a device. As input, the user selects media (images, videos, etc.) from smart glasses or another device and uploads it to the server. The server receives this information and stores it in its storage device along with the user's identification information. As output, the server stores media information linked to individual users.
[0126] Step 2:
[0127] The server optimizes the media information it receives. The server analyzes the stored media information as input. It performs data processing such as resolution adjustment, cropping, and filter application to ensure optimal display for visual devices. The optimized media information is generated as output and saved again.
[0128] Step 3:
[0129] The device obtains optimized media information from the server. As input, the user's smart glasses access the server and request the necessary information. Based on this request, the server sends the corresponding optimized media information as output to the device. The device receives this information and prepares the display.
[0130] Step 4:
[0131] The device displays media information using eye-tracking and gesture recognition. Input consists of the user's gaze and gestures, which the device captures using sensors. Based on the data calculations, the device performs appropriate media control based on the user's gaze direction and gesture content. As output, personalized media information is displayed on the screen in real time.
[0132] Step 5:
[0133] Based on user interaction, media information is dynamically controlled. Input includes user gestures and eye-tracking, which the terminal analyzes and sends instructions to the server. The server receives these instructions and performs actions such as starting, stopping, and adding effects to media playback. The output is media playback that meets the user's expectations.
[0134] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0135] This invention is a system that recognizes a user's emotions in real time and dynamically adjusts the display of media information and interactions based on the recognition results. This system consists of three main components: a server, a terminal, and an emotion engine.
[0136] The device is equipped with an emotion engine that recognizes emotions from the user's facial expressions and voice. The emotion engine analyzes emotional data in real time as the user uses the device and sends the results to a server. The server uses this emotional data to evaluate the user's state and determine the optimal selection and display method of media information.
[0137] The server manages the user's past emotional history based on emotional data and learns the emotional tendencies of individual users. Based on this information, the server can recommend appropriate media information to the user and prompt the user to make a selection.
[0138] The terminal receives instructions from the server and displays media information optimized for the display device. At the same time, it has the function to dynamically adjust effects and playback content according to the user's emotions. For example, if the user expresses the emotion of "joy," it may apply a bright filter to the video or add cheerful music.
[0139] As a concrete example, consider a scenario where a user views weather information as media information. If the emotion engine detects that the user is feeling gloomy because it's a rainy day, the server can use a transparent information overlay to display an additional video of a relaxing landscape. This enables personalized information delivery that aligns with the user's emotions.
[0140] This invention allows users to enjoy digital content experiences tailored to their individual emotional states, further enhancing the value of the device.
[0141] The following describes the processing flow.
[0142] Step 1:
[0143] The user operates the device and selects media information through the application. The device processes this selection information, activates an emotion engine, and captures the user's facial expressions and voice data in real time.
[0144] Step 2:
[0145] The emotion engine analyzes the captured data to identify the user's current emotions. This identification result is generated as an emotion tag within the device.
[0146] Step 3:
[0147] The device sends the generated emotion tag and selected media information to the server. The server receives the emotion tag and analyzes it along with the user's past emotion history.
[0148] Step 4:
[0149] The server determines media information and effects to recommend to the user based on emotion tags. For example, if the user indicates "joy," the server will recommend videos and music with a cheerful tone.
[0150] Step 5:
[0151] The server transmits the selected media information and additional effects to the terminal. The terminal reflects this information on the display device, providing the user with visual and auditory feedback.
[0152] Step 6:
[0153] The user interacts with the media information being displayed. The device, through its emotion engine, recaptures the user's facial expressions and gestures to detect changes in emotion.
[0154] Step 7:
[0155] The device sends newly detected emotion data to the server and awaits instructions to readjust the displayed content. The server processes this information and recommends updating to the optimal display content.
[0156] Step 8:
[0157] The device continuously updates the display of media information in response to instructions from the final server, providing the user with an experience that is always in line with their current emotions.
[0158] (Example 2)
[0159] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0160] Conventional information display systems had the problem of not being able to adjust content to take into account the user's emotional state, and thus failing to provide a personalized experience. Furthermore, static information content made it difficult to adapt in real time to the user's interests and emotions, resulting in users not being able to have a satisfying experience.
[0161] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0162] In this invention, the server includes means for a processing device that analyzes the user's nonverbal expressions to obtain their emotional state, means for dynamically adjusting the selected and displayed information content based on the emotional state, and means for generating and displaying information content corresponding to the user's emotional state using a generative AI model. This enables real-time and personalized information provision based on the user's emotions.
[0163] "Nonverbal expression" refers to means of communication that do not involve language, and includes gestures, facial expressions, tone of voice, and so on.
[0164] "Emotional state" refers to the user's psychological and emotional state, such as joy, sadness, or depression.
[0165] "Information content" is a general term for digital information provided to users, and includes text, images, audio, and video.
[0166] A "generative AI model" refers to an artificial intelligence model that uses machine learning algorithms to generate information tailored to a specific purpose from particular input data.
[0167] A "prompt sentence" is a sentence that serves as an input instruction to an AI model, and is intended to guide it to a specific output result.
[0168] A "processing device" is a set of hardware or software that receives data, analyzes it, and outputs the results.
[0169] This invention is a system that provides information content in real time according to the user's emotional state. The main components consist of a terminal, a server, and an emotion analysis engine.
[0170] The device is equipped with a camera and microphone to capture the user's facial expressions and voice. An emotion analysis engine analyzes this data and identifies the user's emotional state in real time from their nonverbal expressions. Deep learning technology can be used for the specific analysis.
[0171] The analyzed sentiment data is sent to the server via an encrypted protocol. The server records the received sentiment data in a database and can learn the user's past sentiment history using machine learning algorithms. Based on this information, the server selects and provides information content that is appropriate for the user.
[0172] Furthermore, the server can utilize a generative AI model to generate new informational content tailored to the user's emotional state. This generative AI model uses historical emotional data and real-time input data to customize the information best suited to the user.
[0173] The device displays informational content based on instructions from the server. During display, resolution adjustments and effects are added to optimize the user experience. For example, if the user expresses joy, the device enhances the content by applying a bright filter to the video and playing cheerful music.
[0174] As a concrete example, the prompt message could include "When the user expresses feelings of joy, please display a positive news article." This prompt message can be used to encourage the generative AI model to generate appropriate information.
[0175] In this way, users can receive personalized content that responds to their emotions, thereby improving the quality of their digital experience.
[0176] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0177] Step 1:
[0178] The device captures the user's facial expressions and voice using its camera and microphone. The input data consists of real-time video and audio data, and the device's emotion analysis engine analyzes this data to identify the user's emotional state. In this step, a deep learning model is used to perform data calculations that classify the emotional state into categories such as "joy" or "sadness." The output is data containing tags that indicate the user's emotional state.
[0179] Step 2:
[0180] The device sends emotional state data it has acquired to the server. The input data consists of tags indicating the emotional state. The server receives this data and analyzes it by cross-referencing it with past emotional history. Storage in the database and learning of the user's emotional tendencies take place here. The output is analytical data that reflects the user's emotional tendencies.
[0181] Step 3:
[0182] The server uses a generative AI model based on the user's emotional state and past history data to generate appropriate informational content. In this step, prompts are used to input into the generative AI model, generating content optimized for the user. The output is informational content that corresponds to the user's emotional state.
[0183] Step 4:
[0184] The server generates information content and sends it from the server to the terminal. The input data is the generated information content, which the terminal receives. The terminal adjusts the resolution and adds effects to display the information content on the screen. The output is the display data presented to the user in an optimized form.
[0185] Step 5:
[0186] Users view informational content through their devices and perform gestures and other interactions. The device utilizes gesture recognition to dynamically control the informational content (play, pause, add effects) in response to user input. The input data is the user's operation signal, and the output is the state of the content after the operation.
[0187] These steps allow users to receive personalized digital content that responds to their emotions.
[0188] (Application Example 2)
[0189] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0190] Modern information delivery systems require dynamic information delivery tailored to the user's emotional state, but conventional systems have limitations in providing personalized content that responds to each user's individual emotional state. This makes it difficult to provide information that resonates with users' emotions, contributing to decreased user satisfaction.
[0191] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0192] In this invention, the server includes means for a user to register information on a device and dynamically display the information on a display device; means for storing the information on a storage medium and managing the information based on user identification information; and means for analyzing the user's emotions and dynamically adjusting the information displayed based on those emotions. This makes it possible to provide personalized information that is tailored to the user's emotional state.
[0193] "User identification information" refers to information used to identify individual users, and plays a role in uniquely identifying users within the system.
[0194] "Emotional data" refers to information that indicates the emotional state of a user, analyzed from their facial expressions, voice, etc., and is data that specifically represents the user's psychological response.
[0195] "Gesture recognition functionality" is a feature that detects the user's physical movements and uses them as system input, and is a technology that contributes to improving the user interface.
[0196] "Resolution adjustment" is the process of changing the resolution of images or videos in order to optimize the display quality of information on a display device.
[0197] "Cropping" is a technique for editing images and videos by cutting out unnecessary parts to make them the appropriate size and configuration for the display screen.
[0198] A "filter" is a means of adding visual or auditory effects to a medium, and is a technique for improving the sensory quality of information.
[0199] "Personalized information delivery" is the process of providing information in the format most suitable for each user, based on data such as the user's past behavior and individual emotional state.
[0200] "Dynamic adjustment" is a concept that refers to changing the system's output or behavior on the fly in response to real-time data and circumstances.
[0201] The system that realizes this invention recognizes the user's emotions in real time, dynamically adjusts media information based on emotion data, and provides personalized information. The system mainly consists of a server, a terminal, and an emotion engine.
[0202] The device collects the user's facial expressions and voice, and analyzes this data using an emotion engine. The emotion engine uses AI technology to recognize the user's emotions in real time and sends this data to a server. Based on the received emotion data, the server learns the user's individual emotional tendencies and selects the most appropriate media information. In this process, the server also utilizes past emotional history stored on the storage medium to achieve more accurate personalization.
[0203] The server then sends media information selected to match the user's emotions to the terminal. The terminal optimizes this information for user interaction within the display area and performs processing such as resolution adjustment and filter application as needed. This system also features gesture recognition, allowing users to manipulate media information based on their actions.
[0204] For example, if the emotion engine detects that a user is feeling stressed, the system will dynamically display relaxation videos or music on the device to provide a relaxing effect. This creates a personalized entertainment experience tailored to the user's emotional state.
[0205] A concrete example of utilizing a generative AI model is a prompt such as, "The user's current emotion is 'surprise.' Please recommend a movie that is suitable for this emotion." This prompt allows the system to suggest content related to surprise, providing the user with the most relevant information.
[0206] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0207] Step 1:
[0208] The device captures the user's facial expressions and voice in real time. This data is collected using the camera and microphone and sent as input to the emotion engine. The emotion engine analyzes this input and outputs data indicating the user's emotional state.
[0209] Step 2:
[0210] The server receives sentiment data sent from the sentiment engine. Based on this data, the server references the user's past sentiment history and learns the user's sentiment tendencies. This prepares the server to select the most appropriate media information in the next step.
[0211] Step 3:
[0212] The server uses a generative AI model to select appropriate media information based on the user's current emotional state. In this process, emotional data is input to the generative AI model as a prompt, generating information such as, "The user's current emotion is 'surprise.' Please recommend a movie that is suitable for this emotion." The server then generates the optimal media information as output.
[0213] Step 4:
[0214] The server sends selected media information to the terminal. Upon receiving this information, the terminal optimizes it according to the display area. Specifically, it adjusts the resolution and applies filters to prepare the content for display to the user. This improves the user interface.
[0215] Step 5:
[0216] The device displays optimized media information to the user and is capable of recognizing the user's gestures. Through gesture recognition, the user can perform actions such as playing, pausing, and adding effects to the displayed information. This process allows the user to freely manipulate content that aligns with their emotions.
[0217] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0218] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0219] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0220] [Second Embodiment]
[0221] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0222] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0223] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0224] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0225] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0226] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0227] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0228] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0229] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0230] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0231] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0232] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0233] This invention is a system that dynamically displays user-specified media information on a display device such as a mobile phone case and interactively controls the media information in response to user operations. This system mainly consists of a server and terminals, and each component functions in cooperation with the others.
[0234] The server receives media information uploaded by users from their devices and stores it linked to user identification information. This allows for efficient management of multiple user data sets. The server also analyzes the received media information, adjusting the resolution, cropping, and applying filters. Through these processes, the media information is optimized for the display device on the mobile phone case.
[0235] Meanwhile, the device receives optimized media information delivered from the server. The device also has a function to actually display this information on the display device and automatically adjust the screen quality. Furthermore, devices with gesture recognition capabilities can detect user actions such as taps and swipes. Based on this, the device sends gesture instructions to the server and receives commands to control actions such as media playback, pause, and adding effects.
[0236] As a concrete example, let's consider a scenario where a user displays their favorite landscape video on their device and effectively controls it with specific gestures. The user uploads the selected video to the server using the app. The server trims the video, applies the desired filters, and then sends it to the device. The device displays the video and, upon receiving verbal instructions to stop playback, receives instructions from the server and controls the video according to the gestures.
[0237] This invention allows users to enjoy a new experience by customizing the surface of their mobile phone to match their emotions and needs at the moment. Such a system realizes a more intuitive and engaging user experience by reflecting individual sensibilities in the device.
[0238] The following describes the processing flow.
[0239] Step 1:
[0240] The user launches the application, selects the media information they want to display, and presses the upload button. The device processes the selected media information and compresses the data appropriately before sending it to the server.
[0241] Step 2:
[0242] The server receives compressed media data sent from the terminal. Next, it uses user identification information to save the received media data to the corresponding user profile.
[0243] Step 3:
[0244] The server processes the media information into a format suitable for the display device. Specifically, it adjusts the resolution to match the screen size of the mobile phone and crops the display area. It also applies user-specified filters and effects to the media information.
[0245] Step 4:
[0246] The server begins distributing optimized media information to the terminal. It adjusts the streaming speed according to the network conditions to ensure smooth transmission.
[0247] Step 5:
[0248] The terminal interprets data received from the server and reproduces media information on the display device. During display, it automatically adjusts screen brightness and contrast to provide the best possible visual experience.
[0249] Step 6:
[0250] The user performs gestures on the phone case. The device uses its built-in sensors to detect the gestures and sends that information to the server.
[0251] Step 7:
[0252] The server analyzes gesture data and generates operation instructions based on the user's intent. These include stopping playback, pausing, and changing effects. The server then sends the generated instructions back to the terminal.
[0253] Step 8:
[0254] The terminal controls media information on the display device based on operation instructions received from the server. It instantly reflects actions such as playback, pause, and effect changes, providing the user with real-time interaction.
[0255] (Example 1)
[0256] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0257] In recent years, there has been a growing demand for systems that dynamically display information and support user interaction. However, existing systems suffer from inefficient information optimization and management, making it difficult to achieve low-latency information display and real-time communication. Therefore, effective means are needed to provide users with an intuitive and engaging user experience.
[0258] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0259] In this invention, the server includes means for a user to register information with the device and display it dynamically, means for storing the information and managing it based on identification information, means for optimizing the information for the display area of the display device, and means for distributing the information using real-time communication technology. This enables a system that displays information with low latency and high quality, and that is interactively operable by the user.
[0260] "Information" refers to data that a user registers on their device and is processed and displayed in digital format.
[0261] "Device" refers to the collective term for hardware and software used for registering, storing, displaying, managing, and distributing information.
[0262] A "display device" is a device used to visually display information, and includes screens and displays.
[0263] "Motion recognition functionality" refers to a function that allows a device to detect physical actions such as user gestures and taps, and to process information and perform operations in response to those actions.
[0264] "Real-time communication technology" refers to a communication method that enables the instantaneous transmission and reception of data over a network, and its immediate reflection.
[0265] "Distribution" refers to the transmission of information to another device or user, and specifically to the transfer of data over a digital network.
[0266] This invention provides a system for dynamically displaying and interactively manipulating information selected by the user. The system consists primarily of a server and a terminal, which work together in cooperation with each other.
[0267] The server receives information sent from the user's terminal, associates it with identification information, and stores it in a database. Database software or cloud storage services can be used for this process. The server further performs optimization processes such as adjusting the resolution, cropping, and applying filters to the information. Image and video editing software is recommended for this optimization.
[0268] The terminal displays optimized information received from the server onto a display device. The display device utilizes high-quality screen technology, such as an OLED display. The terminal also has a built-in motion recognition function, allowing it to detect user input. For example, the terminal uses sensor technology to detect gestures and taps and sends instruction information to the server. Based on these instructions, the server sends commands such as play, pause, and add effects.
[0269] Users can upload their favorite videos and images to the server through the application. For example, a user can select a landscape image, apply a sepia filter to it, and display it on their device. The displayed content can be interactively managed through user gestures.
[0270] An example of a prompt message is, "Upload a landscape photo, apply a sepia filter, and view it on the terminal."
[0271] This system allows users to customize information according to their individual needs and enjoy a rich digital experience through intuitive operation.
[0272] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0273] Step 1:
[0274] The user selects information and uploads it to the server via their device. The user selects image or video files from their device via the app's interface and sends the selected data to the server by pressing the "Upload" button.
[0275] Input: User selects an image or video file
[0276] Output: Raw media data transferred to the server
[0277] Step 2:
[0278] The server stores the received information along with identification information in a database. The server saves the received files in a specific format (e.g., "User ID_File Name_Date and Time") and organizes them using database software.
[0279] Input: Raw media data sent from the device.
[0280] Output: Media data stored in the database
[0281] Step 3:
[0282] The server retrieves information from the database and performs optimization processes such as resolution adjustment, trimming, and filter application. Using Adobe software APIs, it standardizes the resolution of the image, crops the specified area, and applies the selected filter.
[0283] Input: Media data stored in the database
[0284] Output: Optimized media data
[0285] Step 4:
[0286] The server sends the optimized information to the terminal using real-time communication technology. It utilizes WebSocket to achieve low-latency data transfer.
[0287] Input: Optimized media data
[0288] Output: Media data sent to the terminal
[0289] Step 5:
[0290] The terminal displays the received information on the display device. The terminal uses an OLED display to display media in high quality and automatically adjusts the brightness with an ambient light sensor.
[0291] Input: Media data sent from the server
[0292] Output: Media projected onto the display device
[0293] Step 6:
[0294] Users interact with the displayed information by performing gestures on the device. The device uses sensors to detect actions such as taps and swipes and sends that data as instructions to the server.
[0295] Input: User gestures
[0296] Output: Instruction information sent to the server
[0297] Step 7:
[0298] Based on the received instructions, the server sends the necessary media control commands back to the terminal. This allows for actions such as playback, stopping, and adding effects.
[0299] Input: Instructions sent from the terminal
[0300] Output: Control commands to the terminal
[0301] (Application Example 1)
[0302] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0303] In today's commercial spaces, information provided to consumers is diversifying, but it is difficult for consumers to effectively and intuitively obtain the information they truly need. Especially for busy modern people, there is a need for personalized information delivery systems that allow quick access to necessary information. Furthermore, the sheer volume of information often overwhelms consumers, leading to confusion in their choices. A system is needed to improve this situation and realize efficient and personalized information delivery.
[0304] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0305] In this invention, the server includes means for a user to register media information in a device and dynamically display the media information, means for storing the media information in a storage medium and managing the media information based on user identification information, and means for using a gaze tracking device to detect the user's gaze and actions and perform a customized display of the media information individually. As a result, consumers can obtain personalized information in real time by gaze and gesture in a store.
[0306] "User identification information" is information for uniquely identifying each user and managing associated media data.
[0307] "Media information" is media data such as images, videos, and sounds registered by a user.
[0308] "Display device" is a display device for dynamically displaying media information registered by a user.
[0309] "Gesture recognition function" is a function for detecting a user's operations such as taps and swipes and controlling media information according to the operations.
[0310] "Gaze tracking device" is a device including a sensor for detecting a user's gaze and enabling information display based on it.
[0311] "Customized display using a visual device" is a function for performing a display of media information specialized for individual users based on gaze tracking and gesture recognition.
[0312] "Displaying with low latency" is a process for promptly reflecting media information on a display device in real time to respond immediately to user interactions.
[0313] The system for implementing this invention mainly consists of three parties: a server, a terminal (such as smart glasses), and a user.
[0314] The server has the function of receiving, managing, and optimizing media information registered by the user. Specifically, the server stores media information based on user identification information, adjusts the resolution to suit display on eye-tracking devices, and applies cropping and filters as needed.
[0315] The smart glasses, acting as the terminal, receive optimized media information from a server and display it in real time within the user's field of view. The smart glasses are equipped with eye-tracking sensors and gesture recognition functions, detecting the user's gaze and movements and changing the displayed content accordingly to provide personalized information.
[0316] Through this system, users can obtain real-time information and promotions for products they are interested in using their gaze and gestures within the store. When a user's gaze lingers at a certain point, detailed information related to that product is automatically displayed on the screen. For example, if a user looks at a specific product in a certain section, details and related videos for that product will be displayed.
[0317] By utilizing a generative AI model, the system enables real-time information presentation in response to a prompt question such as: How can media information be displayed and interaction be achieved when a user wearing smart glasses directs their gaze towards a specified product?
[0318] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0319] Step 1:
[0320] The user registers media information using a device. As input, the user selects media (images, videos, etc.) from smart glasses or another device and uploads it to the server. The server receives this information and stores it in its storage device along with the user's identification information. As output, the server stores media information linked to individual users.
[0321] Step 2:
[0322] The server optimizes the media information it receives. The server analyzes the stored media information as input. It performs data processing such as resolution adjustment, cropping, and filter application to ensure optimal display for visual devices. The optimized media information is generated as output and saved again.
[0323] Step 3:
[0324] The device obtains optimized media information from the server. As input, the user's smart glasses access the server and request the necessary information. Based on this request, the server sends the corresponding optimized media information as output to the device. The device receives this information and prepares the display.
[0325] Step 4:
[0326] The device displays media information using eye-tracking and gesture recognition. Input consists of the user's gaze and gestures, which the device captures using sensors. Based on the data calculations, the device performs appropriate media control based on the user's gaze direction and gesture content. As output, personalized media information is displayed on the screen in real time.
[0327] Step 5:
[0328] Based on user interaction, media information is dynamically controlled. Input includes user gestures and eye-tracking, which the terminal analyzes and sends instructions to the server. The server receives these instructions and performs actions such as starting, stopping, and adding effects to media playback. The output is media playback that meets the user's expectations.
[0329] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0330] This invention is a system that recognizes a user's emotions in real time and dynamically adjusts the display of media information and interactions based on the recognition results. This system consists of three main components: a server, a terminal, and an emotion engine.
[0331] The device is equipped with an emotion engine that recognizes emotions from the user's facial expressions and voice. The emotion engine analyzes emotional data in real time as the user uses the device and sends the results to a server. The server uses this emotional data to evaluate the user's state and determine the optimal selection and display method of media information.
[0332] The server manages the user's past emotional history based on emotional data and learns the emotional tendencies of individual users. Based on this information, the server can recommend appropriate media information to the user and prompt the user to make a selection.
[0333] The terminal receives instructions from the server and displays media information optimized for the display device. At the same time, it has the function to dynamically adjust effects and playback content according to the user's emotions. For example, if the user expresses the emotion of "joy," it may apply a bright filter to the video or add cheerful music.
[0334] As a concrete example, consider a scenario where a user views weather information as media information. If the emotion engine detects that the user is feeling gloomy because it's a rainy day, the server can use a transparent information overlay to display an additional video of a relaxing landscape. This enables personalized information delivery that aligns with the user's emotions.
[0335] This invention allows users to enjoy digital content experiences tailored to their individual emotional states, further enhancing the value of the device.
[0336] The following describes the processing flow.
[0337] Step 1:
[0338] The user operates the device and selects media information through the application. The device processes this selection information, activates an emotion engine, and captures the user's facial expressions and voice data in real time.
[0339] Step 2:
[0340] The emotion engine analyzes the captured data to identify the user's current emotions. This identification result is generated as an emotion tag within the device.
[0341] Step 3:
[0342] The device sends the generated emotion tag and selected media information to the server. The server receives the emotion tag and analyzes it along with the user's past emotion history.
[0343] Step 4:
[0344] The server determines media information and effects to recommend to the user based on emotion tags. For example, if the user indicates "joy," the server will recommend videos and music with a cheerful tone.
[0345] Step 5:
[0346] The server transmits the selected media information and additional effects to the terminal. The terminal reflects this information on the display device, providing the user with visual and auditory feedback.
[0347] Step 6:
[0348] The user interacts with the media information being displayed. The device, through its emotion engine, recaptures the user's facial expressions and gestures to detect changes in emotion.
[0349] Step 7:
[0350] The device sends newly detected emotion data to the server and awaits instructions to readjust the displayed content. The server processes this information and recommends updating to the optimal display content.
[0351] Step 8:
[0352] The device continuously updates the display of media information in response to instructions from the final server, providing the user with an experience that is always in line with their current emotions.
[0353] (Example 2)
[0354] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0355] Conventional information display systems had the problem of not being able to adjust content to take into account the user's emotional state, and thus failing to provide a personalized experience. Furthermore, static information content made it difficult to adapt in real time to the user's interests and emotions, resulting in users not being able to have a satisfying experience.
[0356] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0357] In this invention, the server includes means for a processing device that analyzes the user's nonverbal expressions to obtain their emotional state, means for dynamically adjusting the selected and displayed information content based on the emotional state, and means for generating and displaying information content corresponding to the user's emotional state using a generative AI model. This enables real-time and personalized information provision based on the user's emotions.
[0358] "Nonverbal expression" refers to means of communication that do not involve language, and includes gestures, facial expressions, tone of voice, and so on.
[0359] "Emotional state" refers to the user's psychological and emotional state, such as joy, sadness, or depression.
[0360] "Information content" is a general term for digital information provided to users, and includes text, images, audio, and video.
[0361] A "generative AI model" refers to an artificial intelligence model that uses machine learning algorithms to generate information tailored to a specific purpose from particular input data.
[0362] A "prompt sentence" is a sentence that serves as an input instruction to an AI model, and is intended to guide it to a specific output result.
[0363] A "processing device" is a set of hardware or software that receives data, analyzes it, and outputs the results.
[0364] This invention is a system that provides information content in real time according to the user's emotional state. The main components consist of a terminal, a server, and an emotion analysis engine.
[0365] The device is equipped with a camera and microphone to capture the user's facial expressions and voice. An emotion analysis engine analyzes this data and identifies the user's emotional state in real time from their nonverbal expressions. Deep learning technology can be used for the specific analysis.
[0366] The analyzed sentiment data is sent to the server via an encrypted protocol. The server records the received sentiment data in a database and can learn the user's past sentiment history using machine learning algorithms. Based on this information, the server selects and provides information content that is appropriate for the user.
[0367] Furthermore, the server can utilize a generative AI model to generate new informational content tailored to the user's emotional state. This generative AI model uses historical emotional data and real-time input data to customize the information best suited to the user.
[0368] The device displays informational content based on instructions from the server. During display, resolution adjustments and effects are added to optimize the user experience. For example, if the user expresses joy, the device enhances the content by applying a bright filter to the video and playing cheerful music.
[0369] As a concrete example, the prompt message could include "When the user expresses feelings of joy, please display a positive news article." This prompt message can be used to encourage the generative AI model to generate appropriate information.
[0370] In this way, users can receive personalized content that responds to their emotions, thereby improving the quality of their digital experience.
[0371] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0372] Step 1:
[0373] The device captures the user's facial expressions and voice using its camera and microphone. The input data consists of real-time video and audio data, and the device's emotion analysis engine analyzes this data to identify the user's emotional state. In this step, a deep learning model is used to perform data calculations that classify the emotional state into categories such as "joy" or "sadness." The output is data containing tags that indicate the user's emotional state.
[0374] Step 2:
[0375] The device sends emotional state data it has acquired to the server. The input data consists of tags indicating the emotional state. The server receives this data and analyzes it by cross-referencing it with past emotional history. Storage in the database and learning of the user's emotional tendencies take place here. The output is analytical data that reflects the user's emotional tendencies.
[0376] Step 3:
[0377] The server uses a generative AI model based on the user's emotional state and past history data to generate appropriate informational content. In this step, prompts are used to input into the generative AI model, generating content optimized for the user. The output is informational content that corresponds to the user's emotional state.
[0378] Step 4:
[0379] The server generates information content and sends it from the server to the terminal. The input data is the generated information content, which the terminal receives. The terminal adjusts the resolution and adds effects to display the information content on the screen. The output is the display data presented to the user in an optimized form.
[0380] Step 5:
[0381] Users view informational content through their devices and perform gestures and other interactions. The device utilizes gesture recognition to dynamically control the informational content (play, pause, add effects) in response to user input. The input data is the user's operation signal, and the output is the state of the content after the operation.
[0382] These steps allow users to receive personalized digital content that responds to their emotions.
[0383] (Application Example 2)
[0384] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0385] Modern information delivery systems require dynamic information delivery tailored to the user's emotional state, but conventional systems have limitations in providing personalized content that responds to each user's individual emotional state. This makes it difficult to provide information that resonates with users' emotions, contributing to decreased user satisfaction.
[0386] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0387] In this invention, the server includes means for a user to register information on a device and dynamically display the information on a display device; means for storing the information on a storage medium and managing the information based on user identification information; and means for analyzing the user's emotions and dynamically adjusting the information displayed based on those emotions. This makes it possible to provide personalized information that is tailored to the user's emotional state.
[0388] "User identification information" refers to information used to identify individual users, and plays a role in uniquely identifying users within the system.
[0389] "Emotional data" refers to information that indicates the emotional state of a user, analyzed from their facial expressions, voice, etc., and is data that specifically represents the user's psychological response.
[0390] "Gesture recognition functionality" is a feature that detects the user's physical movements and uses them as system input, and is a technology that contributes to improving the user interface.
[0391] "Resolution adjustment" is the process of changing the resolution of images or videos in order to optimize the display quality of information on a display device.
[0392] "Cropping" is a technique for editing images and videos by cutting out unnecessary parts to make them the appropriate size and configuration for the display screen.
[0393] A "filter" is a means of adding visual or auditory effects to a medium, and is a technique for improving the sensory quality of information.
[0394] "Personalized information delivery" is the process of providing information in the format most suitable for each user, based on data such as the user's past behavior and individual emotional state.
[0395] "Dynamic adjustment" is a concept that refers to changing the system's output or behavior on the fly in response to real-time data and circumstances.
[0396] The system that realizes this invention recognizes the user's emotions in real time, dynamically adjusts media information based on emotion data, and provides personalized information. The system mainly consists of a server, a terminal, and an emotion engine.
[0397] The device collects the user's facial expressions and voice, and analyzes this data using an emotion engine. The emotion engine uses AI technology to recognize the user's emotions in real time and sends this data to a server. Based on the received emotion data, the server learns the user's individual emotional tendencies and selects the most appropriate media information. In this process, the server also utilizes past emotional history stored on the storage medium to achieve more accurate personalization.
[0398] The server then sends media information selected to match the user's emotions to the terminal. The terminal optimizes this information for user interaction within the display area and performs processing such as resolution adjustment and filter application as needed. This system also features gesture recognition, allowing users to manipulate media information based on their actions.
[0399] For example, if the emotion engine detects that a user is feeling stressed, the system will dynamically display relaxation videos or music on the device to provide a relaxing effect. This creates a personalized entertainment experience tailored to the user's emotional state.
[0400] A concrete example of utilizing a generative AI model is a prompt such as, "The user's current emotion is 'surprise.' Please recommend a movie that is suitable for this emotion." This prompt allows the system to suggest content related to surprise, providing the user with the most relevant information.
[0401] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0402] Step 1:
[0403] The device captures the user's facial expressions and voice in real time. This data is collected using the camera and microphone and sent as input to the emotion engine. The emotion engine analyzes this input and outputs data indicating the user's emotional state.
[0404] Step 2:
[0405] The server receives sentiment data sent from the sentiment engine. Based on this data, the server references the user's past sentiment history and learns the user's sentiment tendencies. This prepares the server to select the most appropriate media information in the next step.
[0406] Step 3:
[0407] The server uses a generative AI model to select appropriate media information based on the user's current emotional state. In this process, emotional data is input to the generative AI model as a prompt, generating information such as, "The user's current emotion is 'surprise.' Please recommend a movie that is suitable for this emotion." The server then generates the optimal media information as output.
[0408] Step 4:
[0409] The server sends selected media information to the terminal. Upon receiving this information, the terminal optimizes it according to the display area. Specifically, it adjusts the resolution and applies filters to prepare the content for display to the user. This improves the user interface.
[0410] Step 5:
[0411] The device displays optimized media information to the user and is capable of recognizing the user's gestures. Through gesture recognition, the user can perform actions such as playing, pausing, and adding effects to the displayed information. This process allows the user to freely manipulate content that aligns with their emotions.
[0412] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0413] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0414] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0415] [Third Embodiment]
[0416] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0417] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0418] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0419] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0420] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0421] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0422] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0423] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0424] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0425] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0426] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0427] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0428] This invention is a system that dynamically displays user-specified media information on a display device such as a mobile phone case and interactively controls the media information in response to user operations. This system mainly consists of a server and terminals, and each component functions in cooperation with the others.
[0429] The server receives media information uploaded by users from their devices and stores it linked to user identification information. This allows for efficient management of multiple user data sets. The server also analyzes the received media information, adjusting the resolution, cropping, and applying filters. Through these processes, the media information is optimized for the display device on the mobile phone case.
[0430] Meanwhile, the device receives optimized media information delivered from the server. The device also has a function to actually display this information on the display device and automatically adjust the screen quality. Furthermore, devices with gesture recognition capabilities can detect user actions such as taps and swipes. Based on this, the device sends gesture instructions to the server and receives commands to control actions such as media playback, pause, and adding effects.
[0431] As a concrete example, let's consider a scenario where a user displays their favorite landscape video on their device and effectively controls it with specific gestures. The user uploads the selected video to the server using the app. The server trims the video, applies the desired filters, and then sends it to the device. The device displays the video and, upon receiving verbal instructions to stop playback, receives instructions from the server and controls the video according to the gestures.
[0432] This invention allows users to enjoy a new experience by customizing the surface of their mobile phone to match their emotions and needs at the moment. Such a system realizes a more intuitive and engaging user experience by reflecting individual sensibilities in the device.
[0433] The following describes the processing flow.
[0434] Step 1:
[0435] The user launches the application, selects the media information they want to display, and presses the upload button. The device processes the selected media information and compresses the data appropriately before sending it to the server.
[0436] Step 2:
[0437] The server receives compressed media data sent from the terminal. Next, it uses user identification information to save the received media data to the corresponding user profile.
[0438] Step 3:
[0439] The server processes the media information into a format suitable for the display device. Specifically, it adjusts the resolution to match the screen size of the mobile phone and crops the display area. It also applies user-specified filters and effects to the media information.
[0440] Step 4:
[0441] The server begins distributing optimized media information to the terminal. It adjusts the streaming speed according to the network conditions to ensure smooth transmission.
[0442] Step 5:
[0443] The terminal interprets data received from the server and reproduces media information on the display device. During display, it automatically adjusts screen brightness and contrast to provide the best possible visual experience.
[0444] Step 6:
[0445] The user performs gestures on the phone case. The device uses its built-in sensors to detect the gestures and sends that information to the server.
[0446] Step 7:
[0447] The server analyzes gesture data and generates operation instructions based on the user's intent. These include stopping playback, pausing, and changing effects. The server then sends the generated instructions back to the terminal.
[0448] Step 8:
[0449] The terminal controls media information on the display device based on operation instructions received from the server. It instantly reflects actions such as playback, pause, and effect changes, providing the user with real-time interaction.
[0450] (Example 1)
[0451] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0452] In recent years, there has been a growing demand for systems that dynamically display information and support user interaction. However, existing systems suffer from inefficient information optimization and management, making it difficult to achieve low-latency information display and real-time communication. Therefore, effective means are needed to provide users with an intuitive and engaging user experience.
[0453] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0454] In this invention, the server includes means for a user to register information with the device and display it dynamically, means for storing the information and managing it based on identification information, means for optimizing the information for the display area of the display device, and means for distributing the information using real-time communication technology. This enables a system that displays information with low latency and high quality, and that is interactively operable by the user.
[0455] "Information" refers to data that a user registers on their device and is processed and displayed in digital format.
[0456] "Device" refers to the collective term for hardware and software used for registering, storing, displaying, managing, and distributing information.
[0457] A "display device" is a device used to visually display information, and includes screens and displays.
[0458] "Motion recognition functionality" refers to a function that allows a device to detect physical actions such as user gestures and taps, and to process information and perform operations in response to those actions.
[0459] "Real-time communication technology" refers to a communication method that enables the instantaneous transmission and reception of data over a network, and its immediate reflection.
[0460] "Distribution" refers to the transmission of information to another device or user, and specifically to the transfer of data over a digital network.
[0461] This invention provides a system for dynamically displaying and interactively manipulating information selected by the user. The system consists primarily of a server and a terminal, which work together in cooperation with each other.
[0462] The server receives information sent from the user's terminal, associates it with identification information, and stores it in a database. Database software or cloud storage services can be used for this process. The server further performs optimization processes such as adjusting the resolution, cropping, and applying filters to the information. Image and video editing software is recommended for this optimization.
[0463] The terminal displays optimized information received from the server onto a display device. The display device utilizes high-quality screen technology, such as an OLED display. The terminal also has a built-in motion recognition function, allowing it to detect user input. For example, the terminal uses sensor technology to detect gestures and taps and sends instruction information to the server. Based on these instructions, the server sends commands such as play, pause, and add effects.
[0464] Users can upload their favorite videos and images to the server through the application. For example, a user can select a landscape image, apply a sepia filter to it, and display it on their device. The displayed content can be interactively managed through user gestures.
[0465] An example of a prompt message is, "Upload a landscape photo, apply a sepia filter, and view it on the terminal."
[0466] This system allows users to customize information according to their individual needs and enjoy a rich digital experience through intuitive operation.
[0467] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0468] Step 1:
[0469] The user selects information and uploads it to the server via their device. The user selects image or video files from their device via the app's interface and sends the selected data to the server by pressing the "Upload" button.
[0470] Input: User selects an image or video file
[0471] Output: Raw media data transferred to the server
[0472] Step 2:
[0473] The server stores the received information along with identification information in a database. The server saves the received files in a specific format (e.g., "User ID_File Name_Date and Time") and organizes them using database software.
[0474] Input: Raw media data sent from the device.
[0475] Output: Media data stored in the database
[0476] Step 3:
[0477] The server retrieves information from the database and performs optimization processes such as resolution adjustment, cropping, and filter application. Using Adobe software APIs, it standardizes the image resolution, crops a specified area, and applies the selected filter.
[0478] Input: Media data stored in the database
[0479] Output: Optimized media data
[0480] Step 4:
[0481] The server sends optimized information to the terminal using real-time communication technology. WebSocket is used to achieve low-latency data transfer.
[0482] Input: Optimized media data
[0483] Output: Media data sent to the terminal
[0484] Step 5:
[0485] The terminal displays the received information on the display device. The terminal uses an OLED display to show media in high quality and automatically adjusts the brightness using an ambient light sensor.
[0486] Input: Media data sent from the server
[0487] Output: Media displayed on the display device
[0488] Step 6:
[0489] Users interact with the displayed information by performing gestures on the device. The device uses sensors to detect actions such as taps and swipes and sends that data as instructions to the server.
[0490] Input: User gestures
[0491] Output: Instruction information sent to the server
[0492] Step 7:
[0493] Based on the received instructions, the server sends the necessary media control commands back to the terminal. This allows for actions such as playback, stopping, and adding effects.
[0494] Input: Instructions sent from the terminal
[0495] Output: Control commands to the terminal
[0496] (Application Example 1)
[0497] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0498] In today's commercial spaces, information provided to consumers is diversifying, but it is difficult for consumers to effectively and intuitively obtain the information they truly need. Especially for busy modern people, there is a need for personalized information delivery systems that allow quick access to necessary information. Furthermore, the sheer volume of information often overwhelms consumers, leading to confusion in their choices. A system is needed to improve this situation and realize efficient and personalized information delivery.
[0499] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0500] In this invention, the server includes means for a user to register media information with the device and dynamically display the media information, means for storing the media information on a storage medium and managing the media information based on user identification information, and means for detecting the user's gaze and actions using an eye-tracking device and displaying the media information in an individually customized manner. This makes it possible for consumers to obtain personalized information in real time based on their gaze and gestures within a store.
[0501] "User identification information" refers to information used to uniquely identify each user and manage associated media data.
[0502] "Media information" refers to media data such as images, videos, and audio registered by users.
[0503] A "display device" is a display device that dynamically displays media information registered by the user.
[0504] "Gesture recognition" is a function that detects user actions such as taps and swipes and controls media information in response to those actions.
[0505] An "eye-tracking device" is a device that includes sensors that detect the user's gaze and enable the display of information based on that gaze.
[0506] "Individually customized display using visual devices" refers to a function that displays media information tailored to individual users based on eye tracking and gesture recognition.
[0507] "Displaying with low latency" refers to a process that quickly reflects media information on the display device in real time, allowing for immediate response to user interaction.
[0508] A system implementing this invention primarily consists of three components: a server, a terminal (such as smart glasses), and a user.
[0509] The server has the function of receiving, managing, and optimizing media information registered by the user. Specifically, the server stores media information based on user identification information, adjusts the resolution to suit display on eye-tracking devices, and applies cropping and filters as needed.
[0510] The smart glasses, acting as the terminal, receive optimized media information from a server and display it in real time within the user's field of view. The smart glasses are equipped with eye-tracking sensors and gesture recognition functions, detecting the user's gaze and movements and changing the displayed content accordingly to provide personalized information.
[0511] Through this system, users can obtain real-time information and promotions for products they are interested in using their gaze and gestures within the store. When a user's gaze lingers at a certain point, detailed information related to that product is automatically displayed on the screen. For example, if a user looks at a specific product in a certain section, details and related videos for that product will be displayed.
[0512] By utilizing a generative AI model, the system enables real-time information presentation in response to a prompt question such as: How can media information be displayed and interaction be achieved when a user wearing smart glasses directs their gaze towards a specified product?
[0513] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0514] Step 1:
[0515] The user registers media information using a device. As input, the user selects media (images, videos, etc.) from smart glasses or another device and uploads it to the server. The server receives this information and stores it in its storage device along with the user's identification information. As output, the server stores media information linked to individual users.
[0516] Step 2:
[0517] The server optimizes the media information it receives. The server analyzes the stored media information as input. It performs data processing such as resolution adjustment, cropping, and filter application to ensure optimal display for visual devices. The optimized media information is generated as output and saved again.
[0518] Step 3:
[0519] The device obtains optimized media information from the server. As input, the user's smart glasses access the server and request the necessary information. Based on this request, the server sends the corresponding optimized media information as output to the device. The device receives this information and prepares the display.
[0520] Step 4:
[0521] The device displays media information using eye-tracking and gesture recognition. Input consists of the user's gaze and gestures, which the device captures using sensors. Based on the data calculations, the device performs appropriate media control based on the user's gaze direction and gesture content. As output, personalized media information is displayed on the screen in real time.
[0522] Step 5:
[0523] Based on user interaction, media information is dynamically controlled. Input includes user gestures and eye-tracking, which the terminal analyzes and sends instructions to the server. The server receives these instructions and performs actions such as starting, stopping, and adding effects to media playback. The output is media playback that meets the user's expectations.
[0524] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0525] This invention is a system that recognizes a user's emotions in real time and dynamically adjusts the display of media information and interactions based on the recognition results. This system consists of three main components: a server, a terminal, and an emotion engine.
[0526] The device is equipped with an emotion engine that recognizes emotions from the user's facial expressions and voice. The emotion engine analyzes emotional data in real time as the user uses the device and sends the results to a server. The server uses this emotional data to evaluate the user's state and determine the optimal selection and display method of media information.
[0527] The server manages the user's past emotional history based on emotional data and learns the emotional tendencies of individual users. Based on this information, the server can recommend appropriate media information to the user and prompt the user to make a selection.
[0528] The terminal receives instructions from the server and displays media information optimized for the display device. At the same time, it has the function to dynamically adjust effects and playback content according to the user's emotions. For example, if the user expresses the emotion of "joy," it may apply a bright filter to the video or add cheerful music.
[0529] As a concrete example, consider a scenario where a user views weather information as media information. If the emotion engine detects that the user is feeling gloomy because it's a rainy day, the server can use a transparent information overlay to display an additional video of a relaxing landscape. This enables personalized information delivery that aligns with the user's emotions.
[0530] This invention allows users to enjoy digital content experiences tailored to their individual emotional states, further enhancing the value of the device.
[0531] The following describes the processing flow.
[0532] Step 1:
[0533] The user operates the device and selects media information through the application. The device processes this selection information, activates an emotion engine, and captures the user's facial expressions and voice data in real time.
[0534] Step 2:
[0535] The emotion engine analyzes the captured data to identify the user's current emotions. This identification result is generated as an emotion tag within the device.
[0536] Step 3:
[0537] The device sends the generated emotion tag and selected media information to the server. The server receives the emotion tag and analyzes it along with the user's past emotion history.
[0538] Step 4:
[0539] The server determines media information and effects to recommend to the user based on emotion tags. For example, if the user indicates "joy," the server will recommend videos and music with a cheerful tone.
[0540] Step 5:
[0541] The server transmits the selected media information and additional effects to the terminal. The terminal reflects this information on the display device, providing the user with visual and auditory feedback.
[0542] Step 6:
[0543] The user interacts with the media information being displayed. The device, through its emotion engine, recaptures the user's facial expressions and gestures to detect changes in emotion.
[0544] Step 7:
[0545] The device sends newly detected emotion data to the server and awaits instructions to readjust the displayed content. The server processes this information and recommends updating to the optimal display content.
[0546] Step 8:
[0547] The device continuously updates the display of media information in response to instructions from the final server, providing the user with an experience that is always in line with their current emotions.
[0548] (Example 2)
[0549] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0550] Conventional information display systems had the problem of not being able to adjust content to take into account the user's emotional state, and thus failing to provide a personalized experience. Furthermore, static information content made it difficult to adapt in real time to the user's interests and emotions, resulting in users not being able to have a satisfying experience.
[0551] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0552] In this invention, the server includes means for a processing device that analyzes the user's nonverbal expressions to obtain their emotional state, means for dynamically adjusting the selected and displayed information content based on the emotional state, and means for generating and displaying information content corresponding to the user's emotional state using a generative AI model. This enables real-time and personalized information provision based on the user's emotions.
[0553] "Nonverbal expression" refers to means of communication that do not involve language, and includes gestures, facial expressions, tone of voice, and so on.
[0554] "Emotional state" refers to the user's psychological and emotional state, such as joy, sadness, or depression.
[0555] "Information content" is a general term for digital information provided to users, and includes text, images, audio, and video.
[0556] A "generative AI model" refers to an artificial intelligence model that uses machine learning algorithms to generate information tailored to a specific purpose from particular input data.
[0557] A "prompt sentence" is a sentence that serves as an input instruction to an AI model, and is intended to guide it to a specific output result.
[0558] A "processing device" is a set of hardware or software that receives data, analyzes it, and outputs the results.
[0559] This invention is a system that provides information content in real time according to the user's emotional state. The main components consist of a terminal, a server, and an emotion analysis engine.
[0560] The device is equipped with a camera and microphone to capture the user's facial expressions and voice. An emotion analysis engine analyzes this data and identifies the user's emotional state in real time from their nonverbal expressions. Deep learning technology can be used for specific analysis.
[0561] The analyzed sentiment data is sent to the server via an encrypted protocol. The server records the received sentiment data in a database and can learn the user's past sentiment history using machine learning algorithms. Based on this information, the server selects and provides information content that is appropriate for the user.
[0562] Furthermore, the server can utilize a generative AI model to generate new information content tailored to the user's emotional state. This generative AI model uses historical emotional data and real-time input data to customize the information best suited to the user.
[0563] The device displays informational content based on instructions from the server. During display, resolution adjustments and effects are added to optimize the user experience. For example, if the user expresses joy, the device enhances the content by applying a bright filter to the video and playing cheerful music.
[0564] As a concrete example, the prompt might include a message like, "When the user expresses feelings of joy, please display a positive news article." This prompt can be used to encourage the generative AI model to generate appropriate information.
[0565] In this way, users can receive personalized content that responds to their emotions, thereby improving the quality of their digital experience.
[0566] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0567] Step 1:
[0568] The device captures the user's facial expressions and voice using its camera and microphone. The input data consists of real-time video and audio data, and the device's emotion analysis engine analyzes this data to identify the user's emotional state. In this step, a deep learning model is used to perform data calculations that classify the emotional state into categories such as "joy" or "sadness." The output is data containing tags that indicate the user's emotional state.
[0569] Step 2:
[0570] The device sends emotional state data it has acquired to the server. The input data consists of tags indicating the emotional state. The server receives this data and analyzes it by cross-referencing it with past emotional history. Storage in the database and learning of the user's emotional tendencies take place here. The output is analytical data that reflects the user's emotional tendencies.
[0571] Step 3:
[0572] The server uses a generative AI model based on the user's emotional state and past history data to generate appropriate informational content. In this step, prompts are used to input into the generative AI model, generating content optimized for the user. The output is informational content that corresponds to the user's emotional state.
[0573] Step 4:
[0574] The server generates information content and sends it from the server to the terminal. The input data is the generated information content, which the terminal receives. The terminal adjusts the resolution and adds effects to display the information content on the screen. The output is the display data presented to the user in an optimized form.
[0575] Step 5:
[0576] Users view informational content through their devices and perform gestures and other interactions. The device utilizes gesture recognition to dynamically control the informational content (play, pause, add effects) in response to user input. The input data is the user's operation signal, and the output is the state of the content after the operation.
[0577] These steps allow users to receive personalized digital content that responds to their emotions.
[0578] (Application Example 2)
[0579] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0580] Modern information delivery systems require dynamic information delivery tailored to the user's emotional state, but conventional systems have limitations in providing personalized content that responds to each user's individual emotional state. This makes it difficult to provide information that resonates with users' emotions, contributing to decreased user satisfaction.
[0581] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0582] In this invention, the server includes means for a user to register information on a device and dynamically display the information on a display device; means for storing the information on a storage medium and managing the information based on user identification information; and means for analyzing the user's emotions and dynamically adjusting the information displayed based on those emotions. This makes it possible to provide personalized information that is tailored to the user's emotional state.
[0583] "User identification information" refers to information used to identify individual users, and plays a role in uniquely identifying users within the system.
[0584] "Emotional data" refers to information that indicates the emotional state of a user, analyzed from their facial expressions, voice, etc., and is data that specifically represents the user's psychological response.
[0585] "Gesture recognition functionality" is a feature that detects the user's physical movements and uses them as system input, and is a technology that contributes to improving the user interface.
[0586] "Resolution adjustment" is the process of changing the resolution of images or videos in order to optimize the display quality of information on a display device.
[0587] "Cropping" is a technique for editing images and videos by cutting out unnecessary parts to make them the appropriate size and configuration for the display screen.
[0588] A "filter" is a means of adding visual or auditory effects to a medium, and is a technique for improving the sensory quality of information.
[0589] "Personalized information delivery" is the process of providing information in the format most suitable for each user, based on data such as the user's past behavior and individual emotional state.
[0590] "Dynamic adjustment" is a concept that refers to changing the system's output or behavior on the fly in response to real-time data and circumstances.
[0591] The system that realizes this invention recognizes the user's emotions in real time, dynamically adjusts media information based on emotion data, and provides personalized information. The system mainly consists of a server, a terminal, and an emotion engine.
[0592] The device collects the user's facial expressions and voice, and analyzes this data using an emotion engine. The emotion engine uses AI technology to recognize the user's emotions in real time and sends this data to a server. Based on the received emotion data, the server learns the user's individual emotional tendencies and selects the most appropriate media information. In this process, the server also utilizes past emotional history stored on the storage medium to achieve more accurate personalization.
[0593] The server then sends media information selected to match the user's emotions to the terminal. The terminal optimizes this information for user interaction within the display area and performs processing such as resolution adjustment and filter application as needed. This system also features gesture recognition, allowing users to manipulate media information based on their actions.
[0594] For example, if the emotion engine detects that a user is feeling stressed, the system will dynamically display relaxation videos or music on the device to provide a relaxing effect. This creates a personalized entertainment experience tailored to the user's emotional state.
[0595] A concrete example of utilizing a generative AI model is a prompt such as, "The user's current emotion is 'surprise.' Please recommend a movie that is suitable for this emotion." This prompt allows the system to suggest content related to surprise, providing the user with the most relevant information.
[0596] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0597] Step 1:
[0598] The device captures the user's facial expressions and voice in real time. This data is collected using the camera and microphone and sent as input to the emotion engine. The emotion engine analyzes this input and outputs data indicating the user's emotional state.
[0599] Step 2:
[0600] The server receives sentiment data sent from the sentiment engine. Based on this data, the server references the user's past sentiment history and learns the user's sentiment tendencies. This prepares the server to select the most appropriate media information in the next step.
[0601] Step 3:
[0602] The server uses a generative AI model to select appropriate media information based on the user's current emotional state. In this process, emotional data is input to the generative AI model as a prompt, generating information such as, "The user's current emotion is 'surprise.' Please recommend a movie that is suitable for this emotion." The server then generates the optimal media information as output.
[0603] Step 4:
[0604] The server sends selected media information to the terminal. Upon receiving this information, the terminal optimizes it according to the display area. Specifically, it adjusts the resolution and applies filters to prepare the content for display to the user. This improves the user interface.
[0605] Step 5:
[0606] The device displays optimized media information to the user and is capable of recognizing the user's gestures. Through gesture recognition, the user can perform actions such as playing, pausing, and adding effects to the displayed information. This process allows the user to freely manipulate content that aligns with their emotions.
[0607] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0608] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0609] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0610] [Fourth Embodiment]
[0611] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0612] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0613] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0614] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0615] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0616] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0617] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0618] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0619] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0620] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0621] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0622] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0623] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0624] This invention is a system that dynamically displays user-specified media information on a display device such as a mobile phone case and interactively controls the media information in response to user operations. This system mainly consists of a server and terminals, and each component functions in cooperation with the others.
[0625] The server receives media information uploaded by users from their devices and stores it linked to user identification information. This allows for efficient management of multiple user data sets. The server also analyzes the received media information, adjusting the resolution, cropping, and applying filters. Through these processes, the media information is optimized for the display device on the mobile phone case.
[0626] Meanwhile, the device receives optimized media information delivered from the server. The device also has a function to actually display this information on the display device and automatically adjust the screen quality. Furthermore, devices with gesture recognition capabilities can detect user actions such as taps and swipes. Based on this, the device sends gesture instructions to the server and receives commands to control actions such as media playback, pause, and adding effects.
[0627] As a concrete example, let's consider a scenario where a user displays their favorite landscape video on their device and effectively controls it with specific gestures. The user uploads the selected video to the server using the app. The server trims the video, applies the desired filters, and then sends it to the device. The device displays the video and, upon receiving verbal instructions to stop playback, receives instructions from the server and controls the video according to the gestures.
[0628] This invention allows users to enjoy a new experience by customizing the surface of their mobile phone to match their emotions and needs at the moment. Such a system realizes a more intuitive and engaging user experience by reflecting individual sensibilities in the device.
[0629] The following describes the processing flow.
[0630] Step 1:
[0631] The user launches the application, selects the media information they want to display, and presses the upload button. The device processes the selected media information and compresses the data appropriately before sending it to the server.
[0632] Step 2:
[0633] The server receives compressed media data sent from the terminal. Next, it uses user identification information to save the received media data to the corresponding user profile.
[0634] Step 3:
[0635] The server processes the media information into a format suitable for the display device. Specifically, it adjusts the resolution to match the screen size of the mobile phone and crops the display area. It also applies user-specified filters and effects to the media information.
[0636] Step 4:
[0637] The server begins distributing optimized media information to the terminal. It adjusts the streaming speed according to the network conditions to ensure smooth transmission.
[0638] Step 5:
[0639] The terminal interprets data received from the server and reproduces media information on the display device. During display, it automatically adjusts screen brightness and contrast to provide the best possible visual experience.
[0640] Step 6:
[0641] The user performs gestures on the phone case. The device uses its built-in sensors to detect the gestures and sends that information to the server.
[0642] Step 7:
[0643] The server analyzes gesture data and generates operation instructions based on the user's intent. These include stopping playback, pausing, and changing effects. The server then sends the generated instructions back to the terminal.
[0644] Step 8:
[0645] The terminal controls media information on the display device based on operation instructions received from the server. It instantly reflects actions such as playback, pause, and effect changes, providing the user with real-time interaction.
[0646] (Example 1)
[0647] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0648] In recent years, there has been a growing demand for systems that dynamically display information and support user interaction. However, existing systems suffer from inefficient information optimization and management, making it difficult to achieve low-latency information display and real-time communication. Therefore, effective means are needed to provide users with an intuitive and engaging user experience.
[0649] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0650] In this invention, the server includes means for a user to register information with the device and display it dynamically, means for storing the information and managing it based on identification information, means for optimizing the information for the display area of the display device, and means for distributing the information using real-time communication technology. This enables a system that displays information with low latency and high quality, and that is interactively operable by the user.
[0651] "Information" refers to data that a user registers on their device and is processed and displayed in digital format.
[0652] "Device" refers to the collective term for hardware and software used for registering, storing, displaying, managing, and distributing information.
[0653] A "display device" is a device used to visually display information, and includes screens and displays.
[0654] "Motion recognition functionality" refers to a function that allows a device to detect physical actions such as user gestures and taps, and to process information and perform operations in response to those actions.
[0655] "Real-time communication technology" refers to a communication method that enables the instantaneous transmission and reception of data over a network, and its immediate reflection.
[0656] "Distribution" refers to the transmission of information to another device or user, and specifically to the transfer of data over a digital network.
[0657] This invention provides a system for dynamically displaying and interactively manipulating information selected by the user. The system consists primarily of a server and a terminal, which work together in cooperation with each other.
[0658] The server receives information sent from the user's terminal, associates it with identification information, and stores it in a database. Database software or cloud storage services can be used for this process. The server further performs optimization processes such as adjusting the resolution, cropping, and applying filters to the information. Image and video editing software is recommended for this optimization.
[0659] The terminal displays optimized information received from the server onto a display device. The display device utilizes high-quality screen technology, such as an OLED display. The terminal also has a built-in motion recognition function, allowing it to detect user input. For example, the terminal uses sensor technology to detect gestures and taps and sends instruction information to the server. Based on these instructions, the server sends commands such as play, pause, and add effects.
[0660] Users can upload their favorite videos and images to the server through the application. For example, a user can select a landscape image, apply a sepia filter to it, and display it on their device. The displayed content can be interactively managed through user gestures.
[0661] An example of a prompt message is, "Upload a landscape photo, apply a sepia filter, and view it on the terminal."
[0662] This system allows users to customize information according to their individual needs and enjoy a rich digital experience through intuitive operation.
[0663] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0664] Step 1:
[0665] The user selects information and uploads it to the server via their device. The user selects image or video files from their device via the app's interface and sends the selected data to the server by pressing the "Upload" button.
[0666] Input: User selects an image or video file
[0667] Output: Raw media data transferred to the server
[0668] Step 2:
[0669] The server stores the received information along with identification information in a database. The server saves the received files in a specific format (e.g., "User ID_File Name_Date and Time") and organizes them using database software.
[0670] Input: Raw media data sent from the device.
[0671] Output: Media data stored in the database
[0672] Step 3:
[0673] The server retrieves information from the database and performs optimization processes such as resolution adjustment, cropping, and filter application. Using Adobe software APIs, it standardizes the image resolution, crops a specified area, and applies the selected filter.
[0674] Input: Media data stored in the database
[0675] Output: Optimized media data
[0676] Step 4:
[0677] The server sends optimized information to the terminal using real-time communication technology. WebSocket is used to achieve low-latency data transfer.
[0678] Input: Optimized media data
[0679] Output: Media data sent to the terminal
[0680] Step 5:
[0681] The terminal displays the received information on the display device. The terminal uses an OLED display to show media in high quality and automatically adjusts the brightness using an ambient light sensor.
[0682] Input: Media data sent from the server
[0683] Output: Media displayed on the display device
[0684] Step 6:
[0685] Users interact with the displayed information by performing gestures on the device. The device uses sensors to detect actions such as taps and swipes and sends that data as instructions to the server.
[0686] Input: User gestures
[0687] Output: Instruction information sent to the server
[0688] Step 7:
[0689] Based on the received instructions, the server sends the necessary media control commands back to the terminal. This allows for actions such as playback, stopping, and adding effects.
[0690] Input: Instructions sent from the terminal
[0691] Output: Control commands to the terminal
[0692] (Application Example 1)
[0693] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0694] In today's commercial spaces, information provided to consumers is diversifying, but it is difficult for consumers to effectively and intuitively obtain the information they truly need. Especially for busy modern people, there is a need for personalized information delivery systems that allow quick access to necessary information. Furthermore, the sheer volume of information often overwhelms consumers, leading to confusion in their choices. A system is needed to improve this situation and realize efficient and personalized information delivery.
[0695] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0696] In this invention, the server includes means for a user to register media information with the device and dynamically display the media information, means for storing the media information on a storage medium and managing the media information based on user identification information, and means for detecting the user's gaze and actions using an eye-tracking device and displaying the media information in an individually customized manner. This makes it possible for consumers to obtain personalized information in real time based on their gaze and gestures within a store.
[0697] "User identification information" refers to information used to uniquely identify each user and manage associated media data.
[0698] "Media information" refers to media data such as images, videos, and audio registered by users.
[0699] A "display device" is a display device that dynamically displays media information registered by the user.
[0700] "Gesture recognition" is a function that detects user actions such as taps and swipes and controls media information in response to those actions.
[0701] An "eye-tracking device" is a device that includes sensors that detect the user's gaze and enable the display of information based on that gaze.
[0702] "Individually customized display using visual devices" refers to a function that displays media information tailored to individual users based on eye tracking and gesture recognition.
[0703] "Displaying with low latency" refers to a process that quickly reflects media information on the display device in real time, allowing for immediate response to user interaction.
[0704] A system implementing this invention primarily consists of three components: a server, a terminal (such as smart glasses), and a user.
[0705] The server has the function of receiving, managing, and optimizing media information registered by the user. Specifically, the server stores media information based on user identification information, adjusts the resolution to suit display on eye-tracking devices, and applies cropping and filters as needed.
[0706] The smart glasses, acting as the terminal, receive optimized media information from a server and display it in real time within the user's field of view. The smart glasses are equipped with eye-tracking sensors and gesture recognition functions, detecting the user's gaze and movements and changing the displayed content accordingly to provide personalized information.
[0707] Through this system, users can obtain real-time information and promotions for products they are interested in using their gaze and gestures within the store. When a user's gaze lingers at a certain point, detailed information related to that product is automatically displayed on the screen. For example, if a user looks at a specific product in a certain section, details and related videos for that product will be displayed.
[0708] By utilizing a generative AI model, the system enables real-time information presentation in response to a prompt question such as: How can media information be displayed and interaction be achieved when a user wearing smart glasses directs their gaze towards a specified product?
[0709] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0710] Step 1:
[0711] The user registers media information using a device. As input, the user selects media (images, videos, etc.) from smart glasses or another device and uploads it to the server. The server receives this information and stores it in its storage device along with the user's identification information. As output, the server stores media information linked to individual users.
[0712] Step 2:
[0713] The server optimizes the media information it receives. The server analyzes the stored media information as input. It performs data processing such as resolution adjustment, cropping, and filter application to ensure optimal display for visual devices. The optimized media information is generated as output and saved again.
[0714] Step 3:
[0715] The device obtains optimized media information from the server. As input, the user's smart glasses access the server and request the necessary information. Based on this request, the server sends the corresponding optimized media information as output to the device. The device receives this information and prepares the display.
[0716] Step 4:
[0717] The device displays media information using eye-tracking and gesture recognition. Input consists of the user's gaze and gestures, which the device captures using sensors. Based on the data calculations, the device performs appropriate media control based on the user's gaze direction and gesture content. As output, personalized media information is displayed on the screen in real time.
[0718] Step 5:
[0719] Based on user interaction, media information is dynamically controlled. Input includes user gestures and eye-tracking, which the terminal analyzes and sends instructions to the server. The server receives these instructions and performs actions such as starting, stopping, and adding effects to media playback. The output is media playback that meets the user's expectations.
[0720] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0721] This invention is a system that recognizes a user's emotions in real time and dynamically adjusts the display of media information and interactions based on the recognition results. This system consists of three main components: a server, a terminal, and an emotion engine.
[0722] The device is equipped with an emotion engine that recognizes emotions from the user's facial expressions and voice. The emotion engine analyzes emotional data in real time as the user uses the device and sends the results to a server. The server uses this emotional data to evaluate the user's state and determine the optimal selection and display method of media information.
[0723] The server manages the user's past emotional history based on emotional data and learns the emotional tendencies of individual users. Based on this information, the server can recommend appropriate media information to the user and prompt the user to make a selection.
[0724] The terminal receives instructions from the server and displays media information optimized for the display device. At the same time, it has the function to dynamically adjust effects and playback content according to the user's emotions. For example, if the user expresses the emotion of "joy," it may apply a bright filter to the video or add cheerful music.
[0725] As a concrete example, consider a scenario where a user views weather information as media information. If the emotion engine detects that the user is feeling gloomy because it's a rainy day, the server can use a transparent information overlay to display an additional video of a relaxing landscape. This enables personalized information delivery that aligns with the user's emotions.
[0726] This invention allows users to enjoy digital content experiences tailored to their individual emotional states, further enhancing the value of the device.
[0727] The following describes the processing flow.
[0728] Step 1:
[0729] The user operates the device and selects media information through the application. The device processes this selection information, activates an emotion engine, and captures the user's facial expressions and voice data in real time.
[0730] Step 2:
[0731] The emotion engine analyzes the captured data to identify the user's current emotions. This identification result is generated as an emotion tag within the device.
[0732] Step 3:
[0733] The device sends the generated emotion tag and selected media information to the server. The server receives the emotion tag and analyzes it along with the user's past emotion history.
[0734] Step 4:
[0735] The server determines media information and effects to recommend to the user based on emotion tags. For example, if the user indicates "joy," the server will recommend videos and music with a cheerful tone.
[0736] Step 5:
[0737] The server transmits the selected media information and additional effects to the terminal. The terminal reflects this information on the display device, providing the user with visual and auditory feedback.
[0738] Step 6:
[0739] The user interacts with the media information being displayed. The device, through its emotion engine, recaptures the user's facial expressions and gestures to detect changes in emotion.
[0740] Step 7:
[0741] The device sends newly detected emotion data to the server and awaits instructions to readjust the displayed content. The server processes this information and recommends updating to the optimal display content.
[0742] Step 8:
[0743] The device continuously updates the display of media information in response to instructions from the final server, providing the user with an experience that is always in line with their current emotions.
[0744] (Example 2)
[0745] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0746] Conventional information display systems had the problem of not being able to adjust content to take into account the user's emotional state, and thus failing to provide a personalized experience. Furthermore, static information content made it difficult to adapt in real time to the user's interests and emotions, resulting in users not being able to have a satisfying experience.
[0747] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0748] In this invention, the server includes means for a processing device that analyzes the user's nonverbal expressions to obtain their emotional state, means for dynamically adjusting the selected and displayed information content based on the emotional state, and means for generating and displaying information content corresponding to the user's emotional state using a generative AI model. This enables real-time and personalized information provision based on the user's emotions.
[0749] "Nonverbal expression" refers to means of communication that do not involve language, and includes gestures, facial expressions, tone of voice, and so on.
[0750] "Emotional state" refers to the user's psychological and emotional state, such as joy, sadness, or depression.
[0751] "Information content" is a general term for digital information provided to users, and includes text, images, audio, and video.
[0752] A "generative AI model" refers to an artificial intelligence model that uses machine learning algorithms to generate information tailored to a specific purpose from particular input data.
[0753] A "prompt sentence" is a sentence that serves as an input instruction to an AI model, and is intended to guide it to a specific output result.
[0754] A "processing device" is a set of hardware or software that receives data, analyzes it, and outputs the results.
[0755] This invention is a system that provides information content in real time according to the user's emotional state. The main components consist of a terminal, a server, and an emotion analysis engine.
[0756] The device is equipped with a camera and microphone to capture the user's facial expressions and voice. An emotion analysis engine analyzes this data and identifies the user's emotional state in real time from their nonverbal expressions. Deep learning technology can be used for specific analysis.
[0757] The analyzed sentiment data is sent to the server via an encrypted protocol. The server records the received sentiment data in a database and can learn the user's past sentiment history using machine learning algorithms. Based on this information, the server selects and provides information content that is appropriate for the user.
[0758] Furthermore, the server can utilize a generative AI model to generate new information content tailored to the user's emotional state. This generative AI model uses historical emotional data and real-time input data to customize the information best suited to the user.
[0759] The device displays informational content based on instructions from the server. During display, resolution adjustments and effects are added to optimize the user experience. For example, if the user expresses joy, the device enhances the content by applying a bright filter to the video and playing cheerful music.
[0760] As a concrete example, the prompt might include a message like, "When the user expresses feelings of joy, please display a positive news article." This prompt can be used to encourage the generative AI model to generate appropriate information.
[0761] In this way, users can receive personalized content that responds to their emotions, thereby improving the quality of their digital experience.
[0762] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0763] Step 1:
[0764] The device captures the user's facial expressions and voice using its camera and microphone. The input data consists of real-time video and audio data, and the device's emotion analysis engine analyzes this data to identify the user's emotional state. In this step, a deep learning model is used to perform data calculations that classify the emotional state into categories such as "joy" or "sadness." The output is data containing tags that indicate the user's emotional state.
[0765] Step 2:
[0766] The device sends emotional state data it has acquired to the server. The input data consists of tags indicating the emotional state. The server receives this data and analyzes it by cross-referencing it with past emotional history. Storage in the database and learning of the user's emotional tendencies take place here. The output is analytical data that reflects the user's emotional tendencies.
[0767] Step 3:
[0768] The server uses a generative AI model based on the user's emotional state and past history data to generate appropriate informational content. In this step, prompts are used to input into the generative AI model, generating content optimized for the user. The output is informational content that corresponds to the user's emotional state.
[0769] Step 4:
[0770] The server generates information content and sends it from the server to the terminal. The input data is the generated information content, which the terminal receives. The terminal adjusts the resolution and adds effects to display the information content on the screen. The output is the display data presented to the user in an optimized form.
[0771] Step 5:
[0772] Users view informational content through their devices and perform gestures and other interactions. The device utilizes gesture recognition to dynamically control the informational content (play, pause, add effects) in response to user input. The input data is the user's operation signal, and the output is the state of the content after the operation.
[0773] These steps allow users to receive personalized digital content that responds to their emotions.
[0774] (Application Example 2)
[0775] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0776] Modern information delivery systems require dynamic information delivery tailored to the user's emotional state, but conventional systems have limitations in providing personalized content that responds to each user's individual emotional state. This makes it difficult to provide information that resonates with users' emotions, contributing to decreased user satisfaction.
[0777] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0778] In this invention, the server includes means for a user to register information on a device and dynamically display the information on a display device; means for storing the information on a storage medium and managing the information based on user identification information; and means for analyzing the user's emotions and dynamically adjusting the information displayed based on those emotions. This makes it possible to provide personalized information that is tailored to the user's emotional state.
[0779] "User identification information" refers to information used to identify individual users, and plays a role in uniquely identifying users within the system.
[0780] "Emotional data" refers to information that indicates the emotional state of a user, analyzed from their facial expressions, voice, etc., and is data that specifically represents the user's psychological response.
[0781] "Gesture recognition functionality" is a feature that detects the user's physical movements and uses them as system input, and is a technology that contributes to improving the user interface.
[0782] "Resolution adjustment" is the process of changing the resolution of images or videos in order to optimize the display quality of information on a display device.
[0783] "Cropping" is a technique for editing images and videos by cutting out unnecessary parts to make them the appropriate size and configuration for the display screen.
[0784] A "filter" is a means of adding visual or auditory effects to a medium, and is a technique for improving the sensory quality of information.
[0785] "Personalized information delivery" is the process of providing information in the format most suitable for each user, based on data such as the user's past behavior and individual emotional state.
[0786] "Dynamic adjustment" is a concept that refers to changing the system's output or behavior on the fly in response to real-time data and circumstances.
[0787] The system that realizes this invention recognizes the user's emotions in real time, dynamically adjusts media information based on emotion data, and provides personalized information. The system mainly consists of a server, a terminal, and an emotion engine.
[0788] The device collects the user's facial expressions and voice, and analyzes this data using an emotion engine. The emotion engine uses AI technology to recognize the user's emotions in real time and sends this data to a server. Based on the received emotion data, the server learns the user's individual emotional tendencies and selects the most appropriate media information. In this process, the server also utilizes past emotional history stored on the storage medium to achieve more accurate personalization.
[0789] The server then sends media information selected to match the user's emotions to the terminal. The terminal optimizes this information for user interaction within the display area and performs processing such as resolution adjustment and filter application as needed. This system also features gesture recognition, allowing users to manipulate media information based on their actions.
[0790] For example, if the emotion engine detects that a user is feeling stressed, the system will dynamically display relaxation videos or music on the device to provide a relaxing effect. This creates a personalized entertainment experience tailored to the user's emotional state.
[0791] A concrete example of utilizing a generative AI model is a prompt such as, "The user's current emotion is 'surprise.' Please recommend a movie that is suitable for this emotion." This prompt allows the system to suggest content related to surprise, providing the user with the most relevant information.
[0792] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0793] Step 1:
[0794] The device captures the user's facial expressions and voice in real time. This data is collected using the camera and microphone and sent as input to the emotion engine. The emotion engine analyzes this input and outputs data indicating the user's emotional state.
[0795] Step 2:
[0796] The server receives sentiment data sent from the sentiment engine. Based on this data, the server references the user's past sentiment history and learns the user's sentiment tendencies. This prepares the server to select the most appropriate media information in the next step.
[0797] Step 3:
[0798] The server uses a generative AI model to select appropriate media information based on the user's current emotional state. In this process, emotional data is input to the generative AI model as a prompt, generating information such as, "The user's current emotion is 'surprise.' Please recommend a movie that is suitable for this emotion." The server then generates the optimal media information as output.
[0799] Step 4:
[0800] The server sends selected media information to the terminal. Upon receiving this information, the terminal optimizes it according to the display area. Specifically, it adjusts the resolution and applies filters to prepare the content for display to the user. This improves the user interface.
[0801] Step 5:
[0802] The device displays optimized media information to the user and is capable of recognizing the user's gestures. Through gesture recognition, the user can perform actions such as playing, pausing, and adding effects to the displayed information. This process allows the user to freely manipulate content that aligns with their emotions.
[0803] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0804] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0805] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0806] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0807] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0808] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0809] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0810] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0811] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0812] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0813] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0814] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0815] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0816] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0817] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0818] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0819] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0820] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0821] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0822] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0823] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0824] The following is further disclosed regarding the embodiments described above.
[0825] (Claim 1)
[0826] A means for a user to register media information with the device and to dynamically display said media information on a display device,
[0827] Means for storing the media information on a storage medium and managing the media information based on user identification information,
[0828] Means for applying resolution adjustment, cropping, and filtering to optimize the media information for the display area of the display device,
[0829] The terminal device is provided with a gesture recognition function and means for controlling the operation of the media information in accordance with the user's input actions,
[0830] A system that includes this.
[0831] (Claim 2)
[0832] The system according to claim 1, wherein the gesture recognition function includes means for detecting taps, swipes, or other gestures made by the user, and for playing, pausing, and adding effects to media information displayed based on the gesture.
[0833] (Claim 3)
[0834] The system according to claim 1, wherein the terminal device is equipped with means for adjusting the distribution speed of media information in order to display the media information on a display device with low latency.
[0835] "Example 1"
[0836] (Claim 1)
[0837] A means by which a user registers information with the device and dynamically displays the information on a display device,
[0838] Means for storing the information on a storage medium and managing the information based on identification information,
[0839] Means for applying image resolution adjustment, cropping, and filtering to optimize the information for the display area of the display device,
[0840] The terminal device is provided with an action recognition function and means for controlling the operation of the information in accordance with the user's input actions,
[0841] A means for distributing the information using real-time communication technology,
[0842] A system that includes this.
[0843] (Claim 2)
[0844] The system according to claim 1, wherein the motion recognition function includes means for detecting tactile input from a user and for playing, pausing, and adding effects to the information displayed based on the input.
[0845] (Claim 3)
[0846] The system according to claim 1, wherein the terminal device is equipped with means for adjusting the distribution speed of the information in order to display the information on a display device with low latency.
[0847] "Application Example 1"
[0848] (Claim 1)
[0849] A means for a user to register media information with the device and to dynamically display said media information on a display device,
[0850] Means for storing the media information on a storage medium and managing the media information based on user identification information,
[0851] Means for applying resolution adjustment, cropping, and filtering to optimize the media information for the display area of the display device,
[0852] The terminal device is provided with a gesture recognition function and means for controlling the operation of the media information in accordance with the user's input actions,
[0853] A means for detecting the user's gaze and movements using a visual device and displaying the media information in a way that is individually customized,
[0854] A system that includes this.
[0855] (Claim 2)
[0856] The system according to claim 1, wherein the gesture recognition function includes means for detecting taps, swipes, or other gestures made by the user, and for playing, pausing, and adding effects to media information displayed based on the gesture, and further includes means for selectively displaying the media information using the eye-tracking device.
[0857] (Claim 3)
[0858] The system according to claim 1, wherein the terminal device includes means for adjusting the distribution speed of media information in order to display the media information on a display device with low latency, and further includes means for updating the display content in real time by eye tracking.
[0859] "Example 2 of combining an emotion engine"
[0860] (Claim 1)
[0861] A means equipped with a processing device for analyzing the user's nonverbal expressions and obtaining their emotional state,
[0862] Means for dynamically adjusting the selected and displayed information content based on the emotional state,
[0863] A means for optimizing information content to the display characteristics of a display device using a solution structure,
[0864] A means equipped with a memory device for managing the user's emotional history and learning the user's emotional tendencies,
[0865] A means of recommending appropriate informational content to the user based on their emotional history,
[0866] A system that includes this.
[0867] (Claim 2)
[0868] The system according to claim 1, comprising means for generating and displaying information content corresponding to the user's emotional state using a generative AI model.
[0869] (Claim 3)
[0870] The system according to claim 1, comprising means for receiving a prompt message to acquire an emotional state and dynamically generating information content in response to the prompt message.
[0871] "Application example 2 when combining with an emotional engine"
[0872] (Claim 1)
[0873] A means by which a user registers information on a device and dynamically displays the information on a display device,
[0874] Means for storing the information on a storage medium and managing the information based on user identification information,
[0875] Means for applying resolution adjustment, cropping, and filtering to optimize the information for the display area of the display device,
[0876] The terminal device is provided with a gesture recognition function and means for controlling the operation of the information in accordance with the user's input actions,
[0877] A means for analyzing the user's emotions and dynamically adjusting the information displayed based on those emotions,
[0878] A means for making personalized suggestions based on sentiment data and displaying information based on those suggestions,
[0879] A system that includes this.
[0880] (Claim 2)
[0881] The system according to claim 1, wherein the gesture recognition function includes means for detecting taps, swipes, or other gestures made by the user, and for playing, pausing, and adding effects to information displayed based on the gesture.
[0882] (Claim 3)
[0883] The system according to claim 1, wherein the terminal device is equipped with means for adjusting the distribution speed of the information in order to display the information on a display device with low latency. [Explanation of symbols]
[0884] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for a user to register media information with the device and to dynamically display said media information on a display device, Means for storing the media information on a storage medium and managing the media information based on user identification information, Means for applying resolution adjustment, cropping, and filtering to optimize the media information for the display area of the display device, The terminal device is provided with a gesture recognition function and means for controlling the operation of the media information in accordance with the user's input actions, A system that includes this.
2. The system according to claim 1, wherein the gesture recognition function includes means for detecting taps, swipes, or other gestures made by the user, and for playing, pausing, and adding effects to media information displayed based on the gesture.
3. The system according to claim 1, wherein the terminal device is equipped with means for adjusting the distribution speed of media information in order to display the media information on a display device with low latency.