system

The system addresses the challenge of recording and reviewing daily life events by automatically analyzing image, audio, and location data to generate precise and chronological activity logs, improving accuracy through user feedback.

JP2026071024APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing methods for recording daily life events are time-consuming and lack a unified means to integrate various types of information, making it difficult to review past events effectively.

Method used

A system that automatically records daily life events by transmitting image, audio, and location data to a cloud server for analysis, identifying objects, people, and activities, and generating chronological activity logs, with continuous learning to improve analysis accuracy based on user feedback.

Benefits of technology

Provides an efficient and accurate means to record and review daily life events with minimal user effort, enhancing the precision of activity logs through continuous learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071024000001_ABST
    Figure 2026071024000001_ABST
Patent Text Reader

Abstract

This system automatically records the user's daily life by sending image data, audio data, and location information acquired by the device to a cloud server and analyzing this data using a generating AI. [Solution] A system comprising: means for recording image data acquired by a camera on a terminal; means for recording audio data acquired by a voice acquisition device on a terminal; means for acquiring and recording location information using a location detection device; means for transmitting the image data, audio data and location information to a cloud server; means for the cloud server to analyze the received image data and identify an object or person; means for the cloud server to analyze the received audio data and identify a specific event or activity; means for the cloud server to analyze the received location information and identify a location; and means for generating and providing a user activity record based on the analysis results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern times, many people want to record their daily lives in detail, but there is a problem that manually keeping a diary is time-consuming and difficult to continue. In addition, since various types of information (images, voices, location information) exist separately, there is currently a lack of means to confirm them in a unified manner. As a result, it has become difficult to look back on past events, and the needs of users for recording cannot be met.

Means for Solving the Problems

[0005] This invention provides a system that automatically records a user's daily life by transmitting image data, audio data, and location information acquired by a terminal to a cloud server and analyzing this data using a generating AI. Based on the received data, the cloud server identifies objects and people, identifies events and activities, identifies locations, and generates a record of the user's activities. The generated activity record is organized chronologically and made viewable through a user interface, allowing the user to easily review past events without effort. Furthermore, the system aims to improve the precision of the records by performing learning processing to improve the analysis accuracy based on user feedback.

[0006] A "terminal" is a device that a user carries or uses, and that has the function of acquiring, recording, and transmitting images, audio, and location information.

[0007] "Image capture device" refers to a device or function for acquiring image data using an optical sensor, and includes camera functions.

[0008] A "speech acquisition device" is a device used to capture speech data, and a microphone falls into this category.

[0009] A "location detection device" is a function or device that uses GPS or other positioning technologies to determine the current location and collect location information.

[0010] A "cloud server" refers to a group of remote computer servers used to manage, analyze, and store data via the internet, and is responsible for processing data sent from user terminals.

[0011] "Generative AI" refers to an algorithm or system that uses artificial intelligence technology to analyze data and generate new information from images, audio, location information, etc.

[0012] "Image data" refers to visual information collected by a camera or camera, and is expressed in still image or video format.

[0013] "Audio data" refers to sound information recorded by an audio acquisition device, and includes human conversations, ambient sounds, and so on.

[0014] "Location information" refers to data indicating the geographical location of a device, acquired by a location detection device, and is expressed in digital format such as latitude, longitude, and altitude.

[0015] "Analysis" refers to the process of handling data, understanding its content, and extracting important information.

[0016] "Activity logs" are digital records of user activities generated from the analysis of images, audio, and location data. [Brief explanation of the drawing]

[0017] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] Shows an emotion map to which a plurality of emotions are mapped. [Figure 10] Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.

[0019] First, the terms used in the following description will be described.

[0020] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be a single arithmetic device or a combination of a plurality of arithmetic devices. Also, the processor may be a single type of arithmetic device or a combination of a plurality of types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0021] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] This invention relates to a system that automatically records a user's daily life and allows them to easily review it. Specifically, a smart device (hereinafter referred to as "terminal") collects data using various sensors, transmits it to a cloud server for analysis, and records various aspects of the user's life.

[0039] The device continuously collects images taken by the user, audio recordings, and location information. For example, if a user visits a historical site that is a tourist destination, the device records photos of the scene and uses GPS location information to pinpoint the exact location visited. Also, when a user is talking to a friend, the device captures that conversation as audio data.

[0040] The data collected by the device is transmitted to a cloud server via network communication. The cloud server analyzes the received data using advanced generative AI technology, recognizing objects and people from image data, and transcribing conversations from audio data to extract key keywords. Location information is cross-referenced with map data to identify specific facilities and landmarks.

[0041] The server generates information to create a user activity log. This log clearly presents the user's daily activities in chronological order, with important events described in detail. Users can view this log through the application, for example, seeing a record such as "On [Month] [Day], had lunch with an acquaintance at a tourist spot." In this way, users can easily review past events without relying on their memory.

[0042] Furthermore, the accuracy of the analysis is improved through user feedback. Based on corrections and additional information from users, the generated AI model is continuously updated to ensure that more accurate records are reflected in the future.

[0043] In this way, this system provides a means to automatically record daily life with minimal burden on the user and later offer it as a valuable memory.

[0044] The following describes the processing flow.

[0045] Step 1:

[0046] The device collects data using sensors. The device acquires images taken by the user through its camera, records audio with its microphone, and obtains location information using its GPS sensor.

[0047] Step 2:

[0048] The device sends the collected data to the cloud server. At regular intervals, or when sufficient data has accumulated, the device uploads encrypted image data, audio data, and location information to the cloud server via the network.

[0049] Step 3:

[0050] The server analyzes the image data. A generating AI on the cloud server analyzes the image data and identifies the objects and people depicted. For example, it performs facial recognition to identify acquaintances if they are in the image.

[0051] Step 4:

[0052] The server analyzes the audio data. The generative AI converts the audio data into text and extracts important keywords and events from the conversation. This allows it to identify what kind of conversation took place.

[0053] Step 5:

[0054] The server analyzes location information. Using GPS data and comparing it with a map database, it identifies the geographical locations the user has visited. Based on the location information, it recognizes visits to specific facilities or landmarks.

[0055] Step 6:

[0056] The server generates a record of the user's activities. Based on the collected analytical data, it organizes the user's daily activities and generates a chronological activity log. Particularly important events and frequently visited places are described in detail.

[0057] Step 7:

[0058] The server provides the user with a recorded activity log. The user can then open the smartphone application to visually review their daily activity log. This allows the user to easily review the places they visited and the people they met on a specific date.

[0059] (Example 1)

[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0061] In daily life, it is not easy to recall past events without relying on memory, and this is especially difficult in today's information-saturated society. Furthermore, there is a need for technology to efficiently extract and analyze useful information from collected data. Additionally, there is a lack of mechanisms to effectively utilize user feedback to improve analytical accuracy.

[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0063] In this invention, the server includes means for analyzing received image data to identify an object or person, means for analyzing received audio data to identify a specific event or activity, and means for analyzing received location information to identify a place. This makes it possible to record the user's daily life in detail, efficiently review past events, and improve analysis accuracy by utilizing feedback.

[0064] A "terminal" is an electronic device used to collect data from a user's daily life, and includes a camera, an audio acquisition device, and a location detection device.

[0065] A "shooting device" refers to a device used to acquire still images or videos, and is equipped with camera functionality.

[0066] A "sound acquisition device" is a device that records ambient sounds and acquires them as data.

[0067] A "location detection device" is a device that uses GPS or other location information technologies to determine the geographical location of a terminal.

[0068] A "cloud server" is a computing device that receives, analyzes, and stores data via a network, and shares information with other devices.

[0069] A "generative AI model" is an algorithm that uses machine learning techniques to analyze data and extract certain patterns or information.

[0070] "Feedback" refers to additional information or corrections that users provide to the system, which are used to improve analysis accuracy and enhance the system itself.

[0071] "Activity logs" are detailed records of a user's daily activities in chronological order, generated by analyzing user data.

[0072] A "user interface" is a component that provides a visual or manipulative means for a user to access and manipulate records and data.

[0073] A "prompt message" is a sentence that functions as specific instructions or guidelines from the user to support the processing of a generative AI model.

[0074] This invention is a system that allows users to efficiently record their daily lives and easily review them. It mainly consists of a terminal and a cloud server.

[0075] Terminal configuration and operation

[0076] The device collects data such as photos, audio, and location information from everyday life. Specifically, it takes photos using the camera, records audio using the microphone, and obtains location information using the GPS function. These devices are built into the device and automatically acquire data according to the user's movements. For example, if a user is having a picnic in a park on a holiday, the device will take photos of the surrounding scenery, record conversations, and obtain the user's location at that time.

[0077] Data transmission and analysis

[0078] The collected data is transmitted from the device to a cloud server via the network. The server receives this data and analyzes it using a generative AI model. Objects and people are identified from image data, and conversation content is converted to text and keywords extracted from audio data. Location data is used with a map service to identify specific facilities and locations. This analysis result is then generated as a user activity record.

[0079] Activity log and feedback

[0080] The activity logs generated by the server are organized chronologically, and users can view them through a dedicated user interface. These logs include detailed information about specific activities, such as "had a picnic in the park on [date]." Furthermore, the generating AI model continuously learns and improves analysis accuracy as users correct the logs through feedback. For example, adding details such as "had a conversation with a friend in the playground area" will provide more accurate activity logs in the future.

[0081] Specific examples and prompt statements

[0082] As a concrete example, a prompt message such as, "Analyze the photos taken and audio recorded over the weekend and create a record of past activities," is provided. This prompt message serves as a guideline to instruct the server in order to perform the analysis efficiently.

[0083] In this way, the system provides users with an effective and convenient means of recording their daily lives.

[0084] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0085] Step 1:

[0086] The device uses its camera to capture images of the user's surroundings and its microphone to record ambient sounds. It also uses GPS to obtain the user's current location. This step yields three types of input: image data, audio data, and location data. These data are temporarily stored in local storage.

[0087] Step 2:

[0088] The device transmits collected image data, audio data, and location data to a cloud server via the internet. During this process, the data is encrypted and protected. The output is the data transferred to the cloud server.

[0089] Step 3:

[0090] The server analyzes the received image data using a generative AI model. Based on the image's pixel information, it applies an object detection algorithm to identify specific objects or people. The output is the recognized object and its attribute information.

[0091] Step 4:

[0092] The server converts the received audio data into text using speech recognition technology. Furthermore, it analyzes the text content using natural language processing technology, extracting key keywords and phrases. The output consists of the analyzed text information and the extracted keywords.

[0093] Step 5:

[0094] The server compares the received location data with a map API to identify specific facilities or locations. It then utilizes a geographic information system to identify related facility names and landmarks, providing them as output.

[0095] Step 6:

[0096] The server integrates the analysis results from previous sessions and organizes the user's activity log chronologically. This activity log details important events and places visited. The output is an activity log for users to reflect on their daily lives.

[0097] Step 7:

[0098] Users review their activity logs through the application and provide feedback as needed. The server updates the generated AI model based on user corrections and additional information, improving analysis accuracy. The output represents the next accuracy improvement based on the updated AI model.

[0099] In this way, specific actions such as collection, transmission, analysis, and integration are performed at each step, providing users with accurate and detailed records of their activities.

[0100] (Application Example 1)

[0101] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0102] Traditional shopping experiences have faced challenges in providing personalized customer service online. In particular, the provision of personalized information based on customer purchase history and behavior has been insufficient, making it difficult to optimize the shopping experience. Furthermore, there is a need to enhance customer immersion in virtual stores compared to physical stores.

[0103] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0104] In this invention, the server includes means for analyzing image information acquired by the terminal and identifying items related to the purchase, means for analyzing voice information and identifying important keywords, and means for analyzing the customer's behavioral history using location information. This makes it possible to provide a personalized purchasing experience based on the customer's behavioral history.

[0105] A "terminal" is an electronic device that collects and records data when carried or worn by a user.

[0106] A "photography device" is a device that includes a camera or other equipment for acquiring image information.

[0107] A "speech acquisition device" is a device that includes a microphone or other equipment for recording speech information.

[0108] A "location detection device" is a device that includes GPS or similar technologies for measuring the user's current location.

[0109] A "remote data processing device" is a device that performs data analysis and management on the cloud or on a server.

[0110] "Image information" refers to visual data collected by a camera or other imaging device.

[0111] "Audio information" refers to audio data acquired by an audio acquisition device.

[0112] "Location information" refers to geographical coordinate data acquired by a location detection device.

[0113] "Action history" refers to data that records a user's past actions and events and organizes them chronologically.

[0114] "Evaluation information" refers to opinions and feedback provided by users, and is used to improve the accuracy of analysis.

[0115] To implement this invention, the system mainly consists of a terminal and a remote data processing device. The terminal is equipped with a camera, an audio acquisition device, and a location detection device, and collects data in the user's daily life. Specifically, the terminal's camera acquires image information of the user's surroundings, and the audio acquisition device records audio information such as conversations and ambient sounds. The location detection device tracks the user's movement path and acquires location information.

[0116] The data collected by the device is transmitted via network communication to a remote data processing device (cloud server). The server analyzes the image and audio information using Google® Cloud Vision API and Google Cloud Speech-to-Text API. From the image information, objects and features are identified, and specific keywords are extracted from the audio information. Location information is analyzed as geographic coordinate data and recorded on the server as part of the user's activity history.

[0117] The server generates user behavior records using an AI model based on the analyzed data. This allows users to receive personalized information based on their past purchasing behavior and places they have visited. For example, users may receive information about new products related to items they have previously purchased at stores they have visited.

[0118] As an example of a prompt message, we use the format: "Customer ID: 12345 tried on a specific category item at store: brand name. Please suggest relevant new products and promotional information." This allows the system to suggest the most relevant information to the user.

[0119] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0120] Step 1:

[0121] The device uses a camera to acquire image information of the user's surroundings. The input is the user's current visual environment, and the output is the acquired image data. This data is stored on a smart device worn by the user.

[0122] Step 2:

[0123] The terminal uses a voice acquisition device to record sounds around the user. The input is the surrounding sound environment, and the output is audio data. This audio data includes the user's conversation and ambient sounds.

[0124] Step 3:

[0125] Using a location detection device, the terminal obtains the user's location information. The input is the current geographical coordinates, and the output is location data. This location data is used to track the user's movement path.

[0126] Step 4:

[0127] The device transmits collected image data, audio data, and location information to a cloud server. The input is this data, and the output is the status indicating successful transmission to the server. The data is transmitted via network communication.

[0128] Step 5:

[0129] The server analyzes the received image information using the Google Cloud Vision API. The input is image data, and the output is identified object and feature data. This analysis identifies specific objects and attributes.

[0130] Step 6:

[0131] The server uses the Google Cloud Speech-to-Text API to analyze audio information and extract specific keywords. The input is audio data, and the output is the extracted text and keywords. The audio content is converted into text data, and important phrases are highlighted.

[0132] Step 7:

[0133] Based on location information, the server analyzes the user's movement history to identify specific locations. The input is location data, and the output is the result of analyzing the user's movement path. This analysis creates movement patterns based on the locations and routes the user has visited.

[0134] Step 8:

[0135] The server uses a generative AI model to generate personalized user behavior records based on the analyzed data. Inputs include analyzed images, audio, and location information, while output is personalized information tailored to the user. This generated information is based on the user's past purchasing behavior.

[0136] Step 9:

[0137] The user receives generated information using prompts and engages in a personalized purchasing experience. The input is prompts from the server, and the output is the user's purchasing behavior. The user makes a purchase decision based on the information presented.

[0138] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0139] This invention relates to a system that automatically records and analyzes a user's daily life to understand their emotional state, thereby providing the user with a richer reflective experience. This system includes an emotion engine that analyzes data collected by the terminal on a cloud server and reflects the user's emotions in the generated behavioral records.

[0140] The device functions as an integral part of the user's life, collecting image and audio data using its camera and microphone, and obtaining location information using its GPS function. This data collection is performed in the background by the device, so the user does not need to take any special action.

[0141] The collected data is sent to a cloud server via the network. On the cloud server, a generative AI analyzes this data, identifying objects and people, and extracting conversation content from audio data. In addition, an emotion engine analyzes image and audio data to recognize the user's emotional state. For example, based on an image, a smile might be interpreted as "joy," and emotions such as tension or anger might be detected from the tone of voice.

[0142] Based on these analysis results, the server generates a daily activity log of the user and integrates emotional information obtained from the emotion engine. This allows users to understand not only where and what happened, but also how they were feeling at the time. For example, a record might be generated stating, "I had a pleasant conversation at dinner with friends and felt happy."

[0143] Users can view these behavioral records through the application and reflect on past events along with their emotions. The system can improve the accuracy of its emotion engine through user feedback, resulting in more accurate and empathetic records.

[0144] In this way, this system enriches users' memories emotionally and provides a means to understand individual moments at a deeper level.

[0145] The following describes the processing flow.

[0146] Step 1:

[0147] The device collects data. It uses a camera to acquire image data and a microphone to record audio data. It also collects location information through its GPS function. This data is automatically acquired during daily life, without requiring any specific user action.

[0148] Step 2:

[0149] The device sends data to the cloud server. The collected image data, audio data, and location information are transferred to the cloud server in an encrypted state at regular intervals. This ensures that the data is securely managed and awaits analysis.

[0150] Step 3:

[0151] The server analyzes the image data. The generating AI identifies objects and people in the image and further infers the user's emotional state through facial expression analysis. For example, if the user is smiling in the photo, it will estimate "joy."

[0152] Step 4:

[0153] The server analyzes the audio data. Using speech recognition technology, it converts the recorded audio into text and extracts the content of the conversation. At the same time, it detects emotions such as tension, excitement, and anger through speech tone analysis.

[0154] Step 5:

[0155] The server analyzes the location information. It compares GPS data with map information to identify the places the user has visited. This allows for a detailed record of the areas and facilities the user has been in.

[0156] Step 6:

[0157] The server generates activity logs. Based on the analysis results, it compiles the user's daily records into an activity history. By integrating the output of the emotion engine, it generates a rich record that includes the user's emotional state.

[0158] Step 7:

[0159] The server displays the generated activity log in the user interface. Users can freely access this information through the app and reflect on their past events and emotions. Through the feedback function, users can provide information to further improve the accuracy of the record.

[0160] (Example 2)

[0161] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0162] Conventional recording systems often simply record events in a user's daily life based only on time and place, making it difficult to generate records that reflect the emotional state at the time. As a result, important emotional information is often missing when users reflect deeply on past events, leading to a challenge in obtaining a rich reflective experience.

[0163] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0164] In this invention, the server includes an analysis means for recognizing emotional states based on visual and audio data, a means for generating user behavior records and providing them in an integrated manner, and a means for performing learning processing to improve analysis accuracy based on evaluation information provided by the user. This enables users to reflect on past events accompanied by emotions and obtain a richer recording experience.

[0165] A "terminal" is a device that acquires visual data, audio data, and location information in a user's daily life.

[0166] An "image acquisition device" is a device used to collect visual data, and this mainly refers to a camera.

[0167] A "speech acquisition device" is a device used to collect speech data, and this mainly refers to a microphone.

[0168] A "location detection device" is a device used to measure and record a user's location information, and is primarily equipped with GPS functionality.

[0169] An "information processing device" is a computer system that analyzes received data and generates a record of user behavior.

[0170] "Visual data" refers to image information acquired by an image acquisition device.

[0171] "Audio data" refers to sound information acquired by an audio acquisition device.

[0172] "Emotional state" refers to the user's emotional response, analyzed based on visual and audio data.

[0173] "Activity logs" are records of a user's past activities and events, and may also include emotional information.

[0174] "Evaluation information" refers to user-provided feedback, which is used to improve the accuracy of the analysis.

[0175] "Learning process" is a process performed to improve the accuracy of the analysis algorithm based on evaluation information.

[0176] This invention is a system that records a user's daily life from both visual and auditory perspectives, analyzes these records, and adds emotional information. Specifically, a terminal collects visual data, audio data, and location information, and transmits them to an information processing device. The information processing device is equipped with means to perform data analysis using a generative AI model and identify the user's emotional state.

[0177] The device functions as a smartphone or wearable device and is equipped with image acquisition and audio acquisition devices. Specifically, it utilizes a camera for image acquisition and a microphone for audio acquisition. Location information is acquired using GPS functionality. This data is collected in the background by the device, and no special action is required from the user.

[0178] The collected data is securely encrypted and transmitted to the information processing device. The information processing device analyzes the received data using "OpenCV" as image analysis software and "Google Cloud Speech-to-Text" as speech analysis software. This allows it to identify objects and faces from visual data and extract conversations and activity content from audio data. Based on the information obtained through this data processing, the emotion engine determines the user's emotional state. For example, if an image contains a smile, it identifies the emotional state as "joy," and if the tone of voice is elevated, it identifies the emotional state as "excitement."

[0179] The server generates user behavior records based on the analysis results and provides them to the user in an integrated form with emotional information. These behavior records are accessible to the user through the application. A learning process is also performed to improve analysis accuracy based on evaluation information, and user feedback is incorporated to provide more accurate and empathetic records.

[0180] As a concrete example, the prompt message is as follows: "Analyze the emotional state during yesterday's event and generate a record of the behavior." This allows the user to obtain a record that includes detailed emotional information about past events.

[0181] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0182] Step 1:

[0183] The device collects visual data, audio data, and location information. Inputs include real-time visual and audio information from the user's surroundings, and location information obtained via GPS. The device captures images with its camera, records audio with its microphone, and simultaneously acquires its current location. The output is a dataset combining all of this data.

[0184] Step 2:

[0185] The terminal transmits the collected data to the information processing device. The input is a dataset of visual data, audio data, and location information generated in step 1. The terminal securely encrypts the data and transmits it to the information processing device over the network. The output is the encrypted dataset received by the information processing device.

[0186] Step 3:

[0187] The server acts as an information processing device, analyzing the received data. The input consists of datasets of visual data, audio data, and location information transmitted from the terminal. Using a generative AI model, the server first analyzes the visual data with image analysis software to identify objects and people. The audio data is analyzed using an audio analysis program to extract conversations and activity content. The output is the analysis result, including object information, person information, and audio content.

[0188] Step 4:

[0189] The server uses the analysis results to recognize the emotional state. The input consists of object information, person information, and audio content obtained in step 3. The server's emotion engine estimates the user's emotional state based on facial expression information from the visual data and tone information from the audio. For example, if there are many smiles, the emotion of "joy" will be recognized. The output is detailed analysis data with the estimated emotional state added.

[0190] Step 5:

[0191] The server generates user behavior records and integrates emotional information. The input is detailed analytical data including emotional states. Based on this analytical data, the server compiles behavior records chronologically and adds emotional information. The output is the user's behavior record, including emotional information. This record is converted into a format viewable by the user's application.

[0192] Step 6:

[0193] The user views their behavioral records through the application and provides feedback as needed. The input is the behavioral records generated in step 5. The user can access this to reflect on past events and emotions. The user can also provide feedback to improve the accuracy of emotion recognition. The output is the user's feedback information. This feedback is used in the system's learning process.

[0194] (Application Example 2)

[0195] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0196] In modern retail, it is difficult to understand customer purchasing behavior and emotions in real time, which presents a challenge in providing appropriate sales approaches. Furthermore, traditional methods make it difficult to efficiently detect customer interests and dissatisfactions and adjust sales strategies immediately, thus failing to maximize customer satisfaction.

[0197] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0198] In this invention, the server includes means for recording image data acquired by a terminal using a camera, means for recording audio data acquired by a terminal using an audio acquisition device, means for acquiring and recording location information using a location detection device, means for estimating the emotional state of a customer from their image and audio data and providing it in real time, and means for identifying the customer's interests in the store and adjusting the sales strategy accordingly. This enables real-time understanding of the customer's emotional state and the development of flexible sales strategies based on this understanding.

[0199] A "terminal" is a device that acquires image data, audio data, and location information of users and customers.

[0200] A "photography device" is a device used to acquire image data and functions as a camera.

[0201] A "voice acquisition device" is a device for recording voice data and functions as a microphone.

[0202] A "location detection device" is a device that acquires and records a user's location information and functions as a GPS.

[0203] A "cloud server" is a server that receives data sent from a terminal and performs analysis on it.

[0204] "Image data" refers to visual information acquired by a camera or imaging device.

[0205] "Audio data" refers to auditory information acquired by an audio acquisition device.

[0206] "Location information" refers to information indicating a geographical location, obtained by a location detection device.

[0207] "Emotional state" refers to the user's psychological state, which is estimated by analyzing image and audio data.

[0208] "Activity log" refers to a record of a user's daily activities, generated based on analyzed data.

[0209] "Sales strategy" refers to the plans and methods used to effectively sell products in physical stores.

[0210] "Analytical accuracy" refers to the degree to which the results obtained through data analysis are correct.

[0211] This invention provides a system that analyzes customer purchasing behavior and emotional states in physical stores in real time and deploys appropriate sales strategies in real time. This system consists of terminals, a cloud server, and various analysis software.

[0212] The device is a pair of smart glasses worn by the user (customer) and is equipped with a camera and microphone that function as both a camera and an audio acquisition device. This device also has GPS functionality as a location detection device, allowing it to pinpoint the customer's location within the store. This enables the continuous collection of image data, audio data, and location information, which are then transmitted to a cloud server without requiring any user interaction.

[0213] The cloud server is equipped with an image recognition model and a voice analysis model using TENSORFLOW®, which identify objects and people from received image data. It also analyzes voice data to extract conversation content and uses a generative AI model to estimate the customer's emotional state. The estimated emotional state is provided to field service staff in real time in the form of joy, interest, anxiety, etc., and is used to assist customers in the store.

[0214] Based on these data analysis results, the server generates behavioral records to identify customer interests and dissatisfactions. For example, if it is estimated that a customer is standing in front of a shelf for a long time, smiling and looking at a product, this information is immediately notified to a store employee, who can then approach the customer individually. A prompt message might be, "How to build a system that analyzes a customer's facial expression and tone of voice when they pick up a new product using smart glasses, clarifies their emotional state using an emotion engine, and provides the optimal sales approach." This prompt clarifies the analysis procedure within the system, leading to more appropriate customer service.

[0215] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0216] Step 1:

[0217] The device acquires image and audio data of the customer's surroundings through smart glasses. This records the customer's visual and auditory environment within the store. Image data includes store merchandise and the customer's facial expressions, while audio data includes the customer's tone of voice and conversation content. The input is real-world visual and auditory data, and the output is digital image and audio data.

[0218] Step 2:

[0219] The terminal transmits the acquired image and audio data to the cloud server via Wi-Fi. During this process, location information generated by the location detection device is also transmitted. The input from the terminal consists of images, audio, and location information, while the output is the arrival of these data packets to the server.

[0220] Step 3:

[0221] The server analyzes the received image data using TensorFlow to identify objects or people. The analysis uses a deep learning model to extract features from the image and match them with known people or objects. The input is image data sent from the terminal, and the output is a list of identified objects or people.

[0222] Step 4:

[0223] The server analyzes the voice data and uses a generative AI model to estimate the customer's emotional state. The voice analysis engine analyzes the tone, pitch, and volume of the voice to identify emotions. The input is voice data, and the output is recorded as the detected emotional state.

[0224] Step 5:

[0225] The server identifies the customer's location within the store based on the acquired location information. It then analyzes the customer's behavior in specific areas based on the location information to identify their interests. The input is location information, and the output is the customer's precise location within the store.

[0226] Step 6:

[0227] The user (store clerk) receives analysis results from the server in real time and uses that information to interact with customers. The system outputs information about the customer's emotional state and interests to their smart device, leading to appropriate sales responses. The input is the analysis results from the server, and the output is the specific actions taken by the store clerk based on that information.

[0228] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0229] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0230] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0231] [Second Embodiment]

[0232] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0233] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0234] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0235] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0236] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0237] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0238] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0239] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0240] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0241] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0242] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0243] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0244] This invention relates to a system that automatically records a user's daily life and allows them to easily review it. Specifically, a smart device (hereinafter referred to as "terminal") collects data using various sensors, transmits it to a cloud server for analysis, and records various aspects of the user's life.

[0245] The device continuously collects images taken by the user, audio recordings, and location information. For example, if a user visits a historical site that is a tourist destination, the device records photos of the scene and uses GPS location information to pinpoint the exact location visited. Also, when a user is talking to a friend, the device captures that conversation as audio data.

[0246] The data collected by the device is transmitted to a cloud server via network communication. The cloud server analyzes the received data using advanced generative AI technology, recognizing objects and people from image data, and transcribing conversations from audio data to extract key keywords. Location information is cross-referenced with map data to identify specific facilities and landmarks.

[0247] The server generates information to create a user activity log. This log clearly presents the user's daily activities in chronological order, with important events described in detail. Users can view this log through the application, for example, seeing a record such as "On [Month] [Day], had lunch with an acquaintance at a tourist spot." In this way, users can easily review past events without relying on their memory.

[0248] Furthermore, the accuracy of the analysis is improved through user feedback. Based on corrections and additional information from users, the generated AI model is continuously updated to ensure that more accurate records are reflected in the future.

[0249] In this way, this system provides a means to automatically record daily life with minimal burden on the user and later offer it as a valuable memory.

[0250] The following describes the processing flow.

[0251] Step 1:

[0252] The device collects data using sensors. The device acquires images taken by the user through its camera, records audio with its microphone, and obtains location information using its GPS sensor.

[0253] Step 2:

[0254] The device sends the collected data to the cloud server. At regular intervals, or when sufficient data has accumulated, the device uploads encrypted image data, audio data, and location information to the cloud server via the network.

[0255] Step 3:

[0256] The server analyzes the image data. A generating AI on the cloud server analyzes the image data and identifies the objects and people depicted. For example, it performs facial recognition to identify acquaintances if they are in the image.

[0257] Step 4:

[0258] The server analyzes the audio data. The generative AI converts the audio data into text and extracts important keywords and events from the conversation. This allows it to identify what kind of conversation took place.

[0259] Step 5:

[0260] The server analyzes location information. Using GPS data and comparing it with a map database, it identifies the geographical locations the user has visited. Based on the location information, it recognizes visits to specific facilities or landmarks.

[0261] Step 6:

[0262] The server generates a record of the user's activities. Based on the collected analytical data, it organizes the user's daily activities and generates a chronological activity log. Particularly important events and frequently visited places are described in detail.

[0263] Step 7:

[0264] The server provides the user with a recorded activity log. The user can then open the smartphone application to visually review their daily activity log. This allows the user to easily review the places they visited and the people they met on a specific date.

[0265] (Example 1)

[0266] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0267] In daily life, it is not easy to recall past events without relying on memory, and this is especially difficult in today's information-saturated society. Furthermore, there is a need for technology to efficiently extract and analyze useful information from collected data. Additionally, there is a lack of mechanisms to effectively utilize user feedback to improve analytical accuracy.

[0268] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0269] In this invention, the server includes means for analyzing received image data to identify an object or person, means for analyzing received audio data to identify a specific event or activity, and means for analyzing received location information to identify a place. This makes it possible to record the user's daily life in detail, efficiently review past events, and improve analysis accuracy by utilizing feedback.

[0270] A "terminal" is an electronic device used to collect data from a user's daily life, and includes a camera, an audio acquisition device, and a location detection device.

[0271] A "shooting device" refers to a device used to acquire still images or videos, and is equipped with camera functionality.

[0272] A "sound acquisition device" is a device that records ambient sounds and acquires them as data.

[0273] A "location detection device" is a device that uses GPS or other location information technologies to determine the geographical location of a terminal.

[0274] A "cloud server" is a computing device that receives, analyzes, and stores data via a network, and shares information with other devices.

[0275] A "generative AI model" is an algorithm that uses machine learning techniques to analyze data and extract certain patterns or information.

[0276] "Feedback" refers to additional information or corrections that users provide to the system, which are used to improve analysis accuracy and enhance the system itself.

[0277] "Activity logs" are detailed records of a user's daily activities in chronological order, generated by analyzing user data.

[0278] A "user interface" is a component that provides a visual or manipulative means for a user to access and manipulate records and data.

[0279] A "prompt message" is a sentence that functions as specific instructions or guidelines from the user to support the processing of a generative AI model.

[0280] This invention is a system that allows users to efficiently record their daily lives and easily review them. It mainly consists of a terminal and a cloud server.

[0281] Terminal configuration and operation

[0282] The terminal collects data such as photos, voices, and location information in daily life. Specifically, it takes photos using a camera, records voices using a microphone, and obtains location information using the GPS function. These devices are built into the terminal and automatically acquire data according to the user's movements. For example, when the user is having a picnic in the park on a holiday, the terminal takes pictures of the surrounding scenery, records conversations, and obtains the location at that time.

[0283] Data Transmission and Analysis

[0284] The collected data is sent from the terminal to the cloud server via the network. The server receives this and analyzes the data using a generated AI model. It identifies objects and people from image data, extracts keywords by converting the conversation content from voice data into text, and uses a map service to identify specific facilities and locations from location data. This analysis result is generated as the user's activity record.

[0285] Activity Record and Feedback

[0286] The activity record generated by the server is organized in chronological order, and the user can view this through a dedicated user interface. This record details specific activities such as "having a picnic in the park on [month] [day]". Furthermore, by the user correcting the record through feedback, the generated AI model continuously learns and the analysis accuracy improves. For example, by adding details such as "talking with friends in the area with play equipment", more accurate activity records will be provided in subsequent times.

[0287] Specific Examples and Prompt Sentences

[0288] As a specific example, prompt sentences such as "Analyze the photos taken on the weekend and the recorded voices, and create a past activity record." are provided. This prompt sentence serves to give instructions to the server as a guideline for efficiently performing the analysis.

[0289] In this way, the system provides users with an effective and convenient means of recording their daily lives.

[0290] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0291] Step 1:

[0292] The device uses its camera to capture images of the user's surroundings and its microphone to record ambient sounds. It also uses GPS to obtain the user's current location. This step yields three types of input: image data, audio data, and location data. These data are temporarily stored in local storage.

[0293] Step 2:

[0294] The device transmits collected image data, audio data, and location data to a cloud server via the internet. During this process, the data is encrypted and protected. The output is the data transferred to the cloud server.

[0295] Step 3:

[0296] The server analyzes the received image data using a generative AI model. Based on the image's pixel information, it applies an object detection algorithm to identify specific objects or people. The output is the recognized object and its attribute information.

[0297] Step 4:

[0298] The server converts the received audio data into text using speech recognition technology. Furthermore, it analyzes the text content using natural language processing technology, extracting key keywords and phrases. The output consists of the analyzed text information and the extracted keywords.

[0299] Step 5:

[0300] The server matches the received location data with a map API to identify specific facilities or locations. It utilizes a geographic information system to identify relevant facility names and landmarks and provides them as output.

[0301] Step 6:

[0302] The server integrates the previous analysis results and organizes the user's action records in chronological order. These action records detail important events and visited locations. The output is an action record for reviewing the user's daily life.

[0303] Step 7:

[0304] The user checks the action record through the application and provides feedback if necessary. By providing corrections or additional information, the server updates the generated AI model, improving the analysis accuracy. The output is the next accuracy improvement by the updated AI model.

[0305] In this way, specific operations such as collection, transmission, analysis, and integration are performed in each step, providing an accurate and detailed action record for the user.

[0306] (Application Example 1)

[0307] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0308] The conventional shopping experience had the problem that it was difficult to provide individualized customer services online. In particular, personalized information provision based on customers' purchase histories and behaviors was insufficient, making it difficult to optimize the shopping experience. Furthermore, it has been required to enhance customers' immersion in virtual stores compared to physical stores.

[0309] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0310] In this invention, the server includes means for analyzing image information acquired by the terminal and identifying items related to the purchase, means for analyzing voice information and identifying important keywords, and means for analyzing the customer's behavioral history using location information. This makes it possible to provide a personalized purchasing experience based on the customer's behavioral history.

[0311] A "terminal" is an electronic device that collects and records data when carried or worn by a user.

[0312] A "photography device" is a device that includes a camera or other equipment for acquiring image information.

[0313] A "speech acquisition device" is a device that includes a microphone or other equipment for recording speech information.

[0314] A "location detection device" is a device that includes GPS or similar technologies for measuring the user's current location.

[0315] A "remote data processing device" is a device that performs data analysis and management on the cloud or on a server.

[0316] "Image information" refers to visual data collected by a camera or other imaging device.

[0317] "Audio information" refers to audio data acquired by an audio acquisition device.

[0318] "Location information" refers to geographical coordinate data acquired by a location detection device.

[0319] "Action history" refers to data that records a user's past actions and events and organizes them chronologically.

[0320] "Evaluation information" refers to opinions and feedback provided by users, and is used to improve the accuracy of analysis.

[0321] To implement this invention, the system mainly consists of a terminal and a remote data processing device. The terminal is equipped with a camera, an audio acquisition device, and a location detection device, and collects data in the user's daily life. Specifically, the terminal's camera acquires image information of the user's surroundings, and the audio acquisition device records audio information such as conversations and ambient sounds. The location detection device tracks the user's movement path and acquires location information.

[0322] The data collected by the device is transmitted via network communication to a remote data processing device (cloud server). The server analyzes the image and audio information using the Google Cloud Vision API and Google Cloud Speech-to-Text API. From the image information, objects and features are identified, and specific keywords are extracted from the audio information. Location information is analyzed as geographic coordinate data and recorded on the server as part of the user's activity history.

[0323] The server generates user behavior records using an AI model based on the analyzed data. This allows users to receive personalized information based on their past purchasing behavior and places they have visited. For example, users may receive information about new products related to items they have previously purchased at stores they have visited.

[0324] As an example of a prompt message, we use the format: "Customer ID: 12345 tried on a specific category item at store: brand name. Please suggest relevant new products and promotional information." This allows the system to suggest the most relevant information to the user.

[0325] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0326] Step 1:

[0327] The device uses a camera to acquire image information of the user's surroundings. The input is the user's current visual environment, and the output is the acquired image data. This data is stored on a smart device worn by the user.

[0328] Step 2:

[0329] The terminal uses a voice acquisition device to record sounds around the user. The input is the surrounding sound environment, and the output is audio data. This audio data includes the user's conversation and ambient sounds.

[0330] Step 3:

[0331] Using a location detection device, the terminal obtains the user's location information. The input is the current geographical coordinates, and the output is location data. This location data is used to track the user's movement path.

[0332] Step 4:

[0333] The device transmits collected image data, audio data, and location information to a cloud server. The input is this data, and the output is the status indicating successful transmission to the server. The data is transmitted via network communication.

[0334] Step 5:

[0335] The server analyzes the received image information using the Google Cloud Vision API. The input is image data, and the output is identified object and feature data. This analysis identifies specific objects and attributes.

[0336] Step 6:

[0337] The server uses the Google Cloud Speech-to-Text API to analyze audio information and extract specific keywords. The input is audio data, and the output is the extracted text and keywords. The audio content is converted into text data, and important phrases are highlighted.

[0338] Step 7:

[0339] Based on location information, the server analyzes the user's movement history to identify specific locations. The input is location data, and the output is the result of analyzing the user's movement path. This analysis creates movement patterns based on the locations and routes the user has visited.

[0340] Step 8:

[0341] The server uses a generative AI model to generate personalized user behavior records based on the analyzed data. Inputs include analyzed images, audio, and location information, while output is personalized information tailored to the user. This generated information is based on the user's past purchasing behavior.

[0342] Step 9:

[0343] The user receives generated information using prompts and engages in a personalized purchasing experience. The input is prompts from the server, and the output is the user's purchasing behavior. The user makes a purchase decision based on the information presented.

[0344] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0345] This invention relates to a system that automatically records and analyzes a user's daily life to understand their emotional state, thereby providing the user with a richer reflective experience. This system includes an emotion engine that analyzes data collected by the terminal on a cloud server and reflects the user's emotions in the generated behavioral records.

[0346] The device functions as an integral part of the user's life, collecting image and audio data using its camera and microphone, and obtaining location information using its GPS function. This data collection is performed in the background by the device, so the user does not need to take any special action.

[0347] The collected data is sent to a cloud server via the network. On the cloud server, a generative AI analyzes this data, identifying objects and people, and extracting conversation content from audio data. In addition, an emotion engine analyzes image and audio data to recognize the user's emotional state. For example, based on an image, a smile might be interpreted as "joy," and emotions such as tension or anger might be detected from the tone of voice.

[0348] Based on these analysis results, the server generates a daily activity log of the user and integrates emotional information obtained from the emotion engine. This allows users to understand not only where and what happened, but also how they were feeling at the time. For example, a record might be generated stating, "I had a pleasant conversation at dinner with friends and felt happy."

[0349] Users can view these behavioral records through the application and reflect on past events along with their emotions. The system can improve the accuracy of its emotion engine through user feedback, resulting in more accurate and empathetic records.

[0350] In this way, this system enriches users' memories emotionally and provides a means to understand individual moments at a deeper level.

[0351] The following describes the processing flow.

[0352] Step 1:

[0353] The device collects data. It uses a camera to acquire image data and a microphone to record audio data. It also collects location information through its GPS function. This data is automatically acquired during daily life, without requiring any specific user action.

[0354] Step 2:

[0355] The device sends data to the cloud server. The collected image data, audio data, and location information are transferred to the cloud server in an encrypted state at regular intervals. This ensures that the data is securely managed and awaits analysis.

[0356] Step 3:

[0357] The server analyzes the image data. The generating AI identifies objects and people in the image and further infers the user's emotional state through facial expression analysis. For example, if the user is smiling in the photo, it will estimate "joy."

[0358] Step 4:

[0359] The server analyzes the audio data. Using speech recognition technology, it converts the recorded audio into text and extracts the content of the conversation. At the same time, it detects emotions such as tension, excitement, and anger through speech tone analysis.

[0360] Step 5:

[0361] The server analyzes the location information. It compares GPS data with map information to identify the places the user has visited. This allows for a detailed record of the areas and facilities the user has been in.

[0362] Step 6:

[0363] The server generates activity logs. Based on the analysis results, it compiles the user's daily records into an activity history. By integrating the output of the emotion engine, it generates a rich record that includes the user's emotional state.

[0364] Step 7:

[0365] The server displays the generated activity log in the user interface. Users can freely access this information through the app and reflect on their past events and emotions. Through the feedback function, users can provide information to further improve the accuracy of the record.

[0366] (Example 2)

[0367] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0368] Conventional recording systems often simply record events in a user's daily life based only on time and place, making it difficult to generate records that reflect the emotional state at the time. As a result, important emotional information is often missing when users reflect deeply on past events, leading to a challenge in obtaining a rich reflective experience.

[0369] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0370] In this invention, the server includes an analysis means for recognizing emotional states based on visual and audio data, a means for generating user behavior records and providing them in an integrated manner, and a means for performing learning processing to improve analysis accuracy based on evaluation information provided by the user. This enables users to reflect on past events accompanied by emotions and obtain a richer recording experience.

[0371] A "terminal" is a device that acquires visual data, audio data, and location information in a user's daily life.

[0372] An "image acquisition device" is a device used to collect visual data, and this mainly refers to a camera.

[0373] A "speech acquisition device" is a device used to collect speech data, and this mainly refers to a microphone.

[0374] A "location detection device" is a device used to measure and record a user's location information, and is primarily equipped with GPS functionality.

[0375] An "information processing device" is a computer system that analyzes received data and generates a record of user behavior.

[0376] "Visual data" refers to image information acquired by an image acquisition device.

[0377] "Audio data" refers to sound information acquired by an audio acquisition device.

[0378] "Emotional state" refers to the user's emotional response, analyzed based on visual and audio data.

[0379] "Activity logs" are records of a user's past activities and events, and may also include emotional information.

[0380] "Evaluation information" refers to user-provided feedback, which is used to improve the accuracy of the analysis.

[0381] "Learning process" is a process performed to improve the accuracy of the analysis algorithm based on evaluation information.

[0382] This invention is a system that records a user's daily life from both visual and auditory perspectives, analyzes these records, and adds emotional information. Specifically, a terminal collects visual data, audio data, and location information, and transmits them to an information processing device. The information processing device is equipped with means to perform data analysis using a generative AI model and identify the user's emotional state.

[0383] The device functions as a smartphone or wearable device and is equipped with image acquisition and audio acquisition devices. Specifically, it utilizes a camera for image acquisition and a microphone for audio acquisition. Location information is acquired using GPS functionality. This data is collected in the background by the device, and no special action is required from the user.

[0384] The collected data is securely encrypted and transmitted to the information processing device. The information processing device analyzes the received data using "OpenCV" as image analysis software and "Google Cloud Speech-to-Text" as speech analysis software. This allows it to identify objects and faces from visual data and extract conversations and activity content from audio data. Based on the information obtained through this data processing, the emotion engine determines the user's emotional state. For example, if an image contains a smile, it identifies the emotional state as "joy," and if the tone of voice is elevated, it identifies the emotional state as "excitement."

[0385] The server generates user behavior records based on the analysis results and provides them to the user in an integrated form with emotional information. These behavior records are accessible to the user through the application. A learning process is also performed to improve analysis accuracy based on evaluation information, and user feedback is incorporated to provide more accurate and empathetic records.

[0386] As a concrete example, the prompt message is as follows: "Analyze the emotional state during yesterday's event and generate a record of the behavior." This allows the user to obtain a record that includes detailed emotional information about past events.

[0387] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0388] Step 1:

[0389] The device collects visual data, audio data, and location information. Inputs include real-time visual and audio information from the user's surroundings, and location information obtained via GPS. The device captures images with its camera, records audio with its microphone, and simultaneously acquires its current location. The output is a dataset combining all of this data.

[0390] Step 2:

[0391] The terminal transmits the collected data to the information processing device. The input is a dataset of visual data, audio data, and location information generated in step 1. The terminal securely encrypts the data and transmits it to the information processing device over the network. The output is the encrypted dataset received by the information processing device.

[0392] Step 3:

[0393] The server acts as an information processing device, analyzing the received data. The input consists of datasets of visual data, audio data, and location information transmitted from the terminal. Using a generative AI model, the server first analyzes the visual data with image analysis software to identify objects and people. The audio data is analyzed using an audio analysis program to extract conversations and activity content. The output is the analysis result, including object information, person information, and audio content.

[0394] Step 4:

[0395] The server uses the analysis results to recognize the emotional state. The input consists of object information, person information, and audio content obtained in step 3. The server's emotion engine estimates the user's emotional state based on facial expression information from the visual data and tone information from the audio. For example, if there are many smiles, the emotion of "joy" will be recognized. The output is detailed analysis data with the estimated emotional state added.

[0396] Step 5:

[0397] The server generates user behavior records and integrates emotional information. The input is detailed analytical data including emotional states. Based on this analytical data, the server compiles behavior records chronologically and adds emotional information. The output is the user's behavior record, including emotional information. This record is converted into a format viewable by the user's application.

[0398] Step 6:

[0399] The user views their behavioral records through the application and provides feedback as needed. The input is the behavioral records generated in step 5. The user can access this to reflect on past events and emotions. The user can also provide feedback to improve the accuracy of emotion recognition. The output is the user's feedback information. This feedback is used in the system's learning process.

[0400] (Application Example 2)

[0401] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0402] In modern retail, it is difficult to understand customer purchasing behavior and emotions in real time, which presents a challenge in providing appropriate sales approaches. Furthermore, traditional methods make it difficult to efficiently detect customer interests and dissatisfactions and adjust sales strategies immediately, thus failing to maximize customer satisfaction.

[0403] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0404] In this invention, the server includes means for recording image data acquired by a terminal using a camera, means for recording audio data acquired by a terminal using an audio acquisition device, means for acquiring and recording location information using a location detection device, means for estimating the emotional state of a customer from their image and audio data and providing it in real time, and means for identifying the customer's interests in the store and adjusting the sales strategy accordingly. This enables real-time understanding of the customer's emotional state and the development of flexible sales strategies based on this understanding.

[0405] A "terminal" is a device that acquires image data, audio data, and location information of users and customers.

[0406] A "photography device" is a device used to acquire image data and functions as a camera.

[0407] A "voice acquisition device" is a device for recording voice data and functions as a microphone.

[0408] A "location detection device" is a device that acquires and records a user's location information and functions as a GPS.

[0409] A "cloud server" is a server that receives data sent from a terminal and performs analysis on it.

[0410] "Image data" refers to visual information acquired by a camera or imaging device.

[0411] "Audio data" refers to auditory information acquired by an audio acquisition device.

[0412] "Location information" refers to information indicating a geographical location, obtained by a location detection device.

[0413] "Emotional state" refers to the user's psychological state, which is estimated by analyzing image and audio data.

[0414] "Activity log" refers to a record of a user's daily activities, generated based on analyzed data.

[0415] "Sales strategy" refers to the plans and methods used to effectively sell products in physical stores.

[0416] "Analytical accuracy" refers to the degree to which the results obtained through data analysis are correct.

[0417] This invention provides a system that analyzes customer purchasing behavior and emotional states in physical stores in real time and deploys appropriate sales strategies in real time. This system consists of terminals, a cloud server, and various analysis software.

[0418] The device is a pair of smart glasses worn by the user (customer) and is equipped with a camera and microphone that function as both a camera and an audio acquisition device. This device also has GPS functionality as a location detection device, allowing it to pinpoint the customer's location within the store. This enables the continuous collection of image data, audio data, and location information, which are then transmitted to a cloud server without requiring any user interaction.

[0419] The cloud server is equipped with image recognition and speech analysis models using TensorFlow, which identify objects and people from received image data. It also analyzes speech data to extract conversation content and uses a generative AI model to estimate the customer's emotional state. The estimated emotional state is provided to field service staff in real time in the form of joy, interest, anxiety, etc., and is used to assist customers in the store.

[0420] Based on these data analysis results, the server generates behavioral records to identify customer interests and dissatisfactions. For example, if it is estimated that a customer is standing in front of a shelf for a long time, smiling and looking at a product, this information is immediately notified to a store employee, who can then approach the customer individually. A prompt message might be, "How to build a system that analyzes a customer's facial expression and tone of voice when they pick up a new product using smart glasses, clarifies their emotional state using an emotion engine, and provides the optimal sales approach." This prompt clarifies the analysis procedure within the system, leading to more appropriate customer service.

[0421] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0422] Step 1:

[0423] The device acquires image and audio data of the customer's surroundings through smart glasses. This records the customer's visual and auditory environment within the store. Image data includes store merchandise and the customer's facial expressions, while audio data includes the customer's tone of voice and conversation content. The input is real-world visual and auditory data, and the output is digital image and audio data.

[0424] Step 2:

[0425] The terminal transmits the acquired image and audio data to the cloud server via Wi-Fi. During this process, location information generated by the location detection device is also transmitted. The input from the terminal consists of images, audio, and location information, while the output is the arrival of these data packets to the server.

[0426] Step 3:

[0427] The server analyzes the received image data using TensorFlow to identify objects or people. The analysis uses a deep learning model to extract features from the image and match them with known people or objects. The input is image data sent from the terminal, and the output is a list of identified objects or people.

[0428] Step 4:

[0429] The server analyzes the voice data and uses a generative AI model to estimate the customer's emotional state. The voice analysis engine analyzes the tone, pitch, and volume of the voice to identify emotions. The input is voice data, and the output is recorded as the detected emotional state.

[0430] Step 5:

[0431] The server identifies the customer's location within the store based on the acquired location information. It then analyzes the customer's behavior in specific areas based on the location information to identify their interests. The input is location information, and the output is the customer's precise location within the store.

[0432] Step 6:

[0433] The user (store clerk) receives analysis results from the server in real time and uses that information to interact with customers. The system outputs information about the customer's emotional state and interests to their smart device, leading to appropriate sales responses. The input is the analysis results from the server, and the output is the specific actions taken by the store clerk based on that information.

[0434] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0435] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0436] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0437] [Third Embodiment]

[0438] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0439] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0440] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0441] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0442] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0443] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0444] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0445] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0446] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0447] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0448] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0449] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0450] This invention relates to a system that automatically records a user's daily life and allows them to easily review it. Specifically, a smart device (hereinafter referred to as "terminal") collects data using various sensors, transmits it to a cloud server for analysis, and records various aspects of the user's life.

[0451] The device continuously collects images taken by the user, audio recordings, and location information. For example, if a user visits a historical site that is a tourist destination, the device records photos of the scene and uses GPS location information to pinpoint the exact location visited. Also, when a user is talking to a friend, the device captures that conversation as audio data.

[0452] The data collected by the device is transmitted to a cloud server via network communication. The cloud server analyzes the received data using advanced generative AI technology, recognizing objects and people from image data, and transcribing conversations from audio data to extract key keywords. Location information is cross-referenced with map data to identify specific facilities and landmarks.

[0453] The server generates information to create a user activity log. This log clearly presents the user's daily activities in chronological order, with important events described in detail. Users can view this log through the application, for example, seeing a record such as "On [Month] [Day], had lunch with an acquaintance at a tourist spot." In this way, users can easily review past events without relying on their memory.

[0454] Furthermore, the accuracy of the analysis is improved through user feedback. Based on corrections and additional information from users, the generated AI model is continuously updated to ensure that more accurate records are reflected in the future.

[0455] In this way, this system provides a means to automatically record daily life with minimal burden on the user and later offer it as a valuable memory.

[0456] The following describes the processing flow.

[0457] Step 1:

[0458] The device collects data using sensors. The device acquires images taken by the user through its camera, records audio with its microphone, and obtains location information using its GPS sensor.

[0459] Step 2:

[0460] The device sends the collected data to the cloud server. At regular intervals, or when sufficient data has accumulated, the device uploads encrypted image data, audio data, and location information to the cloud server via the network.

[0461] Step 3:

[0462] The server analyzes the image data. A generating AI on the cloud server analyzes the image data and identifies the objects and people depicted. For example, it performs facial recognition to identify acquaintances if they are in the image.

[0463] Step 4:

[0464] The server analyzes the audio data. The generative AI converts the audio data into text and extracts important keywords and events from the conversation. This allows it to identify what kind of conversation took place.

[0465] Step 5:

[0466] The server analyzes location information. Using GPS data and comparing it with a map database, it identifies the geographical locations the user has visited. Based on the location information, it recognizes visits to specific facilities or landmarks.

[0467] Step 6:

[0468] The server generates a record of the user's activities. Based on the collected analytical data, it organizes the user's daily activities and generates a chronological activity log. Particularly important events and frequently visited places are described in detail.

[0469] Step 7:

[0470] The server provides the user with a recorded activity log. The user can then open the smartphone application to visually review their daily activity log. This allows the user to easily review the places they visited and the people they met on a specific date.

[0471] (Example 1)

[0472] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0473] In daily life, it is not easy to recall past events without relying on memory, and this is especially difficult in today's information-saturated society. Furthermore, there is a need for technology to efficiently extract and analyze useful information from collected data. Additionally, there is a lack of mechanisms to effectively utilize user feedback to improve analytical accuracy.

[0474] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0475] In this invention, the server includes means for analyzing received image data to identify an object or person, means for analyzing received audio data to identify a specific event or activity, and means for analyzing received location information to identify a place. This makes it possible to record the user's daily life in detail, efficiently review past events, and improve analysis accuracy by utilizing feedback.

[0476] A "terminal" is an electronic device used to collect data from a user's daily life, and includes a camera, an audio acquisition device, and a location detection device.

[0477] A "shooting device" refers to a device used to acquire still images or videos, and is equipped with camera functionality.

[0478] A "sound acquisition device" is a device that records ambient sounds and acquires them as data.

[0479] A "location detection device" is a device that uses GPS or other location information technologies to determine the geographical location of a terminal.

[0480] A "cloud server" is a computing device that receives, analyzes, and stores data via a network, and shares information with other devices.

[0481] A "generative AI model" is an algorithm that uses machine learning techniques to analyze data and extract certain patterns or information.

[0482] "Feedback" refers to additional information or corrections that users provide to the system, which are used to improve analysis accuracy and enhance the system itself.

[0483] "Activity logs" are detailed records of a user's daily activities in chronological order, generated by analyzing user data.

[0484] A "user interface" is a component that provides a visual or manipulative means for a user to access and manipulate records and data.

[0485] A "prompt message" is a sentence that functions as specific instructions or guidelines from the user to support the processing of a generative AI model.

[0486] This invention is a system that allows users to efficiently record their daily lives and easily review them. It mainly consists of a terminal and a cloud server.

[0487] Terminal configuration and operation

[0488] The device collects data such as photos, audio, and location information from everyday life. Specifically, it takes photos using the camera, records audio using the microphone, and obtains location information using the GPS function. These devices are built into the device and automatically acquire data according to the user's movements. For example, if a user is having a picnic in a park on a holiday, the device will take photos of the surrounding scenery, record conversations, and obtain the user's location at that time.

[0489] Data transmission and analysis

[0490] The collected data is transmitted from the device to a cloud server via the network. The server receives this data and analyzes it using a generative AI model. Objects and people are identified from image data, and conversation content is converted to text and keywords extracted from audio data. Location data is used with a map service to identify specific facilities and locations. This analysis result is then generated as a user activity record.

[0491] Activity log and feedback

[0492] The activity logs generated by the server are organized chronologically, and users can view them through a dedicated user interface. These logs include detailed information about specific activities, such as "had a picnic in the park on [date]." Furthermore, the generating AI model continuously learns and improves analysis accuracy as users correct the logs through feedback. For example, adding details such as "had a conversation with a friend in the playground area" will provide more accurate activity logs in the future.

[0493] Specific examples and prompt statements

[0494] As a concrete example, a prompt message such as, "Analyze the photos taken and audio recorded over the weekend and create a record of past activities," is provided. This prompt message serves as a guideline to instruct the server in order to perform the analysis efficiently.

[0495] In this way, the system provides users with an effective and convenient means of recording their daily lives.

[0496] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0497] Step 1:

[0498] The device uses its camera to capture images of the user's surroundings and its microphone to record ambient sounds. It also uses GPS to obtain the user's current location. This step yields three types of input: image data, audio data, and location data. These data are temporarily stored in local storage.

[0499] Step 2:

[0500] The device transmits collected image data, audio data, and location data to a cloud server via the internet. During this process, the data is encrypted and protected. The output is the data transferred to the cloud server.

[0501] Step 3:

[0502] The server analyzes the received image data using a generative AI model. Based on the image's pixel information, it applies an object detection algorithm to identify specific objects or people. The output is the recognized object and its attribute information.

[0503] Step 4:

[0504] The server converts the received audio data into text using speech recognition technology. Furthermore, it analyzes the text content using natural language processing technology, extracting key keywords and phrases. The output consists of the analyzed text information and the extracted keywords.

[0505] Step 5:

[0506] The server compares the received location data with a map API to identify specific facilities or locations. It then utilizes a geographic information system to identify related facility names and landmarks, providing them as output.

[0507] Step 6:

[0508] The server integrates the analysis results from previous sessions and organizes the user's activity log chronologically. This activity log details important events and places visited. The output is an activity log for users to reflect on their daily lives.

[0509] Step 7:

[0510] Users review their activity logs through the application and provide feedback as needed. The server updates the generated AI model based on user corrections and additional information, improving analysis accuracy. The output represents the next accuracy improvement based on the updated AI model.

[0511] In this way, specific actions such as collection, transmission, analysis, and integration are performed at each step, providing users with accurate and detailed records of their activities.

[0512] (Application Example 1)

[0513] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0514] Traditional shopping experiences have faced challenges in providing personalized customer service online. In particular, the provision of personalized information based on customer purchase history and behavior has been insufficient, making it difficult to optimize the shopping experience. Furthermore, there is a need to enhance customer immersion in virtual stores compared to physical stores.

[0515] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0516] In this invention, the server includes means for analyzing image information acquired by the terminal and identifying items related to the purchase, means for analyzing voice information and identifying important keywords, and means for analyzing the customer's behavioral history using location information. This makes it possible to provide a personalized purchasing experience based on the customer's behavioral history.

[0517] A "terminal" is an electronic device that collects and records data when carried or worn by a user.

[0518] A "photography device" is a device that includes a camera or other equipment for acquiring image information.

[0519] A "speech acquisition device" is a device that includes a microphone or other equipment for recording speech information.

[0520] A "location detection device" is a device that includes GPS or similar technologies for measuring the user's current location.

[0521] A "remote data processing device" is a device that performs data analysis and management on the cloud or on a server.

[0522] "Image information" refers to visual data collected by a camera or other imaging device.

[0523] "Audio information" refers to audio data acquired by an audio acquisition device.

[0524] "Location information" refers to geographical coordinate data acquired by a location detection device.

[0525] "Action history" refers to data that records a user's past actions and events and organizes them chronologically.

[0526] "Evaluation information" refers to opinions and feedback provided by users, and is used to improve the accuracy of analysis.

[0527] To implement this invention, the system mainly consists of a terminal and a remote data processing device. The terminal is equipped with a camera, an audio acquisition device, and a location detection device, and collects data in the user's daily life. Specifically, the terminal's camera acquires image information of the user's surroundings, and the audio acquisition device records audio information such as conversations and ambient sounds. The location detection device tracks the user's movement path and acquires location information.

[0528] The data collected by the device is transmitted via network communication to a remote data processing device (cloud server). The server analyzes the image and audio information using the Google Cloud Vision API and Google Cloud Speech-to-Text API. From the image information, objects and features are identified, and specific keywords are extracted from the audio information. Location information is analyzed as geographic coordinate data and recorded on the server as part of the user's activity history.

[0529] The server generates user behavior records using an AI model based on the analyzed data. This allows users to receive personalized information based on their past purchasing behavior and places they have visited. For example, users may receive information about new products related to items they have previously purchased at stores they have visited.

[0530] As an example of a prompt message, we use the format: "Customer ID: 12345 tried on a specific category item at store: brand name. Please suggest relevant new products and promotional information." This allows the system to suggest the most relevant information to the user.

[0531] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0532] Step 1:

[0533] The device uses a camera to acquire image information of the user's surroundings. The input is the user's current visual environment, and the output is the acquired image data. This data is stored on a smart device worn by the user.

[0534] Step 2:

[0535] The terminal uses a voice acquisition device to record sounds around the user. The input is the surrounding sound environment, and the output is audio data. This audio data includes the user's conversation and ambient sounds.

[0536] Step 3:

[0537] Using a location detection device, the terminal obtains the user's location information. The input is the current geographical coordinates, and the output is location data. This location data is used to track the user's movement path.

[0538] Step 4:

[0539] The device transmits collected image data, audio data, and location information to a cloud server. The input is this data, and the output is the status indicating successful transmission to the server. The data is transmitted via network communication.

[0540] Step 5:

[0541] The server analyzes the received image information using the Google Cloud Vision API. The input is image data, and the output is identified object and feature data. This analysis identifies specific objects and attributes.

[0542] Step 6:

[0543] The server uses the Google Cloud Speech-to-Text API to analyze audio information and extract specific keywords. The input is audio data, and the output is the extracted text and keywords. The audio content is converted into text data, and important phrases are highlighted.

[0544] Step 7:

[0545] Based on location information, the server analyzes the user's movement history to identify specific locations. The input is location data, and the output is the result of analyzing the user's movement path. This analysis creates movement patterns based on the locations and routes the user has visited.

[0546] Step 8:

[0547] The server uses a generative AI model to generate personalized user behavior records based on the analyzed data. Inputs include analyzed images, audio, and location information, while output is personalized information tailored to the user. This generated information is based on the user's past purchasing behavior.

[0548] Step 9:

[0549] The user receives generated information using prompts and engages in a personalized purchasing experience. The input is prompts from the server, and the output is the user's purchasing behavior. The user makes a purchase decision based on the information presented.

[0550] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0551] This invention relates to a system that automatically records and analyzes a user's daily life to understand their emotional state, thereby providing the user with a richer reflective experience. This system includes an emotion engine that analyzes data collected by the terminal on a cloud server and reflects the user's emotions in the generated behavioral records.

[0552] The device functions as an integral part of the user's life, collecting image and audio data using its camera and microphone, and obtaining location information using its GPS function. This data collection is performed in the background by the device, so the user does not need to take any special action.

[0553] The collected data is sent to a cloud server via the network. On the cloud server, a generative AI analyzes this data, identifying objects and people, and extracting conversation content from audio data. In addition, an emotion engine analyzes image and audio data to recognize the user's emotional state. For example, based on an image, a smile might be interpreted as "joy," and emotions such as tension or anger might be detected from the tone of voice.

[0554] Based on these analysis results, the server generates a daily activity log of the user and integrates emotional information obtained from the emotion engine. This allows users to understand not only where and what happened, but also how they were feeling at the time. For example, a record might be generated stating, "I had a pleasant conversation at dinner with friends and felt happy."

[0555] Users can view these behavioral records through the application and reflect on past events along with their emotions. The system can improve the accuracy of its emotion engine through user feedback, resulting in more accurate and empathetic records.

[0556] In this way, this system enriches users' memories emotionally and provides a means to understand individual moments at a deeper level.

[0557] The following describes the processing flow.

[0558] Step 1:

[0559] The device collects data. It uses a camera to acquire image data and a microphone to record audio data. It also collects location information through its GPS function. This data is automatically acquired during daily life, without requiring any specific user action.

[0560] Step 2:

[0561] The device sends data to the cloud server. The collected image data, audio data, and location information are transferred to the cloud server in an encrypted state at regular intervals. This ensures that the data is securely managed and awaits analysis.

[0562] Step 3:

[0563] The server analyzes the image data. The generating AI identifies objects and people in the image and further infers the user's emotional state through facial expression analysis. For example, if the user is smiling in the photo, it will estimate "joy."

[0564] Step 4:

[0565] The server analyzes the audio data. Using speech recognition technology, it converts the recorded audio into text and extracts the content of the conversation. At the same time, it detects emotions such as tension, excitement, and anger through speech tone analysis.

[0566] Step 5:

[0567] The server analyzes the location information. It compares GPS data with map information to identify the places the user has visited. This allows for a detailed record of the areas and facilities the user has been in.

[0568] Step 6:

[0569] The server generates activity logs. Based on the analysis results, it compiles the user's daily records into an activity history. By integrating the output of the emotion engine, it generates a rich record that includes the user's emotional state.

[0570] Step 7:

[0571] The server displays the generated activity log in the user interface. Users can freely access this information through the app and reflect on their past events and emotions. Through the feedback function, users can provide information to further improve the accuracy of the record.

[0572] (Example 2)

[0573] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0574] Conventional recording systems often simply record events in a user's daily life based only on time and place, making it difficult to generate records that reflect the emotional state at the time. As a result, important emotional information is often missing when users reflect deeply on past events, leading to a challenge in obtaining a rich reflective experience.

[0575] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0576] In this invention, the server includes an analysis means for recognizing emotional states based on visual and audio data, a means for generating user behavior records and providing them in an integrated manner, and a means for performing learning processing to improve analysis accuracy based on evaluation information provided by the user. This enables users to reflect on past events accompanied by emotions and obtain a richer recording experience.

[0577] A "terminal" is a device that acquires visual data, audio data, and location information in a user's daily life.

[0578] An "image acquisition device" is a device used to collect visual data, and this mainly refers to a camera.

[0579] A "speech acquisition device" is a device used to collect speech data, and this mainly refers to a microphone.

[0580] A "location detection device" is a device used to measure and record a user's location information, and is primarily equipped with GPS functionality.

[0581] An "information processing device" is a computer system that analyzes received data and generates a record of user behavior.

[0582] "Visual data" refers to image information acquired by an image acquisition device.

[0583] "Audio data" refers to sound information acquired by an audio acquisition device.

[0584] "Emotional state" refers to the user's emotional response, analyzed based on visual and audio data.

[0585] "Activity logs" are records of a user's past activities and events, and may also include emotional information.

[0586] "Evaluation information" refers to user-provided feedback, which is used to improve the accuracy of the analysis.

[0587] "Learning process" is a process performed to improve the accuracy of the analysis algorithm based on evaluation information.

[0588] This invention is a system that records a user's daily life from both visual and auditory perspectives, analyzes these records, and adds emotional information. Specifically, a terminal collects visual data, audio data, and location information, and transmits them to an information processing device. The information processing device is equipped with means to perform data analysis using a generative AI model and identify the user's emotional state.

[0589] The device functions as a smartphone or wearable device and is equipped with image acquisition and audio acquisition devices. Specifically, it utilizes a camera for image acquisition and a microphone for audio acquisition. Location information is acquired using GPS functionality. This data is collected in the background by the device, and no special action is required from the user.

[0590] The collected data is securely encrypted and transmitted to the information processing device. The information processing device analyzes the received data using "OpenCV" as image analysis software and "Google Cloud Speech-to-Text" as speech analysis software. This allows it to identify objects and faces from visual data and extract conversations and activity content from audio data. Based on the information obtained through this data processing, the emotion engine determines the user's emotional state. For example, if an image contains a smile, it identifies the emotional state as "joy," and if the tone of voice is elevated, it identifies the emotional state as "excitement."

[0591] The server generates user behavior records based on the analysis results and provides them to the user in an integrated form with emotional information. These behavior records are accessible to the user through the application. A learning process is also performed to improve analysis accuracy based on evaluation information, and user feedback is incorporated to provide more accurate and empathetic records.

[0592] As a concrete example, the prompt message is as follows: "Analyze the emotional state during yesterday's event and generate a record of the behavior." This allows the user to obtain a record that includes detailed emotional information about past events.

[0593] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0594] Step 1:

[0595] The device collects visual data, audio data, and location information. Inputs include real-time visual and audio information from the user's surroundings, and location information obtained via GPS. The device captures images with its camera, records audio with its microphone, and simultaneously acquires its current location. The output is a dataset combining all of this data.

[0596] Step 2:

[0597] The terminal transmits the collected data to the information processing device. The input is a dataset of visual data, audio data, and location information generated in step 1. The terminal securely encrypts the data and transmits it to the information processing device over the network. The output is the encrypted dataset received by the information processing device.

[0598] Step 3:

[0599] The server acts as an information processing device, analyzing the received data. The input consists of datasets of visual data, audio data, and location information transmitted from the terminal. Using a generative AI model, the server first analyzes the visual data with image analysis software to identify objects and people. The audio data is analyzed using an audio analysis program to extract conversations and activity content. The output is the analysis result, including object information, person information, and audio content.

[0600] Step 4:

[0601] The server uses the analysis results to recognize the emotional state. The input consists of object information, person information, and audio content obtained in step 3. The server's emotion engine estimates the user's emotional state based on facial expression information from the visual data and tone information from the audio. For example, if there are many smiles, the emotion of "joy" will be recognized. The output is detailed analysis data with the estimated emotional state added.

[0602] Step 5:

[0603] The server generates user behavior records and integrates emotional information. The input is detailed analytical data including emotional states. Based on this analytical data, the server compiles behavior records chronologically and adds emotional information. The output is the user's behavior record, including emotional information. This record is converted into a format viewable by the user's application.

[0604] Step 6:

[0605] The user views their behavioral records through the application and provides feedback as needed. The input is the behavioral records generated in step 5. The user can access this to reflect on past events and emotions. The user can also provide feedback to improve the accuracy of emotion recognition. The output is the user's feedback information. This feedback is used in the system's learning process.

[0606] (Application Example 2)

[0607] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0608] In modern retail, it is difficult to understand customer purchasing behavior and emotions in real time, which presents a challenge in providing appropriate sales approaches. Furthermore, traditional methods make it difficult to efficiently detect customer interests and dissatisfactions and adjust sales strategies immediately, thus failing to maximize customer satisfaction.

[0609] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0610] In this invention, the server includes means for recording image data acquired by a terminal using a camera, means for recording audio data acquired by a terminal using an audio acquisition device, means for acquiring and recording location information using a location detection device, means for estimating the emotional state of a customer from their image and audio data and providing it in real time, and means for identifying the customer's interests in the store and adjusting the sales strategy accordingly. This enables real-time understanding of the customer's emotional state and the development of flexible sales strategies based on this understanding.

[0611] A "terminal" is a device that acquires image data, audio data, and location information of users and customers.

[0612] A "photography device" is a device used to acquire image data and functions as a camera.

[0613] A "voice acquisition device" is a device for recording voice data and functions as a microphone.

[0614] A "location detection device" is a device that acquires and records a user's location information and functions as a GPS.

[0615] A "cloud server" is a server that receives data sent from a terminal and performs analysis on it.

[0616] "Image data" refers to visual information acquired by a camera or imaging device.

[0617] "Audio data" refers to auditory information acquired by an audio acquisition device.

[0618] "Location information" refers to information indicating a geographical location, obtained by a location detection device.

[0619] "Emotional state" refers to the user's psychological state, which is estimated by analyzing image and audio data.

[0620] "Activity log" refers to a record of a user's daily activities, generated based on analyzed data.

[0621] "Sales strategy" refers to the plans and methods used to effectively sell products in physical stores.

[0622] "Analytical accuracy" refers to the degree to which the results obtained through data analysis are correct.

[0623] This invention provides a system that analyzes customer purchasing behavior and emotional states in physical stores in real time and deploys appropriate sales strategies in real time. This system consists of terminals, a cloud server, and various analysis software.

[0624] The device is a pair of smart glasses worn by the user (customer) and is equipped with a camera and microphone that function as both a camera and an audio acquisition device. This device also has GPS functionality as a location detection device, allowing it to pinpoint the customer's location within the store. This enables the continuous collection of image data, audio data, and location information, which are then transmitted to a cloud server without requiring any user interaction.

[0625] The cloud server is equipped with image recognition and speech analysis models using TensorFlow, which identify objects and people from received image data. It also analyzes speech data to extract conversation content and uses a generative AI model to estimate the customer's emotional state. The estimated emotional state is provided to field service staff in real time in the form of joy, interest, anxiety, etc., and is used to assist customers in the store.

[0626] Based on these data analysis results, the server generates behavioral records to identify customer interests and dissatisfactions. For example, if it is estimated that a customer is standing in front of a shelf for a long time, smiling and looking at a product, this information is immediately notified to a store employee, who can then approach the customer individually. A prompt message might be, "How to build a system that analyzes a customer's facial expression and tone of voice when they pick up a new product using smart glasses, clarifies their emotional state using an emotion engine, and provides the optimal sales approach." This prompt clarifies the analysis procedure within the system, leading to more appropriate customer service.

[0627] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0628] Step 1:

[0629] The device acquires image and audio data of the customer's surroundings through smart glasses. This records the customer's visual and auditory environment within the store. Image data includes store merchandise and the customer's facial expressions, while audio data includes the customer's tone of voice and conversation content. The input is real-world visual and auditory data, and the output is digital image and audio data.

[0630] Step 2:

[0631] The terminal transmits the acquired image and audio data to the cloud server via Wi-Fi. During this process, location information generated by the location detection device is also transmitted. The input from the terminal consists of images, audio, and location information, while the output is the arrival of these data packets to the server.

[0632] Step 3:

[0633] The server analyzes the received image data using TensorFlow to identify objects or people. The analysis uses a deep learning model to extract features from the image and match them with known people or objects. The input is image data sent from the terminal, and the output is a list of identified objects or people.

[0634] Step 4:

[0635] The server analyzes the voice data and uses a generative AI model to estimate the customer's emotional state. The voice analysis engine analyzes the tone, pitch, and volume of the voice to identify emotions. The input is voice data, and the output is recorded as the detected emotional state.

[0636] Step 5:

[0637] The server identifies the customer's location within the store based on the acquired location information. It then analyzes the customer's behavior in specific areas based on the location information to identify their interests. The input is location information, and the output is the customer's precise location within the store.

[0638] Step 6:

[0639] The user (store clerk) receives analysis results from the server in real time and uses that information to interact with customers. The system outputs information about the customer's emotional state and interests to their smart device, leading to appropriate sales responses. The input is the analysis results from the server, and the output is the specific actions taken by the store clerk based on that information.

[0640] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0641] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0642] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0643] [Fourth Embodiment]

[0644] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0645] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0646] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0647] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0648] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0649] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0650] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0651] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0652] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0653] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0654] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0655] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0656] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0657] This invention relates to a system that automatically records a user's daily life and allows them to easily review it. Specifically, a smart device (hereinafter referred to as "terminal") collects data using various sensors, transmits it to a cloud server for analysis, and records various aspects of the user's life.

[0658] The device continuously collects images taken by the user, audio recordings, and location information. For example, if a user visits a historical site that is a tourist destination, the device records photos of the scene and uses GPS location information to pinpoint the exact location visited. Also, when a user is talking to a friend, the device captures that conversation as audio data.

[0659] The data collected by the device is transmitted to a cloud server via network communication. The cloud server analyzes the received data using advanced generative AI technology, recognizing objects and people from image data, and transcribing conversations from audio data to extract key keywords. Location information is cross-referenced with map data to identify specific facilities and landmarks.

[0660] The server generates information to create a user activity log. This log clearly presents the user's daily activities in chronological order, with important events described in detail. Users can view this log through the application, for example, seeing a record such as "On [Month] [Day], had lunch with an acquaintance at a tourist spot." In this way, users can easily review past events without relying on their memory.

[0661] Furthermore, the accuracy of the analysis is improved through user feedback. Based on corrections and additional information from users, the generated AI model is continuously updated to ensure that more accurate records are reflected in the future.

[0662] In this way, this system provides a means to automatically record daily life with minimal burden on the user and later offer it as a valuable memory.

[0663] The following describes the processing flow.

[0664] Step 1:

[0665] The device collects data using sensors. The device acquires images taken by the user through its camera, records audio with its microphone, and obtains location information using its GPS sensor.

[0666] Step 2:

[0667] The device sends the collected data to the cloud server. At regular intervals, or when sufficient data has accumulated, the device uploads encrypted image data, audio data, and location information to the cloud server via the network.

[0668] Step 3:

[0669] The server analyzes the image data. A generating AI on the cloud server analyzes the image data and identifies the objects and people depicted. For example, it performs facial recognition to identify acquaintances if they are in the image.

[0670] Step 4:

[0671] The server analyzes the audio data. The generative AI converts the audio data into text and extracts important keywords and events from the conversation. This allows it to identify what kind of conversation took place.

[0672] Step 5:

[0673] The server analyzes location information. Using GPS data and comparing it with a map database, it identifies the geographical locations the user has visited. Based on the location information, it recognizes visits to specific facilities or landmarks.

[0674] Step 6:

[0675] The server generates a record of the user's activities. Based on the collected analytical data, it organizes the user's daily activities and generates a chronological activity log. Particularly important events and frequently visited places are described in detail.

[0676] Step 7:

[0677] The server provides the user with a recorded activity log. The user can then open the smartphone application to visually review their daily activity log. This allows the user to easily review the places they visited and the people they met on a specific date.

[0678] (Example 1)

[0679] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0680] In daily life, it is not easy to recall past events without relying on memory, and this is especially difficult in today's information-saturated society. Furthermore, there is a need for technology to efficiently extract and analyze useful information from collected data. Additionally, there is a lack of mechanisms to effectively utilize user feedback to improve analytical accuracy.

[0681] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0682] In this invention, the server includes means for analyzing received image data to identify an object or person, means for analyzing received audio data to identify a specific event or activity, and means for analyzing received location information to identify a place. This makes it possible to record the user's daily life in detail, efficiently review past events, and improve analysis accuracy by utilizing feedback.

[0683] A "terminal" is an electronic device used to collect data from a user's daily life, and includes a camera, an audio acquisition device, and a location detection device.

[0684] A "shooting device" refers to a device used to acquire still images or videos, and is equipped with camera functionality.

[0685] A "sound acquisition device" is a device that records ambient sounds and acquires them as data.

[0686] A "location detection device" is a device that uses GPS or other location information technologies to determine the geographical location of a terminal.

[0687] A "cloud server" is a computing device that receives, analyzes, and stores data via a network, and shares information with other devices.

[0688] A "generative AI model" is an algorithm that uses machine learning techniques to analyze data and extract certain patterns or information.

[0689] "Feedback" refers to additional information or corrections that users provide to the system, which are used to improve analysis accuracy and enhance the system itself.

[0690] "Activity logs" are detailed records of a user's daily activities in chronological order, generated by analyzing user data.

[0691] A "user interface" is a component that provides a visual or manipulative means for a user to access and manipulate records and data.

[0692] A "prompt message" is a sentence that functions as specific instructions or guidelines from the user to support the processing of a generative AI model.

[0693] This invention is a system that allows users to efficiently record their daily lives and easily review them. It mainly consists of a terminal and a cloud server.

[0694] Terminal configuration and operation

[0695] The device collects data such as photos, audio, and location information from everyday life. Specifically, it takes photos using the camera, records audio using the microphone, and obtains location information using the GPS function. These devices are built into the device and automatically acquire data according to the user's movements. For example, if a user is having a picnic in a park on a holiday, the device will take photos of the surrounding scenery, record conversations, and obtain the user's location at that time.

[0696] Data transmission and analysis

[0697] The collected data is transmitted from the device to a cloud server via the network. The server receives this data and analyzes it using a generative AI model. Objects and people are identified from image data, and conversation content is converted to text and keywords extracted from audio data. Location data is used with a map service to identify specific facilities and locations. This analysis result is then generated as a user activity record.

[0698] Activity log and feedback

[0699] The activity logs generated by the server are organized chronologically, and users can view them through a dedicated user interface. These logs include detailed information about specific activities, such as "had a picnic in the park on [date]." Furthermore, the generating AI model continuously learns and improves analysis accuracy as users correct the logs through feedback. For example, adding details such as "had a conversation with a friend in the playground area" will provide more accurate activity logs in the future.

[0700] Specific examples and prompt statements

[0701] As a concrete example, a prompt message such as, "Analyze the photos taken and audio recorded over the weekend and create a record of past activities," is provided. This prompt message serves as a guideline to instruct the server in order to perform the analysis efficiently.

[0702] In this way, the system provides users with an effective and convenient means of recording their daily lives.

[0703] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0704] Step 1:

[0705] The device uses its camera to capture images of the user's surroundings and its microphone to record ambient sounds. It also uses GPS to obtain the user's current location. This step yields three types of input: image data, audio data, and location data. These data are temporarily stored in local storage.

[0706] Step 2:

[0707] The device transmits collected image data, audio data, and location data to a cloud server via the internet. During this process, the data is encrypted and protected. The output is the data transferred to the cloud server.

[0708] Step 3:

[0709] The server analyzes the received image data using a generative AI model. Based on the image's pixel information, it applies an object detection algorithm to identify specific objects or people. The output is the recognized object and its attribute information.

[0710] Step 4:

[0711] The server converts the received audio data into text using speech recognition technology. Furthermore, it analyzes the text content using natural language processing technology, extracting key keywords and phrases. The output consists of the analyzed text information and the extracted keywords.

[0712] Step 5:

[0713] The server compares the received location data with a map API to identify specific facilities or locations. It then utilizes a geographic information system to identify related facility names and landmarks, providing them as output.

[0714] Step 6:

[0715] The server integrates the analysis results from previous sessions and organizes the user's activity log chronologically. This activity log details important events and places visited. The output is an activity log for users to reflect on their daily lives.

[0716] Step 7:

[0717] Users review their activity logs through the application and provide feedback as needed. The server updates the generated AI model based on user corrections and additional information, improving analysis accuracy. The output represents the next accuracy improvement based on the updated AI model.

[0718] In this way, specific actions such as collection, transmission, analysis, and integration are performed at each step, providing users with accurate and detailed records of their activities.

[0719] (Application Example 1)

[0720] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0721] Traditional shopping experiences have faced challenges in providing personalized customer service online. In particular, the provision of personalized information based on customer purchase history and behavior has been insufficient, making it difficult to optimize the shopping experience. Furthermore, there is a need to enhance customer immersion in virtual stores compared to physical stores.

[0722] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0723] In this invention, the server includes means for analyzing image information acquired by the terminal and identifying items related to the purchase, means for analyzing voice information and identifying important keywords, and means for analyzing the customer's behavioral history using location information. This makes it possible to provide a personalized purchasing experience based on the customer's behavioral history.

[0724] A "terminal" is an electronic device that collects and records data when carried or worn by a user.

[0725] A "photography device" is a device that includes a camera or other equipment for acquiring image information.

[0726] A "speech acquisition device" is a device that includes a microphone or other equipment for recording speech information.

[0727] A "location detection device" is a device that includes GPS or similar technologies for measuring the user's current location.

[0728] A "remote data processing device" is a device that performs data analysis and management on the cloud or on a server.

[0729] "Image information" refers to visual data collected by a camera or other imaging device.

[0730] "Audio information" refers to audio data acquired by an audio acquisition device.

[0731] "Location information" refers to geographical coordinate data acquired by a location detection device.

[0732] "Action history" refers to data that records a user's past actions and events and organizes them chronologically.

[0733] "Evaluation information" refers to opinions and feedback provided by users, and is used to improve the accuracy of analysis.

[0734] To implement this invention, the system mainly consists of a terminal and a remote data processing device. The terminal is equipped with a camera, an audio acquisition device, and a location detection device, and collects data in the user's daily life. Specifically, the terminal's camera acquires image information of the user's surroundings, and the audio acquisition device records audio information such as conversations and ambient sounds. The location detection device tracks the user's movement path and acquires location information.

[0735] The data collected by the device is transmitted via network communication to a remote data processing device (cloud server). The server analyzes the image and audio information using the Google Cloud Vision API and Google Cloud Speech-to-Text API. From the image information, objects and features are identified, and specific keywords are extracted from the audio information. Location information is analyzed as geographic coordinate data and recorded on the server as part of the user's activity history.

[0736] The server generates user behavior records using an AI model based on the analyzed data. This allows users to receive personalized information based on their past purchasing behavior and places they have visited. For example, users may receive information about new products related to items they have previously purchased at stores they have visited.

[0737] As an example of a prompt message, we use the format: "Customer ID: 12345 tried on a specific category item at store: brand name. Please suggest relevant new products and promotional information." This allows the system to suggest the most relevant information to the user.

[0738] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0739] Step 1:

[0740] The device uses a camera to acquire image information of the user's surroundings. The input is the user's current visual environment, and the output is the acquired image data. This data is stored on a smart device worn by the user.

[0741] Step 2:

[0742] The terminal uses a voice acquisition device to record sounds around the user. The input is the surrounding sound environment, and the output is audio data. This audio data includes the user's conversation and ambient sounds.

[0743] Step 3:

[0744] Using a location detection device, the terminal obtains the user's location information. The input is the current geographical coordinates, and the output is location data. This location data is used to track the user's movement path.

[0745] Step 4:

[0746] The device transmits collected image data, audio data, and location information to a cloud server. The input is this data, and the output is the status indicating successful transmission to the server. The data is transmitted via network communication.

[0747] Step 5:

[0748] The server analyzes the received image information using the Google Cloud Vision API. The input is image data, and the output is identified object and feature data. This analysis identifies specific objects and attributes.

[0749] Step 6:

[0750] The server uses the Google Cloud Speech-to-Text API to analyze audio information and extract specific keywords. The input is audio data, and the output is the extracted text and keywords. The audio content is converted into text data, and important phrases are highlighted.

[0751] Step 7:

[0752] Based on location information, the server analyzes the user's movement history to identify specific locations. The input is location data, and the output is the result of analyzing the user's movement path. This analysis creates movement patterns based on the locations and routes the user has visited.

[0753] Step 8:

[0754] The server uses a generative AI model to generate personalized user behavior records based on the analyzed data. Inputs include analyzed images, audio, and location information, while output is personalized information tailored to the user. This generated information is based on the user's past purchasing behavior.

[0755] Step 9:

[0756] The user receives generated information using prompts and engages in a personalized purchasing experience. The input is prompts from the server, and the output is the user's purchasing behavior. The user makes a purchase decision based on the information presented.

[0757] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0758] This invention relates to a system that automatically records and analyzes a user's daily life to understand their emotional state, thereby providing the user with a richer reflective experience. This system includes an emotion engine that analyzes data collected by the terminal on a cloud server and reflects the user's emotions in the generated behavioral records.

[0759] The device functions as an integral part of the user's life, collecting image and audio data using its camera and microphone, and obtaining location information using its GPS function. This data collection is performed in the background by the device, so the user does not need to take any special action.

[0760] The collected data is sent to a cloud server via the network. On the cloud server, a generative AI analyzes this data, identifying objects and people, and extracting conversation content from audio data. In addition, an emotion engine analyzes image and audio data to recognize the user's emotional state. For example, based on an image, a smile might be interpreted as "joy," and emotions such as tension or anger might be detected from the tone of voice.

[0761] Based on these analysis results, the server generates a daily activity log of the user and integrates emotional information obtained from the emotion engine. This allows users to understand not only where and what happened, but also how they were feeling at the time. For example, a record might be generated stating, "I had a pleasant conversation at dinner with friends and felt happy."

[0762] Users can view these behavioral records through the application and reflect on past events along with their emotions. The system can improve the accuracy of its emotion engine through user feedback, resulting in more accurate and empathetic records.

[0763] In this way, this system enriches users' memories emotionally and provides a means to understand individual moments at a deeper level.

[0764] The following describes the processing flow.

[0765] Step 1:

[0766] The device collects data. It uses a camera to acquire image data and a microphone to record audio data. It also collects location information through its GPS function. This data is automatically acquired during daily life, without requiring any specific user action.

[0767] Step 2:

[0768] The device sends data to the cloud server. The collected image data, audio data, and location information are transferred to the cloud server in an encrypted state at regular intervals. This ensures that the data is securely managed and awaits analysis.

[0769] Step 3:

[0770] The server analyzes the image data. The generating AI identifies objects and people in the image and further infers the user's emotional state through facial expression analysis. For example, if the user is smiling in the photo, it will estimate "joy."

[0771] Step 4:

[0772] The server analyzes the audio data. Using speech recognition technology, it converts the recorded audio into text and extracts the content of the conversation. At the same time, it detects emotions such as tension, excitement, and anger through speech tone analysis.

[0773] Step 5:

[0774] The server analyzes the location information. It compares GPS data with map information to identify the places the user has visited. This allows for a detailed record of the areas and facilities the user has been in.

[0775] Step 6:

[0776] The server generates activity logs. Based on the analysis results, it compiles the user's daily records into an activity history. By integrating the output of the emotion engine, it generates a rich record that includes the user's emotional state.

[0777] Step 7:

[0778] The server displays the generated activity log in the user interface. Users can freely access this information through the app and reflect on their past events and emotions. Through the feedback function, users can provide information to further improve the accuracy of the record.

[0779] (Example 2)

[0780] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0781] Conventional recording systems often simply record events in a user's daily life based only on time and place, making it difficult to generate records that reflect the emotional state at the time. As a result, important emotional information is often missing when users reflect deeply on past events, leading to a challenge in obtaining a rich reflective experience.

[0782] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0783] In this invention, the server includes an analysis means for recognizing emotional states based on visual and audio data, a means for generating user behavior records and providing them in an integrated manner, and a means for performing learning processing to improve analysis accuracy based on evaluation information provided by the user. This enables users to reflect on past events accompanied by emotions and obtain a richer recording experience.

[0784] A "terminal" is a device that acquires visual data, audio data, and location information in a user's daily life.

[0785] An "image acquisition device" is a device used to collect visual data, and this mainly refers to a camera.

[0786] A "speech acquisition device" is a device used to collect speech data, and this mainly refers to a microphone.

[0787] A "location detection device" is a device used to measure and record a user's location information, and is primarily equipped with GPS functionality.

[0788] An "information processing device" is a computer system that analyzes received data and generates a record of user behavior.

[0789] "Visual data" refers to image information acquired by an image acquisition device.

[0790] "Audio data" refers to sound information acquired by an audio acquisition device.

[0791] "Emotional state" refers to the user's emotional response, analyzed based on visual and audio data.

[0792] "Activity logs" are records of a user's past activities and events, and may also include emotional information.

[0793] "Evaluation information" refers to user-provided feedback, which is used to improve the accuracy of the analysis.

[0794] "Learning process" is a process performed to improve the accuracy of the analysis algorithm based on evaluation information.

[0795] This invention is a system that records a user's daily life from both visual and auditory perspectives, analyzes these records, and adds emotional information. Specifically, a terminal collects visual data, audio data, and location information, and transmits them to an information processing device. The information processing device is equipped with means to perform data analysis using a generative AI model and identify the user's emotional state.

[0796] The device functions as a smartphone or wearable device and is equipped with image acquisition and audio acquisition devices. Specifically, it utilizes a camera for image acquisition and a microphone for audio acquisition. Location information is acquired using GPS functionality. This data is collected in the background by the device, and no special action is required from the user.

[0797] The collected data is securely encrypted and transmitted to the information processing device. The information processing device analyzes the received data using "OpenCV" as image analysis software and "Google Cloud Speech-to-Text" as speech analysis software. This allows it to identify objects and faces from visual data and extract conversations and activity content from audio data. Based on the information obtained through this data processing, the emotion engine determines the user's emotional state. For example, if an image contains a smile, it identifies the emotional state as "joy," and if the tone of voice is elevated, it identifies the emotional state as "excitement."

[0798] The server generates user behavior records based on the analysis results and provides them to the user in an integrated form with emotional information. These behavior records are accessible to the user through the application. A learning process is also performed to improve analysis accuracy based on evaluation information, and user feedback is incorporated to provide more accurate and empathetic records.

[0799] As a concrete example, the prompt message is as follows: "Analyze the emotional state during yesterday's event and generate a record of the behavior." This allows the user to obtain a record that includes detailed emotional information about past events.

[0800] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0801] Step 1:

[0802] The device collects visual data, audio data, and location information. Inputs include real-time visual and audio information from the user's surroundings, and location information obtained via GPS. The device captures images with its camera, records audio with its microphone, and simultaneously acquires its current location. The output is a dataset combining all of this data.

[0803] Step 2:

[0804] The terminal transmits the collected data to the information processing device. The input is a dataset of visual data, audio data, and location information generated in step 1. The terminal securely encrypts the data and transmits it to the information processing device over the network. The output is the encrypted dataset received by the information processing device.

[0805] Step 3:

[0806] The server acts as an information processing device, analyzing the received data. The input consists of datasets of visual data, audio data, and location information transmitted from the terminal. Using a generative AI model, the server first analyzes the visual data with image analysis software to identify objects and people. The audio data is analyzed using an audio analysis program to extract conversations and activity content. The output is the analysis result, including object information, person information, and audio content.

[0807] Step 4:

[0808] The server uses the analysis results to recognize the emotional state. The input consists of object information, person information, and audio content obtained in step 3. The server's emotion engine estimates the user's emotional state based on facial expression information from the visual data and tone information from the audio. For example, if there are many smiles, the emotion of "joy" will be recognized. The output is detailed analysis data with the estimated emotional state added.

[0809] Step 5:

[0810] The server generates user behavior records and integrates emotional information. The input is detailed analytical data including emotional states. Based on this analytical data, the server compiles behavior records chronologically and adds emotional information. The output is the user's behavior record, including emotional information. This record is converted into a format viewable by the user's application.

[0811] Step 6:

[0812] The user views their behavioral records through the application and provides feedback as needed. The input is the behavioral records generated in step 5. The user can access this to reflect on past events and emotions. The user can also provide feedback to improve the accuracy of emotion recognition. The output is the user's feedback information. This feedback is used in the system's learning process.

[0813] (Application Example 2)

[0814] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0815] In modern retail, it is difficult to understand customer purchasing behavior and emotions in real time, which presents a challenge in providing appropriate sales approaches. Furthermore, traditional methods make it difficult to efficiently detect customer interests and dissatisfactions and adjust sales strategies immediately, thus failing to maximize customer satisfaction.

[0816] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0817] In this invention, the server includes means for recording image data acquired by a terminal using a camera, means for recording audio data acquired by a terminal using an audio acquisition device, means for acquiring and recording location information using a location detection device, means for estimating the emotional state of a customer from their image and audio data and providing it in real time, and means for identifying the customer's interests in the store and adjusting the sales strategy accordingly. This enables real-time understanding of the customer's emotional state and the development of flexible sales strategies based on this understanding.

[0818] A "terminal" is a device that acquires image data, audio data, and location information of users and customers.

[0819] A "photography device" is a device used to acquire image data and functions as a camera.

[0820] A "voice acquisition device" is a device for recording voice data and functions as a microphone.

[0821] A "location detection device" is a device that acquires and records a user's location information and functions as a GPS.

[0822] A "cloud server" is a server that receives data sent from a terminal and performs analysis on it.

[0823] "Image data" refers to visual information acquired by a camera or imaging device.

[0824] "Audio data" refers to auditory information acquired by an audio acquisition device.

[0825] "Location information" refers to information indicating a geographical location, obtained by a location detection device.

[0826] "Emotional state" refers to the user's psychological state, which is estimated by analyzing image and audio data.

[0827] "Activity log" refers to a record of a user's daily activities, generated based on analyzed data.

[0828] "Sales strategy" refers to the plans and methods used to effectively sell products in physical stores.

[0829] "Analytical accuracy" refers to the degree to which the results obtained through data analysis are correct.

[0830] This invention provides a system that analyzes customer purchasing behavior and emotional states in physical stores in real time and deploys appropriate sales strategies in real time. This system consists of terminals, a cloud server, and various analysis software.

[0831] The device is a pair of smart glasses worn by the user (customer) and is equipped with a camera and microphone that function as both a camera and an audio acquisition device. This device also has GPS functionality as a location detection device, allowing it to pinpoint the customer's location within the store. This enables the continuous collection of image data, audio data, and location information, which are then transmitted to a cloud server without requiring any user interaction.

[0832] The cloud server is equipped with image recognition and speech analysis models using TensorFlow, which identify objects and people from received image data. It also analyzes speech data to extract conversation content and uses a generative AI model to estimate the customer's emotional state. The estimated emotional state is provided to field service staff in real time in the form of joy, interest, anxiety, etc., and is used to assist customers in the store.

[0833] Based on these data analysis results, the server generates behavioral records to identify customer interests and dissatisfactions. For example, if it is estimated that a customer is standing in front of a shelf for a long time, smiling and looking at a product, this information is immediately notified to a store employee, who can then approach the customer individually. A prompt message might be, "How to build a system that analyzes a customer's facial expression and tone of voice when they pick up a new product using smart glasses, clarifies their emotional state using an emotion engine, and provides the optimal sales approach." This prompt clarifies the analysis procedure within the system, leading to more appropriate customer service.

[0834] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0835] Step 1:

[0836] The device acquires image and audio data of the customer's surroundings through smart glasses. This records the customer's visual and auditory environment within the store. Image data includes store merchandise and the customer's facial expressions, while audio data includes the customer's tone of voice and conversation content. The input is real-world visual and auditory data, and the output is digital image and audio data.

[0837] Step 2:

[0838] The terminal transmits the acquired image and audio data to the cloud server via Wi-Fi. During this process, location information generated by the location detection device is also transmitted. The input from the terminal consists of images, audio, and location information, while the output is the arrival of these data packets to the server.

[0839] Step 3:

[0840] The server analyzes the received image data using TensorFlow to identify objects or people. The analysis uses a deep learning model to extract features from the image and match them with known people or objects. The input is image data sent from the terminal, and the output is a list of identified objects or people.

[0841] Step 4:

[0842] The server analyzes the voice data and uses a generative AI model to estimate the customer's emotional state. The voice analysis engine analyzes the tone, pitch, and volume of the voice to identify emotions. The input is voice data, and the output is recorded as the detected emotional state.

[0843] Step 5:

[0844] The server identifies the customer's location within the store based on the acquired location information. It then analyzes the customer's behavior in specific areas based on the location information to identify their interests. The input is location information, and the output is the customer's precise location within the store.

[0845] Step 6:

[0846] The user (store clerk) receives analysis results from the server in real time and uses that information to interact with customers. The system outputs information about the customer's emotional state and interests to their smart device, leading to appropriate sales responses. The input is the analysis results from the server, and the output is the specific actions taken by the store clerk based on that information.

[0847] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0848] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0849] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0850] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0851] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0852] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0853] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0854] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0855] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0856] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0857] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0858] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0859] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0860] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0861] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0862] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0863] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0864] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0865] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0866] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0867] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0868] The following is further disclosed regarding the embodiments described above.

[0869] (Claim 1)

[0870] A means for the terminal to record image data acquired by the camera,

[0871] A means for recording audio data acquired by a voice acquisition device,

[0872] A means for acquiring and recording location information using a location detection device,

[0873] Means for transmitting the aforementioned image data, audio data, and location information to a cloud server,

[0874] A means for analyzing image data received by a cloud server to identify an object or person,

[0875] A means of analyzing audio data received by a cloud server to identify a specific event or activity,

[0876] A means of analyzing location information received by a cloud server to identify a location,

[0877] A means for generating and providing user behavior records based on the aforementioned analysis results,

[0878] A system that includes this.

[0879] (Claim 2)

[0880] The system according to claim 1, further comprising means for performing a learning process to improve analysis accuracy based on user-provided feedback information.

[0881] (Claim 3)

[0882] The system according to claim 1, further comprising means for organizing the generated activity records chronologically and making them viewable through a user interface.

[0883] "Example 1"

[0884] (Claim 1)

[0885] A means for the terminal to record image data acquired by the camera,

[0886] A means for recording audio data acquired by a voice acquisition device,

[0887] A means for acquiring and recording location information using a location detection device,

[0888] Means for transmitting the aforementioned image data, audio data, and location information,

[0889] A means for analyzing image data received by a cloud server to identify a target or person,

[0890] A means of analyzing audio data received by a cloud server to identify a specific event or activity,

[0891] A means of analyzing location information received by a cloud server to determine the location,

[0892] A means for generating and providing user behavior records based on the aforementioned analysis results,

[0893] A means of performing learning processing to improve analysis accuracy based on user feedback information,

[0894] A means of organizing the generated activity records chronologically and making them viewable through the user interface,

[0895] A system that includes this.

[0896] (Claim 2)

[0897] The system according to claim 1, further comprising means for repeatedly performing the learning process to improve the accuracy of the analysis and updating the generated AI model.

[0898] (Claim 3)

[0899] The system according to claim 1, further comprising means for generating prompt sentences that assist in generating accurate behavioral records based on feedback provided by the user.

[0900] "Application Example 1"

[0901] (Claim 1)

[0902] A means for recording image information acquired by a camera,

[0903] A means for recording audio information acquired by a voice acquisition device,

[0904] A means for acquiring and recording location information using a location detection device,

[0905] Means for transmitting the aforementioned image information, audio information, and location information to a remote data processing device,

[0906] A means for analyzing image information received by a remote data processing device and identifying an object or feature,

[0907] A means for analyzing audio information received by a remote data processing device to identify a specific event or activity,

[0908] A means for analyzing location information received by a remote data processing device to identify a location,

[0909] A means for generating and providing the user's activity history based on the aforementioned analysis results,

[0910] Furthermore, the means for analyzing the user's past purchasing behavior and generating individually optimized recommendation information,

[0911] A system that includes this.

[0912] (Claim 2)

[0913] The system according to claim 1, further comprising means for performing a learning process to improve the accuracy of analysis based on evaluation information provided by the user.

[0914] (Claim 3)

[0915] The system according to claim 1, further comprising means for organizing the generated operation history in chronological order and making it viewable through a user interface.

[0916] "Example 2 of combining an emotion engine"

[0917] (Claim 1)

[0918] A means for recording visual data acquired by an image acquisition device on a terminal,

[0919] A means for recording audio data acquired by a voice acquisition device,

[0920] A means for acquiring and recording location information using a location detection device,

[0921] Means for transmitting the aforementioned visual data, audio data, and location information to an information processing device,

[0922] An information processing device analyzes received visual data and has means for identifying an object or person,

[0923] A means for analyzing audio data received by an information processing device to identify a specific activity or situation,

[0924] A means for analyzing location information received by an information processing device to identify a location,

[0925] An analysis means for recognizing an emotional state based on the aforementioned visual and audio data,

[0926] A means for generating a user behavior record based on the analysis results and providing the integrated emotional state,

[0927] A system that includes this.

[0928] (Claim 2)

[0929] The system according to claim 1, further comprising means for performing a learning process to improve the accuracy of the analysis based on evaluation information provided by the user.

[0930] (Claim 3)

[0931] The system according to claim 1, further comprising means for organizing the generated activity records in chronological order and making them viewable through a user-facing operation screen.

[0932] "Application example 2 when combining with an emotional engine"

[0933] (Claim 1)

[0934] A means for the terminal to record image data acquired by the camera,

[0935] A means for recording audio data acquired by a voice acquisition device,

[0936] A means for acquiring and recording location information using a location detection device,

[0937] Means for transmitting the aforementioned image data, audio data, and location information to a cloud server,

[0938] A means for analyzing image data received by a cloud server to identify an object or person,

[0939] A means of analyzing audio data received by a cloud server to identify a specific event or activity,

[0940] A means of analyzing location information received by a cloud server to identify a location,

[0941] A means of estimating and providing in real time the emotional state of a customer from their image and voice data,

[0942] A means of identifying customer interests within a store and adjusting sales strategies accordingly,

[0943] A means for generating and providing user behavior records based on the aforementioned analysis results,

[0944] A system that includes this.

[0945] (Claim 2)

[0946] The system according to claim 1, further comprising means for performing a learning process to improve analysis accuracy based on user-provided feedback information.

[0947] (Claim 3)

[0948] The system according to claim 1, further comprising means for organizing the generated activity records chronologically and making them viewable through a user interface. [Explanation of Symbols]

[0949] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for the terminal to record image data acquired by the camera, A means for recording audio data acquired by a voice acquisition device, A means for acquiring and recording location information using a location detection device, Means for transmitting the aforementioned image data, audio data, and location information to a cloud server, A means for analyzing image data received by a cloud server to identify an object or person, A means of analyzing audio data received by a cloud server to identify a specific event or activity, A means of analyzing location information received by a cloud server to identify a location, A means for generating and providing user behavior records based on the aforementioned analysis results, A system that includes this.

2. The system according to claim 1, further comprising means for performing a learning process to improve analysis accuracy based on user-provided feedback information.

3. The system according to claim 1, further comprising means for organizing the generated activity records chronologically and making them viewable through a user interface.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A