System
The system allows children to store and recall memories by using audio instructions or hand gestures to save photographs in the cloud, addressing the lack of easy preservation and reminiscence in traditional methods.
Patent Information
- Application Number
- JP2024181361
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-25
- Filing Date
- 2024-10-16
- Publication Date
- 2025-05-12
Smart Images

Figure 2025073085000001_ABST
Abstract
Description
[Technical field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including a description and related instruction sentence regarding the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2022-180282 A Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional methods have been limited in the means by which children can easily store the scenes and experiences they see in their daily lives, and have been insufficient in supporting recollection and recollection. [Means for solving the problem]
[0005] The present invention provides a means for taking pictures by voice commands or hand gestures and storing the taken images in the cloud, allowing children to instantly store scenes and experiences they see in their daily lives. In addition, in response to questions about stored images, related images are obtained and displayed through dialogue or text input to support recollection. This allows children's memories and records of their growth to be easily stored and looked back on later.
[0006] The term "photography instruction" refers to the use of voice instructions or hand gestures to take a picture when the user wants to save a particular scene or experience.
[0007] "Voice commands" refer to a user providing commands to a system using spoken words or phrases.
[0008] "Hand gestures" refer to a user giving instructions to a system using hand and finger movements.
[0009] "Cloud" refers to an environment or service that uses resources such as servers and databases on the Internet to store and process data.
[0010] A "database" is a collection of data used to organize and manage information, and provides a mechanism for efficient data search and storage.
[0011] The term "associated question" refers to a user question or query about a stored image, and indicates information or conditions based on which related images are retrieved from the database.
[0012] The above are definitions of important terms contained in the claims. These definitions make it easier to understand the patent documents and clarify the technical scope. [Brief description of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Diagram 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. FIG. [Diagram 3] FIG. 11 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Diagram 5] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 13 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] 4 is a sequence diagram showing a process flow of the data processing system according to the first embodiment. FIG. [Figure 12] 11 is a sequence diagram showing a process flow of the data processing system in application example 1. FIG. [Figure 13] FIG. 11 is a sequence diagram showing the flow of processing of the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 11 is a sequence diagram showing the flow of processing in the data processing system in application example 2 when combined with an emotion engine. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a signed processor (hereinafter simply referred to as a "processor") may be one arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be one type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.
[0017] In the following embodiments, a signed RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a code is an interface including a communication processor and an antenna. The communication I / F controls communication between multiple computers. An example of a communication standard applied to the communication I / F is a wireless communication standard including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. In addition, in this specification, the same idea as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (e.g., a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (e.g., voice and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs voice according to instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a Complementary Metal-Oxide-Semiconductor (CMOS) image sensor or a Charge Coupled Device (CCD) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Fig. 2, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32. The specific process program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific process program 56 from the storage 32, and executes the read specific process program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific process program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores a reception output program 60. The reception output program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads out the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] The embodiment for carrying out the present invention comprises the following elements.
[0035] 1. Terminal: A device with an interface that accepts voice commands and hand gestures when a user wants to save a particular scene or experience.
[0036] 2. Shooting: The device analyzes voice commands and hand gestures to understand the command to shoot, activates the camera, and shoots the specified scene or experience.
[0037] 3. Transmission means: The terminal has a communication function for transmitting captured images to the cloud. Images can be transmitted to a server via a network.
[0038] 4. Server: Stores the received images in a database, accepts queries about the contents stored by the user, and retrieves and displays related captured images.
[0039] (Specific examples)
[0040] When a user uses the smart device 14 to implement the present invention, the smart device 14 becomes a terminal. The user commands the smart device 14 to take a picture by speaking, such as "save it," or by making a specific gesture. The smart device 14 analyzes the voice command or gesture and understands the command to take a picture. The camera starts up and captures the scene or experience as instructed. The captured image is sent to the cloud using the communication function of the smart device 14. The server stores and saves the received image in a database. When the user asks the smart device 14 a question about the saved content, the smart device 14 sends the question to the server. The server analyzes the received question, retrieves the related captured image from the database, and sends it to the smart device 14. The smart device 14 displays the received captured image to provide the user with support for reminiscence.
[0041] The above is an example of an embodiment of the present invention. Although specific devices and processing procedures may vary depending on the selection and implementation of the implementer, the present invention can be implemented based on this embodiment.
[0042] The process flow will be explained below.
[0043] Step 1: The user speaks into the terminal, saying "Save."
[0044] Step 2: The device receives the voice command and uses voice recognition technology to interpret the command.
[0045] Step 3: Based on the analysis results, the device understands the shooting instructions.
[0046] Step 4: The device activates its camera and captures the specified scene or experience.
[0047] Step 5: The device uses its communication functions to send the captured images to the cloud.
[0048] Step 6: The server stores the received images in a database.
[0049] Step 7: The user asks a question to the terminal.
[0050] Step 8: The terminal sends the query to the server.
[0051] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[0052] Step 10: The server transmits the captured image to the terminal.
[0053] Step 11: The terminal displays the captured image that it has received.
[0054] Examples:
[0055] Step 1: The user issues a voice command to the smart device 14 saying "Save."
[0056] Step 2: The smart device 14 receives the voice instruction and parses the instruction using voice recognition technology.
[0057] Step 3: The smart device 14 understands the shooting instructions based on the analysis results.
[0058] Step 4: The smart device 14 launches a camera app and captures the instructed scene or experience.
[0059] Step 5: The smart device 14 uses its communication function to transmit the captured image to the cloud.
[0060] Step 6: The server stores the received images in a database.
[0061] Step 7: The user asks the smart device 14, "Show me pictures of playing in the park yesterday."
[0062] Step 8: The smart device 14 sends the query to the server.
[0063] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[0064] Step 10: The server transmits the acquired photographed image to the smart device 14.
[0065] Step 11: The smart device 14 displays the received captured image.
[0066] The above is the specific process flow of this system. The user gives voice instructions or gestures, and the device analyzes them and performs appropriate processing, enabling the system to take photos, save images, and answer questions.
[0067] Example 1
[0068] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."
[0069] There is a demand for users to be able to easily save specific scenes and experiences and easily search and display the saved information at a later date. However, with conventional technologies, it is difficult to intuitively give shooting instructions using voice commands or hand gestures, or to efficiently retrieve images saved in the cloud based on a question.
[0070] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0071] In this invention, the server includes a means for receiving a shooting instruction by voice instruction or hand gesture, a means for analyzing the received shooting instruction to start the camera and shoot the scene, a means for transmitting the shot image to the cloud server via the network, a means for storing the received image in a database, and a means for receiving a question about the saved image from the user and acquiring and displaying related images, thereby enabling the user to intuitively store a particular scene or experience and efficiently recall it later.
[0072] A "voice instruction" refers to a specific operation or instruction given by a user to a terminal using voice.
[0073] A "hand gesture" is a specific operation or instruction given by a user to a terminal using hand movements.
[0074] A "photography instruction" is an instruction given by a user to a terminal to capture a specific scene or experience with a camera.
[0075] "Analysis" is the process of understanding the voice instructions and hand gestures received by the device and converting them into appropriate actions.
[0076] "Starting the camera" means starting up the camera function of the device and making it ready to take pictures.
[0077] A "cloud server" is a server used to store and process data over the Internet.
[0078] A database is a structured storage system that organizes and stores large amounts of data and allows for efficient search and retrieval.
[0079] "Network communication technology" refers to technology used for data communication, including Wi-Fi and mobile data.
[0080] A "question" is a query made by a user to obtain information about a previously stored image.
[0081] "Retrieval" is the process of searching and retrieving specific information from a database.
[0082] "Display" means to output the acquired information on the terminal screen so that the user can confirm it.
[0083] The embodiment of the present invention provides a system that allows a user to easily save a particular scene or experience, and easily search and display it later. A specific implementation method of this system will be described below.
[0084] A user uses a dedicated terminal to give voice instructions and hand gestures. This terminal is equipped with voice recognition technology and gesture recognition technology, and for example, voice recognition software provided by each company is used. For gesture recognition, a motion sensor combined with the terminal's camera is used.
[0085] The device analyzes the user's voice instructions and gestures and activates the camera based on those. The camera is capable of capturing high-resolution images. The captured images are sent to a cloud server using the device's network communication function (Wi-Fi or mobile data communication). In this case, HTTP / HTTPS is generally used as the communication protocol.
[0086] The cloud server processes the received image data appropriately, adds metadata (such as the date and time of shooting and the location), and stores the data in a database. The database uses storage provided by each company.
[0087] The user can later ask questions about the stored images through the device, for example by issuing a voice command such as "Show me the scenery from last week." The question is converted into text using the device's voice recognition technology and sent to the server.
[0088] The server analyzes the user's question, searches for and retrieves relevant images from the database, and then transmits the retrieved images to the terminal, which displays the received images on its screen and provides them to the user.
[0089] Examples:
[0090] When a user visits a park, they give a voice command to their smart device saying, "Save this view," or perform a specific hand gesture. The smart device uses its built-in voice recognition software to analyze the voice, activates the camera, and captures the view. The captured image is sent to a cloud server using its Wi-Fi function. The server adds metadata to the received image and stores it in a database.
[0091] A few days later, the user asks the smart device, "Show me the scenery from last week." The smart device converts this voice to text and sends it to the server. The server searches the database for the corresponding image and sends it to the smart device. The device displays the image on the screen, allowing the user to reminisce about the scenery from last week.
[0092] Example prompt:
[0093] Design a system that allows a user to take pictures of scenery in a park and store them for later review. Include instructions for using voice commands or hand gestures to take pictures, store the data in the cloud, and play the stored images at a later date.
[0094] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0095] Step 1:
[0096] The user gives a voice command or a gesture.
[0097] Specific action: The user speaks to the smart device, saying "Save this view," or makes a specific hand gesture.
[0098] Input: Voice information or gesture data.
[0099] Output: Recognized by the terminal as audio or video data.
[0100] Step 2:
[0101] The device analyzes the voice command or gesture.
[0102] Specific operation: The voice recognition software in the terminal analyzes the voice, and the gesture recognition software analyzes the video data.
[0103] Input: The output data of step 1 (audio or video data).
[0104] Output: Parsed instruction (e.g. "Save the view").
[0105] Step 3:
[0106] The device will activate the camera and take a picture of the specified scene.
[0107] Specific operation: Based on the results of voice or gesture analysis, the device will activate the built-in camera to capture the scenery.
[0108] Input: The output data of step 2 (the parsed instructions).
[0109] Output: High resolution image data.
[0110] Step 4:
[0111] The images captured by the device are sent to a cloud server.
[0112] Specific operation: The device uses Wi-Fi or mobile data communication to send the captured image data to a cloud server.
[0113] Input: The output data of step 3 (image data).
[0114] Output: Image data is sent to the server.
[0115] Step 5:
[0116] The server stores the images in a database.
[0117] Specific operation: The server adds metadata (photo date and time and location) to the received image data and stores it in a database.
[0118] Input: The output data of step 4 (image data).
[0119] Output: Data stored in the database.
[0120] Step 6:
[0121] The user queries for a stored image.
[0122] Specific operation: The user speaks to the smart device and asks, "Show me the scenery from last week."
[0123] Input: Audio information.
[0124] Output: Recognized by the device as audio data.
[0125] Step 7:
[0126] The server retrieves relevant images from a database and transmits them to the terminal.
[0127] Specific operation: The server analyzes the question through voice recognition software, searches for and retrieves the corresponding image from the database, and sends the retrieved image to the terminal.
[0128] Input: The output data from step 6 (audio data) and the information in the database.
[0129] Output: The retrieved image data is sent to the terminal.
[0130] Step 8:
[0131] The device displays the image it receives.
[0132] Specific operation: The terminal displays the image data received from the server on the screen.
[0133] Input: The output data of step 7 (image data).
[0134] Output: The image that is displayed to the user.
[0135] (Application example 1)
[0136] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0137] In autonomous vehicles, when passengers want to record a particular experience or scenery, there is a need for a system that allows them to easily search and view the recorded images and videos later while intuitively and conveniently operating the system. In addition, there is an increasing need for passengers to record more experiences as they do not need to concentrate on driving. However, existing technologies to achieve this are limited, and there is no systematic solution.
[0138] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0139] In this invention, the server includes a means for receiving a shooting instruction by voice instruction or hand gesture, a means for transmitting the captured image to the cloud, and a means for retrieving the image stored in the cloud based on an associated question and displaying it on a display device in the vehicle. This allows passengers to easily record a particular experience by voice instruction or gesture only, and store it on the cloud for easy searching and viewing later.
[0140] The term "instruction to shoot" refers to an operation in which the user signals the device to start shooting by voice instruction or hand gesture.
[0141] "Voice instruction" refers to the act of a user speaking into a device to instruct the device to perform a specific action.
[0142] "Hand gestures" refer to instructions that the device recognizes hand movements and actions and uses them to perform specific actions.
[0143] "Cloud" refers to an online storage system for storing data on remote servers via the Internet.
[0144] "Means for sending to the cloud" refers to a method for transmitting images and data captured from a device to a cloud server via a network.
[0145] "Images stored in the cloud" refers to photos and videos that are sent via a network to a remote server and stored in online storage.
[0146] "Related questions" refer to keywords or phrases that users use when searching for specific images or experiences.
[0147] "Means of acquiring and displaying on a display device inside the vehicle" refers to a method of searching for images stored in the cloud and displaying them on a display or the like inside an autonomous vehicle.
[0148] "Display" refers to an electronic device for visually displaying information.
[0149] The main components of the system for realizing this application example are a camera, a microphone, a display, a cloud server, and a terminal with a communication function installed in the vehicle. The detailed configuration and operation of this system are described below.
[0150] 1. Hardware Configuration
[0151] Cameras in vehicles
[0152] The cameras will be installed at specific locations within the train and will be tasked with recording scenes and experiences that passengers indicate they want to capture. The cameras are capable of taking high-resolution images and have the ability to automatically adjust exposure and focus according to the environment.
[0153] microphone
[0154] The microphone is designed to clearly capture the user's voice commands while reducing noise inside the vehicle, allowing the voice recognition technology to accurately interpret them.
[0155] display
[0156] The display is installed on the dashboard or back seat monitor inside the vehicle and is used to visually display stored images and videos.
[0157] 2. Software Configuration
[0158] Voice Recognition Technology
[0159] For voice recognition, a library called "SpeechRecognition (registered trademark)" is used. Voice data is acquired from the microphone, analyzed, and instructions are determined. For example, if the user says "Take a picture," the voice is analyzed and the camera is activated.
[0160] Image capture and storage
[0161] Images captured by the camera are processed using the "OpenCV (registered trademark)" library and temporarily stored in the in-vehicle system before being sent to the cloud server via the network. At this time, the data is sent using the "Requests" library.
[0162] Cloud Server
[0163] The cloud server receives the captured images and stores them in a database, and when a user searches for a specific image, it also retrieves the relevant images from the database based on the relevant query.
[0164] 3. Processing Procedure
[0165] Shooting instructions
[0166] When a user issues a voice command to the camera to "take a picture," the microphone captures the voice and the voice is analyzed using voice recognition technology. The camera then starts up and takes a picture of the scene as instructed.
[0167] Cloud Upload
[0168] The captured images are temporarily stored in the vehicle and then transmitted to a cloud server using the network communication function. The cloud server stores the received images in a database.
[0169] Searching and viewing images
[0170] When the user issues a voice command, such as "Show me today's sunset," the command is again analyzed through voice recognition technology. The cloud server retrieves relevant images from the database and displays them on the vehicle's display.
[0171] Examples
[0172] For example, if a user says "Take a picture" while in the car, the camera will automatically capture the scenery and store the image in the cloud. Later, if the user says "Show me today's sunset," the relevant image will be retrieved from the cloud and displayed on the display in front of the passenger.
[0173] Examples of prompt statements
[0174] Imagine an app that takes a picture of a sunset with an in-car camera, and when the driver says "Show me today's sunset," the app displays the saved image on the in-car display. Please submit a detailed program outline.
[0175] The flow of the specific process in the application example 1 will be described with reference to FIG.
[0176] Step 1:
[0177] The user issues a voice command
[0178] The user says "take a picture" into the microphone, which causes voice data (input) to be captured by the microphone.
[0179] Step 2:
[0180] The device analyzes the voice command
[0181] The terminal analyzes the captured voice data using voice recognition technology (e.g., the SpeechRecognition library) and obtains (outputs) the voice command "Take a picture" as text data.
[0182] Step 3:
[0183] The camera starts and takes a picture
[0184] The device activates the camera based on the results of voice recognition. The camera captures the specified scene and obtains image data (input). The obtained image data is temporarily stored inside the device (output).
[0185] Step 4:
[0186] The device sends image data to the cloud.
[0187] The device transmits the temporarily stored image data to the cloud server using a network communication technology (e.g., the Requests library) (input). The cloud server stores the received image data in a database (output).
[0188] Step 5:
[0189] The user issues a search command
[0190] The user says into the microphone, "Show me today's sunset." This causes the microphone to capture voice data (input) of the search command.
[0191] Step 6:
[0192] The device parses the search instructions
[0193] The device analyzes the captured voice data using voice recognition technology and obtains the voice instruction "Show me today's sunset" as text data (output).
[0194] Step 7:
[0195] The cloud server searches for related images
[0196] The device sends the analyzed text data (input) to the cloud server, which searches the database for related image data that matches the text data and retrieves them (output).
[0197] Step 8:
[0198] Display the image captured by the device
[0199] Relevant image data obtained from the cloud server (input) is sent to the terminal, which then displays the received image data on a display inside the vehicle (output).
[0200] Furthermore, an emotion engine that estimates the emotion of the user may be combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.
[0201] The embodiment for carrying out the present invention comprises the following elements.
[0202] 1. Terminal: A device with an interface that accepts voice commands and hand gestures when a user wants to save a particular scene or experience.
[0203] 2. Shooting: The device analyzes voice commands and hand gestures to understand the command to shoot, activates the camera, and shoots the specified scene or experience.
[0204] 3. Transmission means: The terminal has a communication function for transmitting captured images to the cloud. Images can be transmitted to a server via a network.
[0205] 4. Server: Stores the received images in a database, accepts queries about the contents stored by the user, and retrieves and displays related captured images.
[0206] 5. Emotion engine: The server combines an emotion engine and utilizes technology to recognize the user’s emotions, such as voice and facial expressions. By analyzing the user’s emotions at the time of shooting and associating them with the captured image, the server can automatically classify and search for images based on emotions.
[0207] (Specific examples)
[0208] When a user uses the smart device 14 to implement the present invention, the smart device 14 becomes a terminal. The user commands the smart device 14 to take a picture by giving a voice command such as "save it" or by making a specific gesture. The smart device 14 analyzes the voice command or gesture and understands the command to take a picture. The camera starts up and captures the scene or experience as instructed. The captured image is sent to the cloud using the communication function of the smart device 14. The server stores and saves the received image in a database. When the user asks the smart device 14 a question about the saved content, the smart device 14 sends the question to the server. The server analyzes the received question, retrieves related captured images from the database, and sends them to the smart device 14. The smart device 14 displays the received captured image. The server also combines an emotion engine and uses technology to recognize the user's emotions, such as voice and facial expressions. By analyzing the user's emotions at the time of shooting and associating them with the captured image, automatic classification and search of images based on emotions are performed.
[0209] The above is an example of an embodiment of the present invention. Although specific devices and processing procedures may vary depending on the selection and implementation of the implementer, the present invention can be implemented based on this embodiment.
[0210] The process flow will be explained below.
[0211] Step 1: The user speaks into the terminal, saying "Save."
[0212] Step 2: The device receives the voice command and uses voice recognition technology to interpret the command.
[0213] Step 3: Based on the analysis results, the device understands the shooting instructions.
[0214] Step 4: The device activates its camera and captures the specified scene or experience.
[0215] Step 5: The device uses its communication functions to send the captured images to the cloud.
[0216] Step 6: The server stores the received images in a database.
[0217] Step 7: The user asks a question to the terminal.
[0218] Step 8: The terminal sends the query to the server.
[0219] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[0220] Step 10: The server transmits the captured image to the terminal.
[0221] Step 11: The terminal displays the captured image that it has received.
[0222] Step 12: The server uses the emotion engine to recognize the user's emotions, such as voice and facial expressions.
[0223] Step 13: The server analyzes the user's emotion at the time of taking the photograph and associates it with the photographed image.
[0224] Step 14: The server performs automatic classification and retrieval of images based on emotions.
[0225] Examples:
[0226] Step 1: The user issues a voice command to the smart device 14 saying "Save."
[0227] Step 2: The smart device 14 receives the voice instruction and parses the instruction using voice recognition technology.
[0228] Step 3: The smart device 14 understands the shooting instructions based on the analysis results.
[0229] Step 4: The smart device 14 launches a camera app and captures the instructed scene or experience.
[0230] Step 5: The smart device 14 uses its communication function to transmit the captured image to the cloud.
[0231] Step 6: The server stores the received images in a database.
[0232] Step 7: The user asks the smart device 14, "Show me pictures of playing in the park yesterday."
[0233] Step 8: The smart device 14 sends the query to the server.
[0234] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[0235] Step 10: The server transmits the acquired photographed image to the smart device 14.
[0236] Step 11: The smart device 14 displays the received captured image.
[0237] Step 12: The server uses the emotion engine to recognize the user's emotions, such as voice and facial expressions.
[0238] Step 13: The server analyzes the user's emotion at the time of taking the photograph and associates it with the photographed image.
[0239] Step 14: The server performs automatic classification and retrieval of images based on emotions.
[0240] The above is the specific process flow of this system. The user gives voice instructions or gestures, and the device analyzes them and performs appropriate processing, enabling the system to take photos, save images, answer questions, and classify and search images based on emotions.
[0241] Example 2
[0242] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."
[0243] Conventional image storage systems make it difficult for users to instantly record specific moments and easily search and display them when needed. Conventional systems also lack the ability to classify and search for images based on the user's emotions, and therefore lack the functionality to allow users to replay memories based on their emotions.
[0244] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for a user to give a shooting instruction by voice instruction or hand gesture, a means for a terminal to analyze the voice instruction or gesture, start a camera and shoot a scene, a means for a terminal to transmit the captured image to a cloud server, a means for a server to store the image in a database, a means for a server to accept a related question, obtain related images from the database and transmit them to the terminal, and a means for a server to analyze a user's emotion using an emotion engine and associate the emotion with the image. This allows a user to easily record a particular moment and later search and display images based on the emotion.
[0245] "User" refers to a person who uses the System to store and view specific scenes or experiences.
[0246] The term "terminal" refers to a device that allows a user to give shooting instructions by voice or gestures.
[0247] "Voice instruction" refers to an instruction given by voice to a terminal by a user.
[0248] A "hand gesture" refers to a user giving instructions to a terminal by performing a specific action.
[0249] "Camera" refers to a device built into a terminal that takes photos and videos at the user's command.
[0250] A "cloud server" refers to a remote server accessible via a network, which provides a place to store image data.
[0251] A "database" refers to a system built within a server that systematically stores captured images and related information.
[0252] A "related question" refers to an inquiry that a user makes to the system regarding images that he or she has previously saved and their contents.
[0253] An "emotion engine" is a technology that analyzes voice and facial expressions to recognize the user's emotions, and then classifies and searches for images based on the results.
[0254] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0255] This invention is a system that allows users to easily save specific scenes or experiences, and later search and display those images based on emotions. The system consists of a terminal operated by the user and a server installed on the cloud. The main hardware and software configurations are explained below.
[0256] Terminal
[0257] The terminal is a device that allows the user to give instructions to shoot using voice commands or hand gestures. Specifically, this applies to smartphones, tablets, smart glasses, etc. The terminal has the following main functions:
[0258] 1. Voice Recognition Function:
[0259] The terminal analyzes the user's voice instructions using well-known natural language processing (NLP) techniques.
[0260] 2. Gesture recognition:
[0261] The device is equipped with technology that uses cameras and sensors to recognize the user's hand movements and analyze gestures.
[0262] 3. Camera Function:
[0263] It has a built-in high-resolution camera, and after understanding the shooting instructions, it will capture the specified scene.
[0264] 4. Communication function:
[0265] The captured images are sent to a cloud server using networks such as Wi-Fi, 4G / 5G.
[0266] server
[0267] The server is installed in the cloud and has the following main functions:
[0268] 1. Database Management:
[0269] Image data is efficiently stored and saved in a database. The database can use Amazon Web Services (AWS (registered trademark)) RDS.
[0270] 2. Question analysis function:
[0271] The questions entered by the user are analyzed using natural language processing technology and related images are linked.
[0272] 3. Emotion Engine:
[0273] It uses voice and facial expression analysis technology to recognize user emotions, and uses that data to tag images with emotions for automatic classification and emotion-based search.
[0274] 4. Image sending function:
[0275] The image data retrieved in response to the user's question is sent to the terminal using encryption protocols such as SSL / TLS.
[0276] Examples
[0277] For example, suppose a user says to the device "Save it" during a family trip. The device uses voice recognition technology to analyze this voice command, activates the camera, and captures beautiful scenery. The captured image is then sent to a cloud server using the device's communication function. The sent image is stored in a database on the server. Later, when the user issues a voice command such as "Show me photos of the beach from last summer," the server analyzes the question, searches the database for the corresponding image, and sends it to the user's device. The device displays the received image on the screen. The server's emotion engine also recognizes emotions such as "joy" from the user's voice and facial expression when taking the photo, and associates and saves the emotion tag with the image, so that if the user commands "Show me photos of fun times," an image associated with the joyful emotion will be displayed.
[0278] Example prompts for generative AI models
[0279] "Please explain how the system works, allowing users to take pictures of specific scenes using voice commands and store them in the cloud."
[0280] "How can I use an emotion engine to automatically classify images based on user emotions?"
[0281] "Please explain with a concrete example the process of taking pictures using a smartphone and saving them to the cloud."
[0282] The above is a detailed embodiment of the present invention. This system allows users to easily operate it and store and search images based on emotions, greatly improving the user experience.
[0283] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0284] Step 1:
[0285] The user issues a command to take a picture. Specifically, the user issues a voice command to the smartphone saying "save it" or performs a specific gesture. The input data is a voice signal or a gesture movement, which is input to the terminal. The output is a flag indicating that the command to take a picture has been recognized.
[0286] Step 2:
[0287] The device analyzes the voice command and gestures. The device analyzes the voice command using well-known natural language processing techniques. It also recognizes certain actions using gesture recognition techniques. The input data is the user's voice or gestures, and the output is the analyzed command. Specifically, the voice command "save" is recognized, and an instruction to start the camera is issued.
[0288] Step 3:
[0289] The device activates the camera and captures the scene. The camera understands the specified instructions and takes the picture. The input data is the parsed instructions, and the output is the captured image data. For example, a beautiful landscape during a family trip is captured as a high-resolution photo.
[0290] Step 4:
[0291] The device sends the captured image to the cloud server. The device uses Wi-Fi or 4G / 5G networks to send encrypted image data to the cloud server. The input data is the captured image data, and the output is the image data sent to the cloud server.
[0292] Step 5:
[0293] The server stores the images in a database. The server organizes the image data it receives and stores it in a database. The input data is image data sent via the cloud, and the output is image entries stored in the database. For example, AWS's RDS is used to efficiently store and manage image data.
[0294] Step 6:
[0295] The user asks a question related to the saved content. The user issues a voice command to the smartphone, such as "Show me a photo of the beach from last summer." The input data is the voice command to ask the question, and the output is the content of the question to the server.
[0296] Step 7:
[0297] The server analyzes the question and retrieves related images. Using natural language processing technology, the server analyzes the user's question and searches the database for relevant image data. The input data is the question based on the voice command, and the output is related image data. For example, "photos of the sea from last summer" is searched for.
[0298] Step 8:
[0299] The server sends the relevant images to the user's device. The server encrypts the image data it finds and sends it to the user's device. The input data is the image data retrieved from the database, and the output is the image data sent to the device.
[0300] Step 9:
[0301] The user's device displays the image. The device displays the received image to the user. The input data is the image data received from the server, and the output is the image displayed on the device's display. For example, a photo of a summer beach is displayed on the screen of a smartphone.
[0302] Step 10:
[0303] The server performs emotion analysis using an emotion engine and associates it with the image. The server analyzes the voice and facial expression data to recognize the user's emotion. The image data with the emotion tag is saved for future searches. The input data is the voice and facial expression data at the time of shooting, and the output is image data with the emotion tag. For example, if the user is feeling "joy," that emotion tag is associated with the image.
[0304] (Application example 2)
[0305] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0306] In modern commerce, when consumers visit a store and select products, they need a way to easily save specific scenes and favorite products for later reference. Furthermore, there is no system that can classify and search saved information based on consumers' emotions. This leaves users lacking a way to efficiently manage their shopping experience and receive emotion-based feedback.
[0307] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0308] In this invention, the server includes means for accepting shooting instructions by voice instructions or hand gestures, means for transmitting the captured images to the cloud, means for retrieving and displaying the images stored in the cloud based on associated questions, means for analyzing the user's emotions, and means for organizing and searching for images based on the emotion data.
[0309] This allows users to easily save specific scenes or products and later categorize and search for images based on their emotions, resulting in a more fulfilling shopping experience.
[0310] A "capture command" refers to a voice command or hand gesture made by a user to record a particular scene or experience.
[0311] "Voice instruction" refers to a method in which a user instructs a system to perform a specific operation using voice.
[0312] "Hand gestures" refers to the way in which a user uses hand movements to instruct a system to perform certain actions.
[0313] The term "captured image" refers to a still image captured by a camera based on a user's instruction to capture a picture.
[0314] "Cloud" refers to a technological infrastructure that stores and manages data on remote servers provided via the Internet.
[0315] "Communication means" refers to the technical means and devices for transmitting and receiving data.
[0316] "User's emotions" refers to psychological reactions and sensations analyzed from the user's voice and facial expressions.
[0317] "Emotion data" refers to data obtained by analyzing information regarding a user's emotions.
[0318] An "emotion engine" is a technology that automatically identifies and analyzes a user's emotions from their voice, facial expressions, etc., and makes the results available within the system.
[0319] "Server" refers to a computer system on a network that provides and manages information and data.
[0320] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, an embodiment of the present invention will be described.
[0321] This invention uses a smartphone, smart glasses, or other smart devices as a terminal. The terminal has a built-in camera and microphone, and an interface that accepts voice instructions and hand gestures. This allows users to easily capture and save specific scenes and experiences.
[0322] The device uses voice recognition technology to analyze the user's voice instructions. The "SpeechRecognition" library is used for voice recognition. For example, when the user says "save it," the device will start the camera and take a picture of the scene. It also uses the "OpenCV" library to recognize hand gestures, so it is possible to give a shooting command using specific gestures.
[0323] The captured images are sent to the cloud using the device's communication function. The transmission to the cloud is done using the "Requests" library, and the captured images are stored on the cloud server.
[0324] Images stored on the cloud are retrieved and displayed in response to a user's question. A question is accepted by the device and sent to the cloud server. The cloud server analyzes the question, retrieves the associated image, and sends it to the device.
[0325] In addition, the system analyzes the voice data emitted by the user when taking a photo and uses an emotion engine to recognize the user's emotions. The emotion engine receives the voice data as input and analyzes the emotions. The emotion analysis results are associated with the image and stored in the cloud.
[0326] Image sorting and retrieval based on emotions is performed by the emotion engine on the cloud server. For example, it is possible to search for only images of scenes that the user felt were "fun."
[0327] Examples:
[0328] When a user is shopping in a brick-and-mortar store, they can say "Save this" when they see a product they like. The device that receives this command activates the camera and takes a picture of the scene. This picture is immediately sent to the cloud and saved. If the user then asks, "Show me some fun products," the cloud server uses the emotion engine to search for images associated with "fun" and sends them to the device, allowing the user to check the product again.
[0329] Example prompt:
[0330] "I want to create an application that allows users to easily save their favorite products or specific displays they find in physical stores. The user can take a picture by voice command or gesture, and the captured image will be sent to the cloud. Also, the application will analyze the user's emotions based on the voice data, and later create a product list based on the emotions. Please generate a program for this application."
[0331] The flow of the specific process in the application example 2 will be described with reference to FIG.
[0332] Step 1:
[0333] The device accepts voice commands or hand gestures. The user issues a voice command such as "save" or makes a specific hand gesture. The device receives voice or gestures as input using a microphone or camera, and analyzes them using voice recognition technology (SpeechRecognition library) or gesture recognition technology (OpenCV library). As a result of the analysis, it determines whether there is an instruction to take a photo.
[0334] Input: User's voice commands or hand gestures
[0335] Output: Whether or not shooting is instructed
[0336] Step 2:
[0337] The device starts up the camera and takes a picture of the specified scene. The captured image data is generated and temporarily stored in the device's memory.
[0338] Input:Shooting instructions
[0339] Output: Captured image data
[0340] Step 3:
[0341] The device sends the captured image data to the cloud. The captured image data is sent to the cloud server using the communication function (Requests library). The cloud server stores the received image data in a database.
[0342] Input: Photographed image data
[0343] Output: Image data stored in the cloud
[0344] Step 4:
[0345] The user asks a question about an image stored in the cloud. The user enters the question into the smart device, which then sends the question to the cloud server.
[0346] Input: User question
[0347] Output: A query request to the cloud server
[0348] Step 5:
[0349] The cloud server acquires information related to the stored images, searches the database for the corresponding image data based on the user's question, and transmits the searched image data to the terminal.
[0350] Input: A query request to the cloud server
[0351] Output: Retrieved image data
[0352] Step 6:
[0353] The terminal displays the searched image data. The terminal that receives the image data sent from the cloud server displays it on the screen and provides it to the user.
[0354] Input: Retrieved image data
[0355] Output: Display the image to the user
[0356] Step 7:
[0357] The cloud server uses an emotion engine to analyze the user's voice data and recognize the user's emotions. The emotion engine receives the voice data as input and executes an emotion analysis process. As a result of the analysis, it generates user emotion data and stores it in association with the image data.
[0358] Input: User's voice data
[0359] Output: Emotion data and associated image data
[0360] Step 8:
[0361] The cloud server organizes and searches image data based on the user's emotion data. For example, if a user asks, "Show me some fun products," the cloud server uses the emotion engine to search for images associated with "fun" and sends them to the device.
[0362] Input: emotion data and user questions
[0363] Output: Image data of sentiment-based search results
[0364] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0365] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0366] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0367] [Second embodiment]
[0368] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0369] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0370] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[0371] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0372] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[0373] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[0374] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[0375] Fig. 4 shows an example of main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0376] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0377] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0378] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0379] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".
[0380] The embodiment for carrying out the present invention comprises the following elements.
[0381] 1. Terminal: A device with an interface that accepts voice commands and hand gestures when a user wants to save a particular scene or experience.
[0382] 2. Shooting: The device analyzes voice commands and hand gestures to understand the command to shoot, activates the camera, and shoots the specified scene or experience.
[0383] 3. Transmission means: The terminal has a communication function for transmitting captured images to the cloud. Images can be transmitted to a server via a network.
[0384] 4. Server: Stores the received images in a database, accepts queries about the contents stored by the user, and retrieves and displays related captured images.
[0385] (Specific examples)
[0386] When a user uses the smart glasses 214 to implement the present invention, the smart glasses 214 become a terminal. The user commands the smart glasses 214 to take a picture by giving a voice command such as "save it" or by making a specific gesture. The smart glasses 214 analyze the voice command or gesture and understand the command to take a picture. The camera starts up and captures the instructed scene or experience. The captured image is sent to the cloud using the communication function of the smart glasses 214. The server stores and saves the received image in a database. When the user asks the smart glasses 214 a question about the saved content, the smart glasses 214 sends the question to the server. The server analyzes the received question, retrieves the related captured image from the database, and sends it to the smart glasses 214. The smart glasses 214 display the received captured image to provide the user with support for recollection.
[0387] The above is an example of an embodiment of the present invention. Although specific devices and processing procedures may vary depending on the selection and implementation of the implementer, the present invention can be implemented based on this embodiment.
[0388] The process flow will be explained below.
[0389] Step 1: The user speaks into the terminal, saying "Save."
[0390] Step 2: The device receives the voice command and uses voice recognition technology to interpret the command.
[0391] Step 3: Based on the analysis results, the device understands the shooting instructions.
[0392] Step 4: The device activates its camera and captures the specified scene or experience.
[0393] Step 5: The device uses its communication functions to send the captured images to the cloud.
[0394] Step 6: The server stores the received images in a database.
[0395] Step 7: The user asks a question to the terminal.
[0396] Step 8: The terminal sends the query to the server.
[0397] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[0398] Step 10: The server transmits the captured image to the terminal.
[0399] Step 11: The terminal displays the captured image that it has received.
[0400] Examples:
[0401] Step 1: The user speaks into the smart glasses 214 saying "save."
[0402] Step 2: The smart glasses 214 receive the voice instruction and parse the instruction using voice recognition technology.
[0403] Step 3: The smart glasses 214 understand the shooting instructions based on the analysis results.
[0404] Step 4: The smart glasses 214 launch a camera app and capture the indicated scene or experience.
[0405] Step 5: The smart glasses 214 use their communication capabilities to send the captured images to the cloud.
[0406] Step 6: The server stores the received images in a database.
[0407] Step 7: The user asks the smart glasses 214, "Show me pictures of playing at the park yesterday."
[0408] Step 8: The smart glasses 214 send the query to the server.
[0409] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[0410] Step 10: The server sends the captured image to the smart glasses 214.
[0411] Step 11: The smart glasses 214 display the received captured image.
[0412] The above is the specific process flow of this system. The user gives voice instructions or gestures, and the device analyzes them and performs appropriate processing, enabling the system to take photos, save images, and answer questions.
[0413] Example 1
[0414] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".
[0415] There is a demand for users to be able to easily save specific scenes and experiences and easily search and display the saved information at a later date. However, with conventional technologies, it is difficult to intuitively give shooting instructions using voice commands or hand gestures, or to efficiently retrieve images saved in the cloud based on a question.
[0416] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0417] In this invention, the server includes a means for receiving a shooting instruction by voice instruction or hand gesture, a means for analyzing the received shooting instruction to start the camera and shoot the scene, a means for transmitting the shot image to the cloud server via the network, a means for storing the received image in a database, and a means for receiving a question about the saved image from the user and acquiring and displaying related images, thereby enabling the user to intuitively store a particular scene or experience and efficiently recall it later.
[0418] A "voice instruction" refers to a specific operation or instruction given by a user to a terminal using voice.
[0419] A "hand gesture" is a specific operation or instruction given by a user to a terminal using hand movements.
[0420] A "photography instruction" is an instruction given by a user to a terminal to capture a specific scene or experience with a camera.
[0421] "Analysis" is the process of understanding the voice instructions and hand gestures received by the device and converting them into appropriate actions.
[0422] "Starting the camera" means starting up the camera function of the device and making it ready to take pictures.
[0423] A "cloud server" is a server used to store and process data over the Internet.
[0424] A database is a structured storage system that organizes and stores large amounts of data and allows for efficient search and retrieval.
[0425] "Network communication technology" refers to technology used for data communication, including Wi-Fi and mobile data.
[0426] A "question" is a query made by a user to obtain information about a previously stored image.
[0427] "Retrieval" is the process of searching and retrieving specific information from a database.
[0428] "Display" means to output the acquired information on the terminal screen so that the user can confirm it.
[0429] The embodiment of the present invention provides a system that allows a user to easily save a particular scene or experience, and easily search and display it later. A specific implementation method of this system will be described below.
[0430] A user uses a dedicated terminal to give voice instructions and hand gestures. This terminal is equipped with voice recognition technology and gesture recognition technology, and for example, voice recognition software provided by each company is used. For gesture recognition, a motion sensor combined with the terminal's camera is used.
[0431] The device analyzes the user's voice instructions and gestures and activates the camera based on those. The camera is capable of capturing high-resolution images. The captured images are sent to a cloud server using the device's network communication function (Wi-Fi or mobile data communication). In this case, HTTP / HTTPS is generally used as the communication protocol.
[0432] The cloud server appropriately processes the received image data, adds metadata (such as the date and time of shooting and the location), and stores the data in a database. The database uses storage provided by each company.
[0433] The user can later ask questions about the stored images through the device, for example by issuing a voice command such as "Show me the scenery from last week." The question is converted into text using the device's voice recognition technology and sent to the server.
[0434] The server analyzes the user's question, searches for and retrieves relevant images from the database, and then transmits the retrieved images to the terminal, which displays the received images on its screen and provides them to the user.
[0435] Examples:
[0436] When a user visits a park, they give a voice command to their smart device saying, "Save this view," or perform a specific hand gesture. The smart device uses its built-in voice recognition software to analyze the voice, activates the camera, and captures the view. The captured image is sent to a cloud server using its Wi-Fi function. The server adds metadata to the received image and stores it in a database.
[0437] A few days later, the user asks the smart device, "Show me the scenery from last week." The smart device converts this voice to text and sends it to the server. The server searches the database for the corresponding image and sends it to the smart device. The device displays the image on the screen, allowing the user to reminisce about the scenery from last week.
[0438] Example prompt:
[0439] Design a system that allows a user to take pictures of scenery in a park and store them for later review. Include instructions for using voice commands or hand gestures to take pictures, store the data in the cloud, and play the stored images at a later date.
[0440] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0441] Step 1:
[0442] The user gives a voice command or a gesture.
[0443] Specific action: The user speaks to the smart device, saying "Save this view," or makes a specific hand gesture.
[0444] Input: Voice information or gesture data.
[0445] Output: Recognized by the terminal as audio or video data.
[0446] Step 2:
[0447] The device analyzes the voice command or gesture.
[0448] Specific operation: The voice recognition software in the terminal analyzes the voice, and the gesture recognition software analyzes the video data.
[0449] Input: The output data of step 1 (audio or video data).
[0450] Output: Parsed instruction (e.g. "Save the view").
[0451] Step 3:
[0452] The device will activate the camera and take a picture of the specified scene.
[0453] Specific operation: Based on the results of voice or gesture analysis, the device will activate the built-in camera to capture the scenery.
[0454] Input: The output data of step 2 (the parsed instructions).
[0455] Output: High resolution image data.
[0456] Step 4:
[0457] The images captured by the device are sent to a cloud server.
[0458] Specific operation: The device uses Wi-Fi or mobile data communication to send the captured image data to a cloud server.
[0459] Input: The output data of step 3 (image data).
[0460] Output: Image data is sent to the server.
[0461] Step 5:
[0462] The server stores the images in a database.
[0463] Specific operation: The server adds metadata (photo date and time and location) to the received image data and stores it in a database.
[0464] Input: The output data of step 4 (image data).
[0465] Output: Data stored in the database.
[0466] Step 6:
[0467] The user queries for a stored image.
[0468] Specific operation: The user speaks to the smart device and asks, "Show me the scenery from last week."
[0469] Input: Audio information.
[0470] Output: Recognized by the device as audio data.
[0471] Step 7:
[0472] The server retrieves relevant images from a database and transmits them to the terminal.
[0473] Specific operation: The server analyzes the question through voice recognition software, searches for and retrieves the corresponding image from the database, and sends the retrieved image to the terminal.
[0474] Input: The output data from step 6 (audio data) and the information in the database.
[0475] Output: The retrieved image data is sent to the terminal.
[0476] Step 8:
[0477] The device displays the image it receives.
[0478] Specific operation: The terminal displays the image data received from the server on the screen.
[0479] Input: The output data of step 7 (image data).
[0480] Output: The image that is displayed to the user.
[0481] (Application example 1)
[0482] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".
[0483] In autonomous vehicles, when passengers want to record a particular experience or scenery, there is a need for a system that allows them to easily search and view the recorded images and videos later while intuitively and conveniently operating the system. In addition, there is an increasing need for passengers to record more experiences as they do not need to concentrate on driving. However, existing technologies to achieve this are limited, and there is no systematic solution.
[0484] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0485] In this invention, the server includes a means for receiving a shooting instruction by voice instruction or hand gesture, a means for transmitting the captured image to the cloud, and a means for retrieving the image stored in the cloud based on an associated question and displaying it on a display device in the vehicle. This allows passengers to easily record a particular experience by voice instruction or gesture only, and store it on the cloud for easy searching and viewing later.
[0486] The term "instruction to shoot" refers to an operation in which the user signals the device to start shooting by voice instruction or hand gesture.
[0487] "Voice instruction" refers to the act of a user speaking into a device to instruct the device to perform a specific action.
[0488] "Hand gestures" refer to instructions that the device recognizes hand movements and actions and uses them to perform specific actions.
[0489] "Cloud" refers to an online storage system for storing data on remote servers via the Internet.
[0490] "Means for sending to the cloud" refers to a method for transmitting images and data captured from a device to a cloud server via a network.
[0491] "Images stored in the cloud" refers to photos and videos that are sent via a network to a remote server and stored in online storage.
[0492] "Related questions" refer to keywords or phrases that users use when searching for specific images or experiences.
[0493] "Means of acquiring and displaying on a display device inside the vehicle" refers to a method of searching for images stored in the cloud and displaying them on a display or the like inside an autonomous vehicle.
[0494] "Display" refers to an electronic device for visually displaying information.
[0495] The main components of the system for realizing this application example are a camera, a microphone, a display, a cloud server, and a terminal with a communication function installed in the vehicle. The detailed configuration and operation of this system are described below.
[0496] 1. Hardware Configuration
[0497] Cameras in vehicles
[0498] The cameras will be installed at specific locations within the train and will be tasked with recording scenes and experiences that passengers indicate they want to capture. The cameras are capable of taking high-resolution images and have the ability to automatically adjust exposure and focus according to the environment.
[0499] microphone
[0500] The microphone is designed to clearly capture the user's voice commands while reducing noise inside the vehicle, allowing the voice recognition technology to accurately interpret them.
[0501] display
[0502] The display is installed on the dashboard or back seat monitor inside the vehicle and is used to visually display stored images and videos.
[0503] 2. Software Configuration
[0504] Voice Recognition Technology
[0505] For voice recognition, a library called "SpeechRecognition" is used. Voice data is acquired from the microphone, analyzed, and instructions are determined. For example, if the user says "Take a picture," the voice is analyzed and the camera is activated.
[0506] Image capture and storage
[0507] Images captured by the camera are processed using the "OpenCV" library and temporarily stored in the in-vehicle system before being sent to a cloud server via the network, using the "Requests" library to send the data.
[0508] Cloud Server
[0509] The cloud server receives the captured images and stores them in a database, and when a user searches for a specific image, it also retrieves the relevant images from the database based on the relevant query.
[0510] 3. Processing Procedure
[0511] Shooting instructions
[0512] When a user issues a voice command to the camera to "take a picture," the microphone captures the voice and the voice is analyzed using voice recognition technology. The camera then starts up and takes a picture of the scene as instructed.
[0513] Cloud Upload
[0514] The captured images are temporarily stored in the vehicle and then transmitted to a cloud server using the network communication function. The cloud server stores the received images in a database.
[0515] Searching and viewing images
[0516] When the user issues a voice command, such as "Show me today's sunset," the command is again analyzed through voice recognition technology. The cloud server retrieves relevant images from the database and displays them on the vehicle's display.
[0517] Examples
[0518] For example, if a user says "Take a picture" while in the car, the camera will automatically capture the scenery and store the image in the cloud. Later, if the user says "Show me today's sunset," the relevant image will be retrieved from the cloud and displayed on the display in front of the passenger.
[0519] Examples of prompt statements
[0520] Imagine an app that takes a picture of a sunset with an in-car camera, and when the driver says "Show me today's sunset," the app displays the saved image on the in-car display. Please submit a detailed program outline.
[0521] The flow of the specific process in the application example 1 will be described with reference to FIG.
[0522] Step 1:
[0523] The user issues a voice command
[0524] The user says "take a picture" into the microphone, which causes voice data (input) to be captured by the microphone.
[0525] Step 2:
[0526] The device analyzes the voice command
[0527] The terminal analyzes the captured voice data using voice recognition technology (e.g., the SpeechRecognition library) and obtains (outputs) the voice command "Take a picture" as text data.
[0528] Step 3:
[0529] The camera starts and takes a picture
[0530] The device activates the camera based on the results of voice recognition. The camera captures the specified scene and obtains image data (input). The obtained image data is temporarily stored inside the device (output).
[0531] Step 4:
[0532] The device sends image data to the cloud.
[0533] The device transmits the temporarily stored image data to the cloud server using a network communication technology (e.g., the Requests library) (input). The cloud server stores the received image data in a database (output).
[0534] Step 5:
[0535] The user issues a search command
[0536] The user says into the microphone, "Show me today's sunset." This causes the microphone to capture voice data (input) of the search command.
[0537] Step 6:
[0538] The device parses the search instructions
[0539] The device analyzes the captured voice data using voice recognition technology and obtains the voice instruction "Show me today's sunset" as text data (output).
[0540] Step 7:
[0541] The cloud server searches for related images
[0542] The device sends the analyzed text data (input) to the cloud server, which searches the database for related image data that matches the text data and retrieves them (output).
[0543] Step 8:
[0544] Display the image captured by the device
[0545] Relevant image data obtained from the cloud server (input) is sent to the terminal, which then displays the received image data on a display inside the vehicle (output).
[0546] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0547] The embodiment for carrying out the present invention comprises the following elements.
[0548] 1. Terminal: A device with an interface that accepts voice commands and hand gestures when a user wants to save a particular scene or experience.
[0549] 2. Shooting: The device analyzes voice commands and hand gestures to understand the command to shoot, activates the camera, and shoots the specified scene or experience.
[0550] 3. Transmission means: The terminal has a communication function for transmitting captured images to the cloud. Images can be transmitted to a server via a network.
[0551] 4. Server: Stores the received images in a database, accepts queries about the contents stored by the user, and retrieves and displays related captured images.
[0552] 5. Emotion engine: The server combines an emotion engine and utilizes technology to recognize the user’s emotions, such as voice and facial expressions. By analyzing the user’s emotions at the time of shooting and associating them with the captured image, the server can automatically classify and search for images based on emotions.
[0553] (Specific examples)
[0554] When a user uses the smart glasses 214 to implement the present invention, the smart glasses 214 become a terminal. The user commands the smart glasses 214 to take a picture by giving a voice command such as "save it" or by making a specific gesture. The smart glasses 214 analyze the voice command or gesture and understand the command to take a picture. The camera starts up and captures the scene or experience as instructed. The captured image is sent to the cloud using the communication function of the smart glasses 214. The server stores and saves the received image in a database. When the user asks the smart glasses 214 a question about the saved content, the smart glasses 214 sends the question to the server. The server analyzes the received question, retrieves related captured images from the database, and sends them to the smart glasses 214. The smart glasses 214 display the received captured images. The server also combines an emotion engine and utilizes technology to recognize the user's emotions, such as voice and facial expressions. By analyzing the user's emotions at the time of shooting and associating them with the captured image, automatic classification and search of images based on emotions are performed.
[0555] The above is an example of an embodiment of the present invention. Although specific devices and processing procedures may vary depending on the selection and implementation of the implementer, the present invention can be implemented based on this embodiment.
[0556] The process flow will be explained below.
[0557] Step 1: The user speaks into the terminal, saying "Save."
[0558] Step 2: The device receives the voice command and uses voice recognition technology to interpret the command.
[0559] Step 3: Based on the analysis results, the device understands the shooting instructions.
[0560] Step 4: The device activates its camera and captures the specified scene or experience.
[0561] Step 5: The device uses its communication functions to send the captured images to the cloud.
[0562] Step 6: The server stores the received images in a database.
[0563] Step 7: The user asks a question to the terminal.
[0564] Step 8: The terminal sends the query to the server.
[0565] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[0566] Step 10: The server transmits the captured image to the terminal.
[0567] Step 11: The terminal displays the captured image that it has received.
[0568] Step 12: The server uses the emotion engine to recognize the user's emotions, such as voice and facial expressions.
[0569] Step 13: The server analyzes the user's emotion at the time of taking the photograph and associates it with the photographed image.
[0570] Step 14: The server performs automatic classification and retrieval of images based on emotions.
[0571] Examples:
[0572] Step 1: The user speaks into the smart glasses 214 saying "save."
[0573] Step 2: The smart glasses 214 receive the voice instruction and parse the instruction using voice recognition technology.
[0574] Step 3: The smart glasses 214 understand the shooting instructions based on the analysis results.
[0575] Step 4: The smart glasses 214 launch a camera app and capture the indicated scene or experience.
[0576] Step 5: The smart glasses 214 use their communication capabilities to send the captured images to the cloud.
[0577] Step 6: The server stores the received images in a database.
[0578] Step 7: The user asks the smart glasses 214, "Show me pictures of playing at the park yesterday."
[0579] Step 8: The smart glasses 214 send the query to the server.
[0580] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[0581] Step 10: The server sends the captured image to the smart glasses 214.
[0582] Step 11: The smart glasses 214 display the received captured image.
[0583] Step 12: The server uses the emotion engine to recognize the user's emotions, such as voice and facial expressions.
[0584] Step 13: The server analyzes the user's emotion at the time of taking the photograph and associates it with the photographed image.
[0585] Step 14: The server performs automatic classification and retrieval of images based on emotions.
[0586] The above is the specific process flow of this system. The user gives voice instructions or gestures, and the device analyzes them and performs appropriate processing, enabling the system to take photos, save images, answer questions, and classify and search images based on emotions.
[0587] Example 2
[0588] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".
[0589] Conventional image storage systems make it difficult for users to instantly record specific moments and easily search and display them when needed. Conventional systems also lack the ability to classify and search for images based on the user's emotions, and therefore lack the functionality to allow users to replay memories based on their emotions.
[0590] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for a user to give a shooting instruction by voice instruction or hand gesture, a means for a terminal to analyze the voice instruction or gesture, start a camera and shoot a scene, a means for a terminal to transmit the captured image to a cloud server, a means for a server to store the image in a database, a means for a server to accept a related question, obtain related images from the database and transmit them to the terminal, and a means for a server to analyze a user's emotion using an emotion engine and associate the emotion with the image. This allows a user to easily record a particular moment and later search and display images based on the emotion.
[0591] "User" refers to a person who uses the System to store and view specific scenes or experiences.
[0592] The term "terminal" refers to a device that allows a user to give shooting instructions by voice or gestures.
[0593] "Voice instruction" refers to an instruction given by voice to a terminal by a user.
[0594] A "hand gesture" refers to a user giving instructions to a terminal by performing a specific action.
[0595] "Camera" refers to a device built into a terminal that takes photos and videos at the user's command.
[0596] A "cloud server" refers to a remote server accessible via a network, which provides a place to store image data.
[0597] A "database" refers to a system built within a server that systematically stores captured images and related information.
[0598] A "related question" refers to an inquiry that a user makes to the system regarding images that he or she has previously saved and their contents.
[0599] An "emotion engine" is a technology that analyzes voice and facial expressions to recognize the user's emotions, and then classifies and searches for images based on the results.
[0600] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0601] This invention is a system that allows users to easily save specific scenes or experiences, and later search and display those images based on emotions. The system consists of a terminal operated by the user and a server installed on the cloud. The main hardware and software configurations are explained below.
[0602] Terminal
[0603] The terminal is a device that allows the user to give instructions to shoot using voice commands or hand gestures. Specifically, this applies to smartphones, tablets, smart glasses, etc. The terminal has the following main functions:
[0604] 1. Voice Recognition Function:
[0605] The terminal analyzes the user's voice instructions using well-known natural language processing (NLP) techniques.
[0606] 2. Gesture recognition:
[0607] The device is equipped with technology that uses cameras and sensors to recognize the user's hand movements and analyze gestures.
[0608] 3. Camera Function:
[0609] It has a built-in high-resolution camera, and after understanding the shooting instructions, it will capture the specified scene.
[0610] 4. Communication function:
[0611] The captured images are sent to a cloud server using networks such as Wi-Fi, 4G / 5G.
[0612] server
[0613] The server is installed in the cloud and has the following main functions:
[0614] 1. Database Management:
[0615] Image data is efficiently stored and saved in a database. The database can use Amazon Web Services (AWS) RDS.
[0616] 2. Question analysis function:
[0617] The questions entered by the user are analyzed using natural language processing technology and related images are linked.
[0618] 3. Emotion Engine:
[0619] It uses voice and facial expression analysis technology to recognize user emotions, and uses that data to tag images with emotions for automatic classification and emotion-based search.
[0620] 4. Image sending function:
[0621] The image data retrieved in response to the user's question is sent to the terminal using encryption protocols such as SSL / TLS.
[0622] Examples
[0623] For example, suppose a user says to the device "Save it" during a family trip. The device uses voice recognition technology to analyze this voice command, activates the camera, and captures beautiful scenery. The captured image is then sent to a cloud server using the device's communication function. The sent image is stored in a database on the server. Later, when the user issues a voice command such as "Show me photos of the beach from last summer," the server analyzes the question, searches the database for the corresponding image, and sends it to the user's device. The device displays the received image on the screen. The server's emotion engine also recognizes emotions such as "joy" from the user's voice and facial expression when taking the photo, and associates and saves the emotion tag with the image, so that if the user commands "Show me photos of fun times," an image associated with the joyful emotion will be displayed.
[0624] Example prompts for generative AI models
[0625] "Please explain how the system works, allowing users to take pictures of specific scenes using voice commands and store them in the cloud."
[0626] "How can I use an emotion engine to automatically classify images based on user emotions?"
[0627] "Please explain with a concrete example the process of taking pictures using a smartphone and saving them to the cloud."
[0628] The above is a detailed embodiment of the present invention. This system allows users to easily operate it and store and search images based on emotions, greatly improving the user experience.
[0629] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0630] Step 1:
[0631] The user issues a command to take a picture. Specifically, the user issues a voice command to the smartphone saying "save it" or performs a specific gesture. The input data is a voice signal or a gesture movement, which is input to the terminal. The output is a flag indicating that the command to take a picture has been recognized.
[0632] Step 2:
[0633] The device analyzes the voice command and gestures. The device analyzes the voice command using well-known natural language processing techniques. It also recognizes certain actions using gesture recognition techniques. The input data is the user's voice or gestures, and the output is the analyzed command. Specifically, the voice command "save" is recognized, and an instruction to start the camera is issued.
[0634] Step 3:
[0635] The device activates the camera and captures the scene. The camera understands the specified instructions and takes the picture. The input data is the parsed instructions, and the output is the captured image data. For example, a beautiful landscape during a family trip is captured as a high-resolution photo.
[0636] Step 4:
[0637] The device sends the captured image to the cloud server. The device uses Wi-Fi or 4G / 5G networks to send encrypted image data to the cloud server. The input data is the captured image data, and the output is the image data sent to the cloud server.
[0638] Step 5:
[0639] The server stores the images in a database. The server organizes the image data it receives and stores it in a database. The input data is image data sent via the cloud, and the output is image entries stored in the database. For example, AWS's RDS is used to efficiently store and manage image data.
[0640] Step 6:
[0641] The user asks a question related to the saved content. The user issues a voice command to the smartphone, such as "Show me a photo of the beach from last summer." The input data is the voice command to ask the question, and the output is the content of the question to the server.
[0642] Step 7:
[0643] The server analyzes the question and retrieves related images. Using natural language processing technology, the server analyzes the user's question and searches the database for relevant image data. The input data is the question based on the voice command, and the output is related image data. For example, "photos of the sea from last summer" is searched for.
[0644] Step 8:
[0645] The server sends the relevant images to the user's device. The server encrypts the image data it finds and sends it to the user's device. The input data is the image data retrieved from the database, and the output is the image data sent to the device.
[0646] Step 9:
[0647] The user's device displays the image. The device displays the received image to the user. The input data is the image data received from the server, and the output is the image displayed on the device's display. For example, a photo of a summer beach is displayed on the screen of a smartphone.
[0648] Step 10:
[0649] The server performs emotion analysis using an emotion engine and associates it with the image. The server analyzes the voice and facial expression data to recognize the user's emotion. The image data with the emotion tag is saved for future searches. The input data is the voice and facial expression data at the time of shooting, and the output is image data with the emotion tag. For example, if the user is feeling "joy," that emotion tag is associated with the image.
[0650] (Application example 2)
[0651] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".
[0652] In modern commerce, when consumers visit a store and select products, they need a way to easily save specific scenes and favorite products for later reference. Furthermore, there is no system that can classify and search saved information based on consumers' emotions. This leaves users lacking a way to efficiently manage their shopping experience and receive emotion-based feedback.
[0653] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0654] In this invention, the server includes means for accepting shooting instructions by voice instructions or hand gestures, means for transmitting the captured images to the cloud, means for retrieving and displaying the images stored in the cloud based on associated questions, means for analyzing the user's emotions, and means for organizing and searching for images based on the emotion data.
[0655] This allows users to easily save specific scenes or products and later categorize and search for images based on their emotions, resulting in a more fulfilling shopping experience.
[0656] A "capture command" refers to a voice command or hand gesture made by a user to record a particular scene or experience.
[0657] "Voice instruction" refers to a method in which a user instructs a system to perform a specific operation using voice.
[0658] "Hand gestures" refers to the way in which a user uses hand movements to instruct a system to perform certain actions.
[0659] The term "captured image" refers to a still image captured by a camera based on a user's instruction to capture a picture.
[0660] "Cloud" refers to a technological infrastructure that stores and manages data on remote servers provided via the Internet.
[0661] "Communication means" refers to the technical means and devices for transmitting and receiving data.
[0662] "User's emotions" refers to psychological reactions and sensations analyzed from the user's voice and facial expressions.
[0663] "Emotion data" refers to data obtained by analyzing information regarding a user's emotions.
[0664] An "emotion engine" is a technology that automatically identifies and analyzes a user's emotions from their voice, facial expressions, etc., and makes the results available within the system.
[0665] "Server" refers to a computer system on a network that provides and manages information and data.
[0666] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, an embodiment of the present invention will be described.
[0667] This invention uses a smartphone, smart glasses, or other smart devices as a terminal. The terminal has a built-in camera and microphone, and an interface that accepts voice instructions and hand gestures. This allows users to easily capture and save specific scenes and experiences.
[0668] The device uses voice recognition technology to analyze the user's voice instructions. The "SpeechRecognition" library is used for voice recognition. For example, when the user says "save it," the device will start the camera and take a picture of the scene. It also uses the "OpenCV" library to recognize hand gestures, so it is possible to give a shooting command using specific gestures.
[0669] The captured images are sent to the cloud using the device's communication function. The transmission to the cloud is done using the "Requests" library, and the captured images are stored on the cloud server.
[0670] Images stored on the cloud are retrieved and displayed in response to a user's question. A question is accepted by the device and sent to the cloud server. The cloud server analyzes the question, retrieves the associated image, and sends it to the device.
[0671] In addition, the system analyzes the voice data emitted by the user when taking a photo and uses an emotion engine to recognize the user's emotions. The emotion engine receives the voice data as input and analyzes the emotions. The emotion analysis results are associated with the image and stored in the cloud.
[0672] Image sorting and retrieval based on emotions is performed by the emotion engine on the cloud server. For example, it is possible to search for only images of scenes that the user felt were "fun."
[0673] Examples:
[0674] When a user is shopping in a brick-and-mortar store, they can say "Save this" when they see a product they like. The device that receives this command activates the camera and takes a picture of the scene. This picture is immediately sent to the cloud and saved. If the user then asks, "Show me some fun products," the cloud server uses the emotion engine to search for images associated with "fun" and sends them to the device, allowing the user to check the product again.
[0675] Example prompt:
[0676] "I want to create an application that allows users to easily save their favorite products or specific displays they find in physical stores. The user can take a picture by voice command or gesture, and the captured image will be sent to the cloud. Also, the application will analyze the user's emotions based on the voice data, and later create a product list based on the emotions. Please generate a program for this application."
[0677] The flow of the specific process in the application example 2 will be described with reference to FIG.
[0678] Step 1:
[0679] The device accepts voice commands or hand gestures. The user issues a voice command such as "save" or makes a specific hand gesture. The device receives voice or gestures as input using a microphone or camera, and analyzes them using voice recognition technology (SpeechRecognition library) or gesture recognition technology (OpenCV library). As a result of the analysis, it determines whether there is an instruction to take a photo.
[0680] Input: User's voice commands or hand gestures
[0681] Output: Whether or not shooting is instructed
[0682] Step 2:
[0683] The device starts up the camera and takes a picture of the specified scene. The captured image data is generated and temporarily stored in the device's memory.
[0684] Input:Shooting instructions
[0685] Output: Captured image data
[0686] Step 3:
[0687] The device sends the captured image data to the cloud. The captured image data is sent to the cloud server using the communication function (Requests library). The cloud server stores the received image data in a database.
[0688] Input: Photographed image data
[0689] Output: Image data stored in the cloud
[0690] Step 4:
[0691] The user asks a question about an image stored in the cloud. The user enters the question into the smart device, which then sends the question to the cloud server.
[0692] Input: User question
[0693] Output: A query request to the cloud server
[0694] Step 5:
[0695] The cloud server acquires information related to the stored images, searches the database for the corresponding image data based on the user's question, and transmits the searched image data to the terminal.
[0696] Input: A query request to the cloud server
[0697] Output: Retrieved image data
[0698] Step 6:
[0699] The terminal displays the searched image data. The terminal that receives the image data sent from the cloud server displays it on the screen and provides it to the user.
[0700] Input: Retrieved image data
[0701] Output: Display the image to the user
[0702] Step 7:
[0703] The cloud server uses an emotion engine to analyze the user's voice data and recognize the user's emotions. The emotion engine receives the voice data as input and executes an emotion analysis process. As a result of the analysis, it generates user emotion data and stores it in association with the image data.
[0704] Input: User's voice data
[0705] Output: Emotion data and associated image data
[0706] Step 8:
[0707] The cloud server organizes and searches image data based on the user's emotion data. For example, if a user asks, "Show me some fun products," the cloud server uses the emotion engine to search for images associated with "fun" and sends them to the device.
[0708] Input: emotion data and user questions
[0709] Output: Image data of sentiment-based search results
[0710] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0711] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0712] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0713] [Third embodiment]
[0714] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0715] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0716] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[0717] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0718] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[0719] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).
[0720] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[0721] Fig. 6 shows an example of main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0722] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0723] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0724] In the headset type terminal 314, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0725] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server", and the headset type terminal 314 will be referred to as the "terminal".
[0726] The embodiment for carrying out the present invention comprises the following elements.
[0727] 1. Terminal: A device with an interface that accepts voice commands and hand gestures when a user wants to save a particular scene or experience.
[0728] 2. Shooting: The device analyzes voice commands and hand gestures to understand the command to shoot, activates the camera, and shoots the specified scene or experience.
[0729] 3. Transmission means: The terminal has a communication function for transmitting captured images to the cloud. Images can be transmitted to a server via a network.
[0730] 4. Server: Stores the received images in a database, accepts queries about the contents stored by the user, and retrieves and displays related captured images.
[0731] (Specific examples)
[0732] When a user uses the headset terminal 314 to implement the present invention, the headset terminal 314 becomes a terminal. The user commands the headset terminal 314 to take a picture by giving a voice command such as "save it" or by making a specific gesture. The headset terminal 314 analyzes the voice command or gesture and understands the command to take a picture. The camera starts up and takes a picture of the instructed scene or experience. The captured image is sent to the cloud using the communication function of the headset terminal 314. The server stores and saves the received image in a database. When the user asks the headset terminal 314 a question about the saved content, the headset terminal 314 sends the question to the server. The server analyzes the received question, retrieves related captured images from the database, and sends them to the headset terminal 314. The headset terminal 314 displays the received captured images to provide the user with support for reminiscence.
[0733] The above is an example of an embodiment of the present invention. Although specific devices and processing procedures may vary depending on the selection and implementation of the implementer, the present invention can be implemented based on this embodiment.
[0734] The process flow will be explained below.
[0735] Step 1: The user speaks into the terminal, saying "Save."
[0736] Step 2: The device receives the voice command and uses voice recognition technology to interpret the command.
[0737] Step 3: Based on the analysis results, the device understands the shooting instructions.
[0738] Step 4: The device activates its camera and captures the specified scene or experience.
[0739] Step 5: The device uses its communication functions to send the captured images to the cloud.
[0740] Step 6: The server stores the received images in a database.
[0741] Step 7: The user asks a question to the terminal.
[0742] Step 8: The terminal sends the query to the server.
[0743] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[0744] Step 10: The server transmits the captured image to the terminal.
[0745] Step 11: The terminal displays the captured image that it has received.
[0746] Examples:
[0747] Step 1: The user issues a voice command to the headset type terminal 314 saying "save."
[0748] Step 2: The headset terminal 314 receives the voice instruction and uses voice recognition technology to interpret the instruction.
[0749] Step 3: The headset type terminal 314 understands the shooting instructions based on the analysis results.
[0750] Step 4: The headset type terminal 314 starts a camera application and captures the instructed scene or experience.
[0751] Step 5: The headset type terminal 314 uses the communication function to transmit the captured image to the cloud.
[0752] Step 6: The server stores the received images in a database.
[0753] Step 7: The user asks the headset type terminal 314, "Show me pictures of playing in the park yesterday."
[0754] Step 8: The headset terminal 314 sends a query to the server.
[0755] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[0756] Step 10: The server transmits the captured image to the headset type terminal 314.
[0757] Step 11: The headset type terminal 314 displays the received captured image.
[0758] The above is the specific process flow of this system. The user gives voice instructions or gestures, and the device analyzes them and performs appropriate processing, enabling the system to take photos, save images, and answer questions.
[0759] Example 1
[0760] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".
[0761] There is a demand for users to be able to easily save specific scenes and experiences and easily search and display the saved information at a later date. However, with conventional technologies, it is difficult to intuitively give shooting instructions using voice commands or hand gestures, or to efficiently retrieve images saved in the cloud based on a question.
[0762] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0763] In this invention, the server includes a means for receiving a shooting instruction by voice instruction or hand gesture, a means for analyzing the received shooting instruction to start the camera and shoot the scene, a means for transmitting the shot image to the cloud server via the network, a means for storing the received image in a database, and a means for receiving a question about the saved image from the user and acquiring and displaying related images, thereby enabling the user to intuitively store a particular scene or experience and efficiently recall it later.
[0764] A "voice instruction" refers to a specific operation or instruction given by a user to a terminal using voice.
[0765] A "hand gesture" is a specific operation or instruction given by a user to a terminal using hand movements.
[0766] A "photography instruction" is an instruction given by a user to a terminal to capture a specific scene or experience with a camera.
[0767] "Analysis" is the process of understanding the voice instructions and hand gestures received by the device and converting them into appropriate actions.
[0768] "Starting the camera" means starting up the camera function of the device and making it ready to take pictures.
[0769] A "cloud server" is a server used to store and process data over the Internet.
[0770] A database is a structured storage system that organizes and stores large amounts of data and allows for efficient search and retrieval.
[0771] "Network communication technology" refers to technology used for data communication, including Wi-Fi and mobile data.
[0772] A "question" is a query made by a user to obtain information about a previously stored image.
[0773] "Retrieval" is the process of searching and retrieving specific information from a database.
[0774] "Display" means to output the acquired information on the terminal screen so that the user can confirm it.
[0775] The embodiment of the present invention provides a system that allows a user to easily save a particular scene or experience, and easily search and display it later. A specific implementation method of this system will be described below.
[0776] A user uses a dedicated terminal to give voice instructions and hand gestures. This terminal is equipped with voice recognition technology and gesture recognition technology, and for example, voice recognition software provided by each company is used. For gesture recognition, a motion sensor combined with the terminal's camera is used.
[0777] The device analyzes the user's voice instructions and gestures and activates the camera based on those. The camera is capable of capturing high-resolution images. The captured images are sent to a cloud server using the device's network communication function (Wi-Fi or mobile data communication). In this case, HTTP / HTTPS is generally used as the communication protocol.
[0778] The cloud server processes the received image data appropriately, adds metadata (such as the date and time of shooting and the location), and stores the data in a database. The database uses storage provided by each company.
[0779] The user can later ask questions about the stored images through the device, for example by issuing a voice command such as "Show me the scenery from last week." The question is converted into text using the device's voice recognition technology and sent to the server.
[0780] The server analyzes the user's question, searches for and retrieves relevant images from the database, and then transmits the retrieved images to the terminal, which displays the received images on its screen and provides them to the user.
[0781] Examples:
[0782] When a user visits a park, they give a voice command to their smart device saying, "Save this view," or perform a specific hand gesture. The smart device uses its built-in voice recognition software to analyze the voice, activates the camera, and captures the view. The captured image is sent to a cloud server using its Wi-Fi function. The server adds metadata to the received image and stores it in a database.
[0783] A few days later, the user asks the smart device, "Show me the scenery from last week." The smart device converts this voice to text and sends it to the server. The server searches the database for the corresponding image and sends it to the smart device. The device displays the image on the screen, allowing the user to reminisce about the scenery from last week.
[0784] Example prompt:
[0785] Design a system that allows a user to take pictures of scenery in a park and store them for later review. Include instructions for using voice commands or hand gestures to take pictures, store the data in the cloud, and play the stored images at a later date.
[0786] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0787] Step 1:
[0788] The user gives a voice command or a gesture.
[0789] Specific action: The user speaks to the smart device, saying "Save this view," or makes a specific hand gesture.
[0790] Input: Voice information or gesture data.
[0791] Output: Recognized by the terminal as audio or video data.
[0792] Step 2:
[0793] The device analyzes the voice command or gesture.
[0794] Specific operation: The voice recognition software in the terminal analyzes the voice, and the gesture recognition software analyzes the video data.
[0795] Input: The output data of step 1 (audio or video data).
[0796] Output: Parsed instruction (e.g. "Save the view").
[0797] Step 3:
[0798] The device will activate the camera and take a picture of the specified scene.
[0799] Specific operation: Based on the results of voice or gesture analysis, the device will activate the built-in camera to capture the scenery.
[0800] Input: The output data of step 2 (the parsed instructions).
[0801] Output: High resolution image data.
[0802] Step 4:
[0803] The images captured by the device are sent to a cloud server.
[0804] Specific operation: The device uses Wi-Fi or mobile data communication to send the captured image data to a cloud server.
[0805] Input: The output data of step 3 (image data).
[0806] Output: Image data is sent to the server.
[0807] Step 5:
[0808] The server stores the images in a database.
[0809] Specific operation: The server adds metadata (photo date and time and location) to the received image data and stores it in a database.
[0810] Input: The output data of step 4 (image data).
[0811] Output: Data stored in the database.
[0812] Step 6:
[0813] The user queries for a stored image.
[0814] Specific operation: The user speaks to the smart device and asks, "Show me the scenery from last week."
[0815] Input: Audio information.
[0816] Output: Recognized by the device as audio data.
[0817] Step 7:
[0818] The server retrieves relevant images from a database and transmits them to the terminal.
[0819] Specific operation: The server analyzes the question through voice recognition software, searches for and retrieves the corresponding image from the database, and sends the retrieved image to the terminal.
[0820] Input: The output data from step 6 (audio data) and the information in the database.
[0821] Output: The retrieved image data is sent to the terminal.
[0822] Step 8:
[0823] The device displays the image it receives.
[0824] Specific operation: The terminal displays the image data received from the server on the screen.
[0825] Input: The output data of step 7 (image data).
[0826] Output: The image that is displayed to the user.
[0827] (Application example 1)
[0828] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0829] In autonomous vehicles, when passengers want to record a particular experience or scenery, there is a need for a system that allows them to easily search and view the recorded images and videos later while intuitively and conveniently operating the system. In addition, there is an increasing need for passengers to record more experiences as they do not need to concentrate on driving. However, existing technologies to achieve this are limited, and there is no systematic solution.
[0830] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0831] In this invention, the server includes a means for receiving a shooting instruction by voice instruction or hand gesture, a means for transmitting the captured image to the cloud, and a means for retrieving the image stored in the cloud based on an associated question and displaying it on a display device in the vehicle. This allows passengers to easily record a particular experience by voice instruction or gesture only, and store it on the cloud for easy searching and viewing later.
[0832] The term "instruction to shoot" refers to an operation in which the user signals the device to start shooting by voice instruction or hand gesture.
[0833] "Voice instruction" refers to the act of a user speaking into a device to instruct the device to perform a specific action.
[0834] "Hand gestures" refer to instructions that the device recognizes hand movements and actions and uses them to perform specific actions.
[0835] "Cloud" refers to an online storage system for storing data on remote servers via the Internet.
[0836] "Means for sending to the cloud" refers to a method for transmitting images and data captured from a device to a cloud server via a network.
[0837] "Images stored in the cloud" refers to photos and videos that are sent via a network to a remote server and stored in online storage.
[0838] "Related questions" refer to keywords or phrases that users use when searching for specific images or experiences.
[0839] "Means of acquiring and displaying on a display device inside the vehicle" refers to a method of searching for images stored in the cloud and displaying them on a display or the like inside an autonomous vehicle.
[0840] "Display" refers to an electronic device for visually displaying information.
[0841] The main components of the system for realizing this application example are a camera, a microphone, a display, a cloud server, and a terminal with a communication function installed in the vehicle. The detailed configuration and operation of this system are described below.
[0842] 1. Hardware Configuration
[0843] Cameras in vehicles
[0844] The cameras will be installed at specific locations within the train and will be tasked with recording scenes and experiences that passengers indicate they want to capture. The cameras are capable of taking high-resolution images and have the ability to automatically adjust exposure and focus according to the environment.
[0845] microphone
[0846] The microphone is designed to clearly capture the user's voice commands while reducing noise inside the vehicle, allowing the voice recognition technology to accurately interpret them.
[0847] display
[0848] The display is installed on the dashboard or back seat monitor inside the vehicle and is used to visually display stored images and videos.
[0849] 2. Software Configuration
[0850] Voice Recognition Technology
[0851] For voice recognition, a library called "SpeechRecognition" is used. Voice data is acquired from the microphone, analyzed, and instructions are determined. For example, if the user says "Take a picture," the voice is analyzed and the camera is activated.
[0852] Image capture and storage
[0853] Images captured by the camera are processed using the "OpenCV" library and temporarily stored in the in-vehicle system before being sent to a cloud server via the network, using the "Requests" library to send the data.
[0854] Cloud Server
[0855] The cloud server receives the captured images and stores them in a database, and when a user searches for a specific image, it also retrieves the relevant images from the database based on the relevant query.
[0856] 3. Processing Procedure
[0857] Shooting instructions
[0858] When a user issues a voice command to the camera to "take a picture," the microphone captures the voice and the voice is analyzed using voice recognition technology. The camera then starts up and takes a picture of the scene as instructed.
[0859] Cloud Upload
[0860] The captured images are temporarily stored in the vehicle and then transmitted to a cloud server using the network communication function. The cloud server stores the received images in a database.
[0861] Searching and viewing images
[0862] When the user issues a voice command, such as "Show me today's sunset," the command is again analyzed through voice recognition technology. The cloud server retrieves relevant images from the database and displays them on the vehicle's display.
[0863] Examples
[0864] For example, if a user says "Take a picture" while in the car, the camera will automatically capture the scenery and store the image in the cloud. Later, if the user says "Show me today's sunset," the relevant image will be retrieved from the cloud and displayed on the display in front of the passenger.
[0865] Examples of prompt statements
[0866] Imagine an app that takes a picture of a sunset with an in-car camera, and when the driver says "Show me today's sunset," the app displays the saved image on the in-car display. Please submit a detailed program outline.
[0867] The flow of the specific process in the application example 1 will be described with reference to FIG.
[0868] Step 1:
[0869] The user issues a voice command
[0870] The user says "take a picture" into the microphone, which causes voice data (input) to be captured by the microphone.
[0871] Step 2:
[0872] The device analyzes the voice command
[0873] The terminal analyzes the captured voice data using voice recognition technology (e.g., the SpeechRecognition library) and obtains (outputs) the voice command "Take a picture" as text data.
[0874] Step 3:
[0875] The camera starts and takes a picture
[0876] The device activates the camera based on the results of voice recognition. The camera captures the specified scene and obtains image data (input). The obtained image data is temporarily stored inside the device (output).
[0877] Step 4:
[0878] The device sends image data to the cloud.
[0879] The device transmits the temporarily stored image data to the cloud server using a network communication technology (e.g., the Requests library) (input). The cloud server stores the received image data in a database (output).
[0880] Step 5:
[0881] The user issues a search command
[0882] The user says into the microphone, "Show me today's sunset." This causes the microphone to capture voice data (input) of the search command.
[0883] Step 6:
[0884] The device parses the search instructions
[0885] The device analyzes the captured voice data using voice recognition technology and obtains the voice instruction "Show me today's sunset" as text data (output).
[0886] Step 7:
[0887] The cloud server searches for related images
[0888] The device sends the analyzed text data (input) to the cloud server, which searches the database for related image data that matches the text data and retrieves them (output).
[0889] Step 8:
[0890] Display the image captured by the device
[0891] Relevant image data obtained from the cloud server (input) is sent to the terminal, which then displays the received image data on a display inside the vehicle (output).
[0892] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0893] The embodiment for carrying out the present invention comprises the following elements.
[0894] 1. Terminal: A device with an interface that accepts voice commands and hand gestures when a user wants to save a particular scene or experience.
[0895] 2. Shooting: The device analyzes voice commands and hand gestures to understand the command to shoot, activates the camera, and shoots the specified scene or experience.
[0896] 3. Transmission means: The terminal has a communication function for transmitting captured images to the cloud. Images can be transmitted to a server via a network.
[0897] 4. Server: Stores the received images in a database, accepts queries about the contents stored by the user, and retrieves and displays related captured images.
[0898] 5. Emotion engine: The server combines an emotion engine and utilizes technology to recognize the user’s emotions, such as voice and facial expressions. By analyzing the user’s emotions at the time of shooting and associating them with the captured image, the server can automatically classify and search for images based on emotions.
[0899] (Specific examples)
[0900] When a user uses the headset terminal 314 to implement the present invention, the headset terminal 314 becomes a terminal. The user commands the headset terminal 314 to take a picture by giving a voice command such as "save it" or by making a specific gesture. The headset terminal 314 analyzes the voice command or gesture and understands the command to take a picture. The camera starts up and takes a picture of the instructed scene or experience. The captured image is sent to the cloud using the communication function of the headset terminal 314. The server stores and saves the received image in a database. When the user asks the headset terminal 314 a question about the saved content, the headset terminal 314 sends the question to the server. The server analyzes the received question, retrieves related captured images from the database, and sends them to the headset terminal 314. The headset terminal 314 displays the received captured image. The server also combines an emotion engine and utilizes technology to recognize the user's emotions, such as voice and facial expressions. By analyzing the user's emotions at the time of shooting and associating them with the captured image, automatic classification and search of images based on emotions are performed.
[0901] The above is an example of an embodiment of the present invention. Although specific devices and processing procedures may vary depending on the selection and implementation of the implementer, the present invention can be implemented based on this embodiment.
[0902] The process flow will be explained below.
[0903] Step 1: The user speaks into the terminal, saying "Save."
[0904] Step 2: The device receives the voice command and uses voice recognition technology to interpret the command.
[0905] Step 3: Based on the analysis results, the device understands the shooting instructions.
[0906] Step 4: The device activates its camera and captures the specified scene or experience.
[0907] Step 5: The device uses its communication functions to send the captured images to the cloud.
[0908] Step 6: The server stores the received images in a database.
[0909] Step 7: The user asks a question to the terminal.
[0910] Step 8: The terminal sends the query to the server.
[0911] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[0912] Step 10: The server transmits the captured image to the terminal.
[0913] Step 11: The terminal displays the captured image that it has received.
[0914] Step 12: The server uses the emotion engine to recognize the user's emotions, such as voice and facial expressions.
[0915] Step 13: The server analyzes the user's emotion at the time of taking the photograph and associates it with the photographed image.
[0916] Step 14: The server performs automatic classification and retrieval of images based on emotions.
[0917] Examples:
[0918] Step 1: The user issues a voice command to the headset type terminal 314 saying "save."
[0919] Step 2: The headset terminal 314 receives the voice instruction and uses voice recognition technology to interpret the instruction.
[0920] Step 3: The headset type terminal 314 understands the shooting instructions based on the analysis results.
[0921] Step 4: The headset type terminal 314 starts a camera application and captures the instructed scene or experience.
[0922] Step 5: The headset type terminal 314 uses the communication function to transmit the captured image to the cloud.
[0923] Step 6: The server stores the received images in a database.
[0924] Step 7: The user asks the headset type terminal 314, "Show me pictures of playing in the park yesterday."
[0925] Step 8: The headset terminal 314 sends the query to the server.
[0926] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[0927] Step 10: The server transmits the captured image to the headset type terminal 314.
[0928] Step 11: The headset type terminal 314 displays the received captured image.
[0929] Step 12: The server uses the emotion engine to recognize the user's emotions, such as voice and facial expressions.
[0930] Step 13: The server analyzes the user's emotion at the time of taking the photograph and associates it with the photographed image.
[0931] Step 14: The server performs automatic classification and retrieval of images based on emotions.
[0932] The above is the specific process flow of this system. The user gives voice instructions or gestures, and the device analyzes them and performs appropriate processing, enabling the system to take photos, save images, answer questions, and classify and search images based on emotions.
[0933] Example 2
[0934] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".
[0935] Conventional image storage systems make it difficult for users to instantly record specific moments and easily search and display them when needed. Conventional systems also lack the ability to classify and search for images based on the user's emotions, and therefore lack the functionality to allow users to replay memories based on their emotions.
[0936] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for a user to give a shooting instruction by voice instruction or hand gesture, a means for a terminal to analyze the voice instruction or gesture, start a camera and shoot a scene, a means for a terminal to transmit the captured image to a cloud server, a means for a server to store the image in a database, a means for a server to accept a related question, obtain related images from the database and transmit them to the terminal, and a means for a server to analyze a user's emotion using an emotion engine and associate the emotion with the image. This allows a user to easily record a particular moment and later search and display images based on the emotion.
[0937] "User" refers to a person who uses the System to store and view specific scenes or experiences.
[0938] The term "terminal" refers to a device that allows a user to give shooting instructions by voice or gestures.
[0939] "Voice instruction" refers to an instruction given by voice to a terminal by a user.
[0940] A "hand gesture" refers to a user giving instructions to a terminal by performing a specific action.
[0941] "Camera" refers to a device built into a terminal that takes photos and videos at the user's command.
[0942] A "cloud server" refers to a remote server accessible via a network, which provides a place to store image data.
[0943] A "database" refers to a system built within a server that systematically stores captured images and related information.
[0944] A "related question" refers to an inquiry that a user makes to the system regarding images that he or she has previously saved and their contents.
[0945] An "emotion engine" is a technology that analyzes voice and facial expressions to recognize the user's emotions, and then classifies and searches for images based on the results.
[0946] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0947] This invention is a system that allows users to easily save specific scenes or experiences, and later search and display those images based on emotions. The system consists of a terminal operated by the user and a server installed on the cloud. The main hardware and software configurations are explained below.
[0948] Terminal
[0949] The terminal is a device that allows the user to give instructions to shoot using voice commands or hand gestures. Specifically, this applies to smartphones, tablets, smart glasses, etc. The terminal has the following main functions:
[0950] 1. Voice Recognition Function:
[0951] The terminal analyzes the user's voice instructions using well-known natural language processing (NLP) techniques.
[0952] 2. Gesture recognition:
[0953] The device is equipped with technology that uses cameras and sensors to recognize the user's hand movements and analyze gestures.
[0954] 3. Camera Function:
[0955] It has a built-in high-resolution camera, and after understanding the shooting instructions, it will capture the specified scene.
[0956] 4. Communication function:
[0957] The captured images are sent to a cloud server using networks such as Wi-Fi, 4G / 5G.
[0958] server
[0959] The server is installed in the cloud and has the following main functions:
[0960] 1. Database Management:
[0961] Image data is efficiently stored and saved in a database. The database can use Amazon Web Services (AWS) RDS.
[0962] 2. Question analysis function:
[0963] The questions entered by the user are analyzed using natural language processing technology and related images are linked.
[0964] 3. Emotion Engine:
[0965] It uses voice and facial expression analysis technology to recognize user emotions, and uses that data to tag images with emotions for automatic classification and emotion-based search.
[0966] 4. Image sending function:
[0967] The image data retrieved in response to the user's question is sent to the terminal using encryption protocols such as SSL / TLS.
[0968] Examples
[0969] For example, suppose a user says to the device "Save it" during a family trip. The device uses voice recognition technology to analyze this voice command, activates the camera, and captures beautiful scenery. The captured image is then sent to a cloud server using the device's communication function. The sent image is stored in a database on the server. Later, when the user issues a voice command such as "Show me photos of the beach from last summer," the server analyzes the question, searches the database for the corresponding image, and sends it to the user's device. The device displays the received image on the screen. The server's emotion engine also recognizes emotions such as "joy" from the user's voice and facial expression when taking the photo, and associates and saves the emotion tag with the image, so that if the user commands "Show me photos of fun times," an image associated with the joyful emotion will be displayed.
[0970] Example prompts for generative AI models
[0971] "Please explain how the system works, allowing users to take pictures of specific scenes using voice commands and store them in the cloud."
[0972] "How can I use an emotion engine to automatically classify images based on user emotions?"
[0973] "Please explain with a concrete example the process of taking pictures using a smartphone and saving them to the cloud."
[0974] The above is a detailed embodiment of the present invention. This system allows users to easily operate it and store and search images based on emotions, greatly improving the user experience.
[0975] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0976] Step 1:
[0977] The user issues a command to take a picture. Specifically, the user issues a voice command to the smartphone saying "save it" or performs a specific gesture. The input data is a voice signal or a gesture movement, which is input to the terminal. The output is a flag indicating that the command to take a picture has been recognized.
[0978] Step 2:
[0979] The device analyzes the voice command and gestures. The device analyzes the voice command using well-known natural language processing techniques. It also recognizes certain actions using gesture recognition techniques. The input data is the user's voice or gestures, and the output is the analyzed command. Specifically, the voice command "save" is recognized, and an instruction to start the camera is issued.
[0980] Step 3:
[0981] The device activates the camera and captures the scene. The camera understands the specified instructions and takes the picture. The input data is the parsed instructions, and the output is the captured image data. For example, a beautiful landscape during a family trip is captured as a high-resolution photo.
[0982] Step 4:
[0983] The device sends the captured image to the cloud server. The device uses Wi-Fi or 4G / 5G networks to send encrypted image data to the cloud server. The input data is the captured image data, and the output is the image data sent to the cloud server.
[0984] Step 5:
[0985] The server stores the images in a database. The server organizes the image data it receives and stores it in a database. The input data is image data sent via the cloud, and the output is image entries stored in the database. For example, AWS's RDS is used to efficiently store and manage image data.
[0986] Step 6:
[0987] The user asks a question related to the saved content. The user issues a voice command to the smartphone, such as "Show me a photo of the beach from last summer." The input data is the voice command to ask the question, and the output is the content of the question to the server.
[0988] Step 7:
[0989] The server analyzes the question and retrieves related images. Using natural language processing technology, the server analyzes the user's question and searches the database for relevant image data. The input data is the question based on the voice command, and the output is related image data. For example, "photos of the sea from last summer" is searched for.
[0990] Step 8:
[0991] The server sends the relevant images to the user's device. The server encrypts the image data it finds and sends it to the user's device. The input data is the image data retrieved from the database, and the output is the image data sent to the device.
[0992] Step 9:
[0993] The user's device displays the image. The device displays the received image to the user. The input data is the image data received from the server, and the output is the image displayed on the device's display. For example, a photo of a summer beach is displayed on the screen of a smartphone.
[0994] Step 10:
[0995] The server performs emotion analysis using an emotion engine and associates it with the image. The server analyzes the voice and facial expression data to recognize the user's emotion. The image data with the emotion tag is saved for future searches. The input data is the voice and facial expression data at the time of shooting, and the output is image data with the emotion tag. For example, if the user is feeling "joy," that emotion tag is associated with the image.
[0996] (Application example 2)
[0997] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server", and the headset type terminal 314 will be referred to as a "terminal".
[0998] In modern commerce, when consumers visit a store and select products, they need a way to easily save specific scenes and favorite products for later reference. Furthermore, there is no system that can classify and search saved information based on consumers' emotions. This leaves users lacking a way to efficiently manage their shopping experience and receive emotion-based feedback.
[0999] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1000] In this invention, the server includes means for accepting shooting instructions by voice instructions or hand gestures, means for transmitting the captured images to the cloud, means for retrieving and displaying the images stored in the cloud based on associated questions, means for analyzing the user's emotions, and means for organizing and searching for images based on the emotion data.
[1001] This allows users to easily save specific scenes or products and later categorize and search for images based on their emotions, resulting in a more fulfilling shopping experience.
[1002] A "capture command" refers to a voice command or hand gesture made by a user to record a particular scene or experience.
[1003] "Voice instruction" refers to a method in which a user instructs a system to perform a specific operation using voice.
[1004] "Hand gestures" refers to the way in which a user uses hand movements to instruct a system to perform certain actions.
[1005] The term "captured image" refers to a still image captured by a camera based on a user's instruction to capture a picture.
[1006] "Cloud" refers to a technological infrastructure that stores and manages data on remote servers provided via the Internet.
[1007] "Communication means" refers to the technical means and devices for transmitting and receiving data.
[1008] "User's emotions" refers to psychological reactions and sensations analyzed from the user's voice and facial expressions.
[1009] "Emotion data" refers to data obtained by analyzing information regarding a user's emotions.
[1010] An "emotion engine" is a technology that automatically identifies and analyzes a user's emotions from their voice, facial expressions, etc., and makes the results available within the system.
[1011] "Server" refers to a computer system on a network that provides and manages information and data.
[1012] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, an embodiment of the present invention will be described.
[1013] This invention uses a smartphone, smart glasses, or other smart devices as a terminal. The terminal has a built-in camera and microphone, and an interface that accepts voice instructions and hand gestures. This allows users to easily capture and save specific scenes and experiences.
[1014] The device uses voice recognition technology to analyze the user's voice instructions. The "SpeechRecognition" library is used for voice recognition. For example, when the user says "save it," the device will start the camera and take a picture of the scene. It also uses the "OpenCV" library to recognize hand gestures, so it is possible to give a shooting command using specific gestures.
[1015] The captured images are sent to the cloud using the device's communication function. The transmission to the cloud is done using the "Requests" library, and the captured images are stored on the cloud server.
[1016] Images stored on the cloud are retrieved and displayed in response to a user's question. A question is accepted by the device and sent to the cloud server. The cloud server analyzes the question, retrieves the associated image, and sends it to the device.
[1017] In addition, the system analyzes the voice data emitted by the user when taking a photo and uses an emotion engine to recognize the user's emotions. The emotion engine receives the voice data as input and analyzes the emotions. The emotion analysis results are associated with the image and stored in the cloud.
[1018] Image sorting and retrieval based on emotions is performed by the emotion engine on the cloud server. For example, it is possible to search for only images of scenes that the user felt were "fun."
[1019] Examples:
[1020] When a user is shopping in a brick-and-mortar store, they can say "Save this" when they see a product they like. The device that receives this command activates the camera and takes a picture of the scene. This picture is immediately sent to the cloud and saved. If the user then asks, "Show me some fun products," the cloud server uses the emotion engine to search for images associated with "fun" and sends them to the device, allowing the user to check the product again.
[1021] Example prompt:
[1022] "I want to create an application that allows users to easily save their favorite products or specific displays they find in physical stores. The user can take a picture by voice command or gesture, and the captured image will be sent to the cloud. Also, the application will analyze the user's emotions based on the voice data, and later create a product list based on the emotions. Please generate a program for this application."
[1023] The flow of the specific process in the application example 2 will be described with reference to FIG.
[1024] Step 1:
[1025] The device accepts voice commands or hand gestures. The user issues a voice command such as "save" or makes a specific hand gesture. The device receives voice or gestures as input using a microphone or camera, and analyzes them using voice recognition technology (SpeechRecognition library) or gesture recognition technology (OpenCV library). As a result of the analysis, it determines whether there is an instruction to take a photo.
[1026] Input: User's voice commands or hand gestures
[1027] Output: Whether or not shooting is instructed
[1028] Step 2:
[1029] The device starts up the camera and takes a picture of the specified scene. The captured image data is generated and temporarily stored in the device's memory.
[1030] Input:Shooting instructions
[1031] Output: Captured image data
[1032] Step 3:
[1033] The device sends the captured image data to the cloud. The captured image data is sent to the cloud server using the communication function (Requests library). The cloud server stores the received image data in a database.
[1034] Input: Photographed image data
[1035] Output: Image data stored in the cloud
[1036] Step 4:
[1037] The user asks a question about an image stored in the cloud. The user enters the question into the smart device, which then sends the question to the cloud server.
[1038] Input: User question
[1039] Output: A query request to the cloud server
[1040] Step 5:
[1041] The cloud server acquires information related to the stored images, searches the database for the corresponding image data based on the user's question, and transmits the searched image data to the terminal.
[1042] Input: A query request to the cloud server
[1043] Output: Retrieved image data
[1044] Step 6:
[1045] The terminal displays the searched image data. The terminal that receives the image data sent from the cloud server displays it on the screen and provides it to the user.
[1046] Input: Retrieved image data
[1047] Output: Display the image to the user
[1048] Step 7:
[1049] The cloud server uses an emotion engine to analyze the user's voice data and recognize the user's emotions. The emotion engine receives the voice data as input and executes an emotion analysis process. As a result of the analysis, it generates user emotion data and stores it in association with the image data.
[1050] Input: User's voice data
[1051] Output: Emotion data and associated image data
[1052] Step 8:
[1053] The cloud server organizes and searches image data based on the user's emotion data. For example, if a user asks, "Show me some fun products," the cloud server uses the emotion engine to search for images associated with "fun" and sends them to the device.
[1054] Input: emotion data and user questions
[1055] Output: Image data of sentiment-based search results
[1056] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1057] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1058] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1059] [Fourth embodiment]
[1060] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1061] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1062] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[1063] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. In addition, the microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1064] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[1065] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).
[1066] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[1067] The control target 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, legs, etc. The posture and behavior of the robot 414 are controlled by controlling the motors of the arms, hands, legs, etc. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1068] Fig. 8 shows an example of main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1069] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1070] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1071] In the robot 414, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1072] Next, a description will be given of the specific processing by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".
[1073] The embodiment for carrying out the present invention comprises the following elements.
[1074] 1. Terminal: A device with an interface that accepts voice commands and hand gestures when a user wants to save a particular scene or experience.
[1075] 2. Shooting: The device analyzes voice commands and hand gestures to understand the command to shoot, activates the camera, and shoots the specified scene or experience.
[1076] 3. Transmission means: The terminal has a communication function for transmitting captured images to the cloud. Images can be transmitted to a server via a network.
[1077] 4. Server: Stores the received images in a database, accepts queries about the contents stored by the user, and retrieves and displays related captured images.
[1078] (Specific examples)
[1079] When a user uses the robot 414 to implement the present invention, the robot 414 becomes a terminal. The user commands the robot 414 to take a picture by giving a voice command such as "save it" or by making a specific gesture. The robot 414 analyzes the voice command or gesture and understands the command to take a picture. The camera starts up and takes a picture of the instructed scene or experience. The captured image is sent to the cloud using the communication function of the robot 414. The server stores and saves the received image in a database. When the user asks the robot 414 a question about the saved content, the robot 414 sends the question to the server. The server analyzes the received question, retrieves the related captured image from the database, and sends it to the robot 414. The robot 414 displays the received captured image to provide the user with support for reminiscence.
[1080] The above is an example of an embodiment of the present invention. Although specific devices and processing procedures may vary depending on the selection and implementation of the implementer, the present invention can be implemented based on this embodiment.
[1081] The process flow will be explained below.
[1082] Step 1: The user speaks into the terminal, saying "Save."
[1083] Step 2: The device receives the voice command and uses voice recognition technology to interpret the command.
[1084] Step 3: Based on the analysis results, the device understands the shooting instructions.
[1085] Step 4: The device activates its camera and captures the specified scene or experience.
[1086] Step 5: The device uses its communication functions to send the captured images to the cloud.
[1087] Step 6: The server stores the received images in a database.
[1088] Step 7: The user asks a question to the terminal.
[1089] Step 8: The terminal sends the query to the server.
[1090] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[1091] Step 10: The server transmits the captured image to the terminal.
[1092] Step 11: The terminal displays the captured image that it has received.
[1093] Examples:
[1094] Step 1: The user commands the robot 414 to "save."
[1095] Step 2: The robot 414 receives the voice instruction and uses voice recognition technology to interpret the instruction.
[1096] Step 3: The robot 414 understands the shooting instructions based on the analysis results.
[1097] Step 4: The robot 414 launches a camera app and captures the instructed scene or experience.
[1098] Step 5: The robot 414 uses its communication capabilities to send the captured images to the cloud.
[1099] Step 6: The server stores the received images in a database.
[1100] Step 7: The user asks the robot 414, "Show me pictures of playing in the park yesterday."
[1101] Step 8: The robot 414 sends the question to the server.
[1102] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[1103] Step 10: The server transmits the captured image to the robot 414.
[1104] Step 11: The robot 414 displays the captured image that it has received.
[1105] The above is the specific process flow of this system. The user gives voice instructions or gestures, and the device analyzes them and performs appropriate processing, enabling the system to take photos, save images, and answer questions.
[1106] Example 1
[1107] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."
[1108] There is a demand for users to be able to easily save specific scenes and experiences and easily search and display the saved information at a later date. However, with conventional technologies, it is difficult to intuitively give shooting instructions using voice commands or hand gestures, or to efficiently retrieve images saved in the cloud based on a question.
[1109] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1110] In this invention, the server includes a means for receiving a shooting instruction by voice instruction or hand gesture, a means for analyzing the received shooting instruction to start the camera and shoot the scene, a means for transmitting the shot image to the cloud server via the network, a means for storing the received image in a database, and a means for receiving a question about the saved image from the user and acquiring and displaying related images, thereby enabling the user to intuitively store a particular scene or experience and efficiently recall it later.
[1111] A "voice instruction" refers to a specific operation or instruction given by a user to a terminal using voice.
[1112] A "hand gesture" is a specific operation or instruction given by a user to a terminal using hand movements.
[1113] A "photography instruction" is an instruction given by a user to a terminal to capture a specific scene or experience with a camera.
[1114] "Analysis" is the process of understanding the voice instructions and hand gestures received by the device and converting them into appropriate actions.
[1115] "Starting the camera" means starting up the camera function of the device and making it ready to take pictures.
[1116] A "cloud server" is a server used to store and process data over the Internet.
[1117] A database is a structured storage system that organizes and stores large amounts of data and allows for efficient search and retrieval.
[1118] "Network communication technology" refers to technology used for data communication, including Wi-Fi and mobile data.
[1119] A "question" is a query made by a user to obtain information about a previously stored image.
[1120] "Retrieval" is the process of searching and retrieving specific information from a database.
[1121] "Display" means to output the acquired information on the terminal screen so that the user can confirm it.
[1122] The embodiment of the present invention provides a system that allows a user to easily save a particular scene or experience, and easily search and display it later. A specific implementation method of this system will be described below.
[1123] A user uses a dedicated terminal to give voice instructions and hand gestures. This terminal is equipped with voice recognition technology and gesture recognition technology, and for example, voice recognition software provided by each company is used. For gesture recognition, a motion sensor combined with the terminal's camera is used.
[1124] The device analyzes the user's voice instructions and gestures and activates the camera based on those. The camera is capable of capturing high-resolution images. The captured images are sent to a cloud server using the device's network communication function (Wi-Fi or mobile data communication). In this case, HTTP / HTTPS is generally used as the communication protocol.
[1125] The cloud server appropriately processes the received image data, adds metadata (such as the date and time of shooting and the location), and stores the data in a database. The database uses storage provided by each company.
[1126] The user can later ask questions about the stored images through the device, for example by issuing a voice command such as "Show me the scenery from last week." The question is converted into text using the device's voice recognition technology and sent to the server.
[1127] The server analyzes the user's question, searches for and retrieves relevant images from the database, and then transmits the retrieved images to the terminal, which displays the received images on its screen and provides them to the user.
[1128] Examples:
[1129] When a user visits a park, they give a voice command to their smart device saying, "Save this view," or perform a specific hand gesture. The smart device uses its built-in voice recognition software to analyze the voice, activates the camera, and captures the view. The captured image is sent to a cloud server using its Wi-Fi function. The server adds metadata to the received image and stores it in a database.
[1130] A few days later, the user asks the smart device, "Show me the scenery from last week." The smart device converts this voice to text and sends it to the server. The server searches the database for the corresponding image and sends it to the smart device. The device displays the image on the screen, allowing the user to reminisce about the scenery from last week.
[1131] Example prompt:
[1132] Design a system that allows a user to take pictures of scenery in a park and store them for later review. Include instructions for using voice commands or hand gestures to take pictures, store the data in the cloud, and play the stored images at a later date.
[1133] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1134] Step 1:
[1135] The user gives a voice command or a gesture.
[1136] Specific action: The user speaks to the smart device, saying "Save this view," or makes a specific hand gesture.
[1137] Input: Voice information or gesture data.
[1138] Output: Recognized by the terminal as audio or video data.
[1139] Step 2:
[1140] The device analyzes the voice command or gesture.
[1141] Specific operation: The voice recognition software in the terminal analyzes the voice, and the gesture recognition software analyzes the video data.
[1142] Input: The output data of step 1 (audio or video data).
[1143] Output: Parsed instruction (e.g. "Save the view").
[1144] Step 3:
[1145] The device will activate the camera and take a picture of the specified scene.
[1146] Specific operation: Based on the results of voice or gesture analysis, the device will activate the built-in camera to capture the scenery.
[1147] Input: The output data of step 2 (the parsed instructions).
[1148] Output: High resolution image data.
[1149] Step 4:
[1150] The images captured by the device are sent to a cloud server.
[1151] Specific operation: The device uses Wi-Fi or mobile data communication to send the captured image data to a cloud server.
[1152] Input: The output data of step 3 (image data).
[1153] Output: Image data is sent to the server.
[1154] Step 5:
[1155] The server stores the images in a database.
[1156] Specific operation: The server adds metadata (photo date and time and location) to the received image data and stores it in a database.
[1157] Input: The output data of step 4 (image data).
[1158] Output: Data stored in the database.
[1159] Step 6:
[1160] The user queries for a stored image.
[1161] Specific operation: The user speaks to the smart device and asks, "Show me the scenery from last week."
[1162] Input: Audio information.
[1163] Output: Recognized by the device as audio data.
[1164] Step 7:
[1165] The server retrieves relevant images from a database and transmits them to the terminal.
[1166] Specific operation: The server analyzes the question through voice recognition software, searches for and retrieves the corresponding image from the database, and sends the retrieved image to the terminal.
[1167] Input: The output data from step 6 (audio data) and the information in the database.
[1168] Output: The retrieved image data is sent to the terminal.
[1169] Step 8:
[1170] The device displays the image it receives.
[1171] Specific operation: The terminal displays the image data received from the server on the screen.
[1172] Input: The output data of step 7 (image data).
[1173] Output: The image that is displayed to the user.
[1174] (Application example 1)
[1175] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1176] In autonomous vehicles, when passengers want to record a particular experience or scenery, there is a need for a system that allows them to easily search and view the recorded images and videos later while intuitively and conveniently operating the system. In addition, there is an increasing need for passengers to record more experiences as they do not need to concentrate on driving. However, existing technologies to achieve this are limited, and there is no systematic solution.
[1177] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1178] In this invention, the server includes a means for receiving a shooting instruction by voice instruction or hand gesture, a means for transmitting the captured image to the cloud, and a means for retrieving the image stored in the cloud based on an associated question and displaying it on a display device in the vehicle. This allows passengers to easily record a particular experience by voice instruction or gesture only, and store it on the cloud for easy searching and viewing later.
[1179] The term "instruction to shoot" refers to an operation in which the user signals the device to start shooting by voice instruction or hand gesture.
[1180] "Voice instruction" refers to the act of a user speaking into a device to instruct the device to perform a specific action.
[1181] "Hand gestures" refer to instructions that the device recognizes hand movements and actions and uses them to perform specific actions.
[1182] "Cloud" refers to an online storage system for storing data on remote servers via the Internet.
[1183] "Means for sending to the cloud" refers to a method for transmitting images and data captured from a device to a cloud server via a network.
[1184] "Images stored in the cloud" refers to photos and videos that are sent via a network to a remote server and stored in online storage.
[1185] "Related questions" refer to keywords or phrases that users use when searching for specific images or experiences.
[1186] "Means of acquiring and displaying on a display device inside the vehicle" refers to a method of searching for images stored in the cloud and displaying them on a display or the like inside an autonomous vehicle.
[1187] "Display" refers to an electronic device for visually displaying information.
[1188] The main components of the system for realizing this application example are a camera, a microphone, a display, a cloud server, and a terminal with a communication function installed in the vehicle. The detailed configuration and operation of this system are described below.
[1189] 1. Hardware Configuration
[1190] Cameras in vehicles
[1191] The cameras will be installed at specific locations within the train and will be tasked with recording scenes and experiences that passengers indicate they want to capture. The cameras are capable of taking high-resolution images and have the ability to automatically adjust exposure and focus according to the environment.
[1192] microphone
[1193] The microphone is designed to clearly capture the user's voice commands while reducing noise inside the vehicle, allowing the voice recognition technology to accurately interpret them.
[1194] display
[1195] The display is installed on the dashboard or back seat monitor inside the vehicle and is used to visually display stored images and videos.
[1196] 2. Software Configuration
[1197] Voice Recognition Technology
[1198] For voice recognition, a library called "SpeechRecognition" is used. Voice data is acquired from the microphone, analyzed, and instructions are determined. For example, if the user says "Take a picture," the voice is analyzed and the camera is activated.
[1199] Image capture and storage
[1200] Images captured by the camera are processed using the "OpenCV" library and temporarily stored in the in-vehicle system before being sent to a cloud server via the network, using the "Requests" library to send the data.
[1201] Cloud Server
[1202] The cloud server receives the captured images and stores them in a database, and when a user searches for a specific image, it also retrieves the relevant images from the database based on the relevant query.
[1203] 3. Processing Procedure
[1204] Shooting instructions
[1205] When a user issues a voice command to the camera to "take a picture," the microphone captures the voice and the voice is analyzed using voice recognition technology. The camera then starts up and takes a picture of the scene as instructed.
[1206] Cloud Upload
[1207] The captured images are temporarily stored in the vehicle and then transmitted to a cloud server using the network communication function. The cloud server stores the received images in a database.
[1208] Searching and viewing images
[1209] When the user issues a voice command, such as "Show me today's sunset," the command is again analyzed through voice recognition technology. The cloud server retrieves relevant images from the database and displays them on the vehicle's display.
[1210] Examples
[1211] For example, if a user says "Take a picture" while in the car, the camera will automatically capture the scenery and store the image in the cloud. Later, if the user says "Show me today's sunset," the relevant image will be retrieved from the cloud and displayed on the display in front of the passenger.
[1212] Examples of prompt statements
[1213] Imagine an app that takes a picture of a sunset with an in-car camera, and when the driver says "Show me today's sunset," the app displays the saved image on the in-car display. Please submit a detailed program outline.
[1214] The flow of the specific process in the application example 1 will be described with reference to FIG.
[1215] Step 1:
[1216] The user issues a voice command
[1217] The user says "take a picture" into the microphone, which causes voice data (input) to be captured by the microphone.
[1218] Step 2:
[1219] The device analyzes the voice command
[1220] The terminal analyzes the captured voice data using voice recognition technology (e.g., the SpeechRecognition library) and obtains (outputs) the voice command "Take a picture" as text data.
[1221] Step 3:
[1222] The camera starts and takes a picture
[1223] The device activates the camera based on the results of voice recognition. The camera captures the specified scene and obtains image data (input). The obtained image data is temporarily stored inside the device (output).
[1224] Step 4:
[1225] The device sends image data to the cloud.
[1226] The device transmits the temporarily stored image data to the cloud server using a network communication technology (e.g., the Requests library) (input). The cloud server stores the received image data in a database (output).
[1227] Step 5:
[1228] The user issues a search command
[1229] The user says into the microphone, "Show me today's sunset." This causes the microphone to capture voice data (input) of the search command.
[1230] Step 6:
[1231] The device parses the search instructions
[1232] The device analyzes the captured voice data using voice recognition technology and obtains the voice instruction "Show me today's sunset" as text data (output).
[1233] Step 7:
[1234] The cloud server searches for related images
[1235] The device sends the analyzed text data (input) to the cloud server, which searches the database for related image data that matches the text data and retrieves them (output).
[1236] Step 8:
[1237] Display the image captured by the device
[1238] Relevant image data obtained from the cloud server (input) is sent to the terminal, which then displays the received image data on a display inside the vehicle (output).
[1239] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1240] The embodiment for carrying out the present invention comprises the following elements.
[1241] 1. Terminal: A device with an interface that accepts voice commands and hand gestures when a user wants to save a particular scene or experience.
[1242] 2. Shooting: The device analyzes voice commands and hand gestures to understand the command to shoot, activates the camera, and shoots the specified scene or experience.
[1243] 3. Transmission means: The terminal has a communication function for transmitting captured images to the cloud. Images can be transmitted to a server via a network.
[1244] 4. Server: Stores the received images in a database, accepts queries about the contents stored by the user, and retrieves and displays related captured images.
[1245] 5. Emotion engine: The server combines an emotion engine and utilizes technology to recognize the user’s emotions, such as voice and facial expressions. By analyzing the user’s emotions at the time of shooting and associating them with the captured image, the server can automatically classify and search for images based on emotions.
[1246] (Specific examples)
[1247] When a user uses the robot 414 to implement the present invention, the robot 414 becomes a terminal. The user commands the robot 414 to take a picture by giving a voice command such as "save it" or by making a specific gesture. The robot 414 analyzes the voice command or gesture and understands the command to take a picture. The camera starts up and takes a picture of the instructed scene or experience. The captured image is sent to the cloud using the communication function of the robot 414. The server stores and saves the received image in a database. When the user asks the robot 414 a question about the saved content, the robot 414 sends the question to the server. The server analyzes the received question, retrieves related captured images from the database, and sends them to the robot 414. The robot 414 displays the received captured images. The server also combines an emotion engine and utilizes technology to recognize the user's emotions, such as voice and facial expressions. By analyzing the user's emotions at the time of shooting and associating them with the captured image, automatic classification and search of images based on emotions are performed.
[1248] The above is an example of an embodiment of the present invention. Although specific devices and processing procedures may vary depending on the selection and implementation of the implementer, the present invention can be implemented based on this embodiment.
[1249] The process flow will be explained below.
[1250] Step 1: The user speaks into the terminal, saying "Save."
[1251] Step 2: The device receives the voice command and uses voice recognition technology to interpret the command.
[1252] Step 3: Based on the analysis results, the device understands the shooting instructions.
[1253] Step 4: The device activates its camera and captures the specified scene or experience.
[1254] Step 5: The device uses its communication functions to send the captured images to the cloud.
[1255] Step 6: The server stores the received images in a database.
[1256] Step 7: The user asks a question to the terminal.
[1257] Step 8: The terminal sends the query to the server.
[1258] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[1259] Step 10: The server transmits the captured image to the terminal.
[1260] Step 11: The terminal displays the captured image that it has received.
[1261] Step 12: The server uses the emotion engine to recognize the user's emotions, such as voice and facial expressions.
[1262] Step 13: The server analyzes the user's emotion at the time of taking the photograph and associates it with the photographed image.
[1263] Step 14: The server performs automatic classification and retrieval of images based on emotions.
[1264] Examples:
[1265] Step 1: The user commands the robot 414 to "save."
[1266] Step 2: The robot 414 receives the voice instruction and uses voice recognition technology to interpret the instruction.
[1267] Step 3: The robot 414 understands the shooting instructions based on the analysis results.
[1268] Step 4: The robot 414 launches a camera app and captures the instructed scene or experience.
[1269] Step 5: The robot 414 uses its communication capabilities to send the captured images to the cloud.
[1270] Step 6: The server stores the received images in a database.
[1271] Step 7: The user asks the robot 414, "Show me pictures of playing in the park yesterday."
[1272] Step 8: The robot 414 sends the question to the server.
[1273] Step 9: The server analyzes the received query and retrieves the relevant captured images from the database.
[1274] Step 10: The server transmits the captured image to the robot 414.
[1275] Step 11: The robot 414 displays the captured image that it has received.
[1276] Step 12: The server uses the emotion engine to recognize the user's emotions, such as voice and facial expressions.
[1277] Step 13: The server analyzes the user's emotion at the time of taking the photograph and associates it with the photographed image.
[1278] Step 14: The server performs automatic classification and retrieval of images based on emotions.
[1279] The above is the specific process flow of this system. The user gives voice instructions or gestures, and the device analyzes them and performs appropriate processing, enabling the system to take photos, save images, answer questions, and classify and search images based on emotions.
[1280] Example 2
[1281] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."
[1282] Conventional image storage systems make it difficult for users to instantly record specific moments and easily search and display them when needed. Conventional systems also lack the ability to classify and search for images based on the user's emotions, and therefore lack the functionality to allow users to replay memories based on their emotions.
[1283] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for a user to give a shooting instruction by voice instruction or hand gesture, a means for a terminal to analyze the voice instruction or gesture, start a camera and shoot a scene, a means for a terminal to transmit the captured image to a cloud server, a means for a server to store the image in a database, a means for a server to accept a related question, obtain related images from the database and transmit them to the terminal, and a means for a server to analyze a user's emotion using an emotion engine and associate the emotion with the image. This allows a user to easily record a particular moment and later search and display images based on the emotion.
[1284] "User" refers to a person who uses the System to store and view specific scenes or experiences.
[1285] The term "terminal" refers to a device that allows a user to give shooting instructions by voice or gestures.
[1286] "Voice instruction" refers to an instruction given by voice to a terminal by a user.
[1287] A "hand gesture" refers to a user giving instructions to a terminal by performing a specific action.
[1288] "Camera" refers to a device built into a terminal that takes photos and videos at the user's command.
[1289] A "cloud server" refers to a remote server accessible via a network, which provides a place to store image data.
[1290] A "database" refers to a system built within a server that systematically stores captured images and related information.
[1291] A "related question" refers to an inquiry that a user makes to the system regarding images that he or she has previously saved and their contents.
[1292] An "emotion engine" is a technology that analyzes voice and facial expressions to recognize the user's emotions, and then classifies and searches for images based on the results.
[1293] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[1294] This invention is a system that allows users to easily save specific scenes or experiences, and later search and display those images based on emotions. The system consists of a terminal operated by the user and a server installed on the cloud. The main hardware and software configurations are explained below.
[1295] Terminal
[1296] The terminal is a device that allows the user to give instructions to shoot using voice commands or hand gestures. Specifically, this applies to smartphones, tablets, smart glasses, etc. The terminal has the following main functions:
[1297] 1. Voice Recognition Function:
[1298] The terminal analyzes the user's voice instructions using well-known natural language processing (NLP) techniques.
[1299] 2. Gesture recognition:
[1300] The device is equipped with technology that uses cameras and sensors to recognize the user's hand movements and analyze gestures.
[1301] 3. Camera Function:
[1302] It has a built-in high-resolution camera, and after understanding the shooting instructions, it will capture the specified scene.
[1303] 4. Communication function:
[1304] The captured images are sent to a cloud server using networks such as Wi-Fi, 4G / 5G.
[1305] server
[1306] The server is installed in the cloud and has the following main functions:
[1307] 1. Database Management:
[1308] Image data is efficiently stored and saved in a database. The database can use Amazon Web Services (AWS) RDS.
[1309] 2. Question analysis function:
[1310] The questions entered by the user are analyzed using natural language processing technology and related images are linked.
[1311] 3. Emotion Engine:
[1312] It uses voice and facial expression analysis technology to recognize user emotions, and uses that data to tag images with emotions for automatic classification and emotion-based search.
[1313] 4. Image transmission function:
[1314] The image data retrieved in response to the user's question is sent to the terminal using encryption protocols such as SSL / TLS.
[1315] Examples
[1316] For example, suppose a user says to the device "Save it" during a family trip. The device uses voice recognition technology to analyze this voice command, activates the camera, and captures beautiful scenery. The captured image is then sent to a cloud server using the device's communication function. The sent image is stored in a database on the server. Later, when the user issues a voice command such as "Show me photos of the beach from last summer," the server analyzes the question, searches the database for the corresponding image, and sends it to the user's device. The device displays the received image on the screen. The server's emotion engine also recognizes emotions such as "joy" from the user's voice and facial expression when taking the photo, and associates and saves the emotion tag with the image, so that if the user commands "Show me photos of fun times," an image associated with the joyful emotion will be displayed.
[1317] Example prompts for generative AI models
[1318] "Please explain how the system works, allowing users to take pictures of specific scenes using voice commands and store them in the cloud."
[1319] "How can I use an emotion engine to automatically classify images based on user emotions?"
[1320] "Please explain with a concrete example the process of taking pictures using a smartphone and saving them to the cloud."
[1321] The above is a detailed embodiment of the present invention. This system allows users to easily operate it and store and search images based on emotions, greatly improving the user experience.
[1322] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1323] Step 1:
[1324] The user issues a command to take a picture. Specifically, the user issues a voice command to the smartphone saying "save it" or performs a specific gesture. The input data is a voice signal or a gesture movement, which is input to the terminal. The output is a flag indicating that the command to take a picture has been recognized.
[1325] Step 2:
[1326] The device analyzes the voice command and gestures. The device analyzes the voice command using well-known natural language processing techniques. It also recognizes certain actions using gesture recognition techniques. The input data is the user's voice or gestures, and the output is the analyzed command. Specifically, the voice command "save" is recognized, and an instruction to start the camera is issued.
[1327] Step 3:
[1328] The device activates the camera and captures the scene. The camera understands the specified instructions and takes the picture. The input data is the parsed instructions, and the output is the captured image data. For example, a beautiful landscape during a family trip is captured as a high-resolution photo.
[1329] Step 4:
[1330] The device sends the captured image to the cloud server. The device uses Wi-Fi or 4G / 5G networks to send encrypted image data to the cloud server. The input data is the captured image data, and the output is the image data sent to the cloud server.
[1331] Step 5:
[1332] The server stores the images in a database. The server organizes the image data it receives and stores it in a database. The input data is image data sent via the cloud, and the output is image entries stored in the database. For example, AWS's RDS is used to efficiently store and manage image data.
[1333] Step 6:
[1334] The user asks a question related to the saved content. The user issues a voice command to the smartphone, such as "Show me a photo of the beach from last summer." The input data is the voice command to ask the question, and the output is the content of the question to the server.
[1335] Step 7:
[1336] The server analyzes the question and retrieves related images. Using natural language processing technology, the server analyzes the user's question and searches the database for relevant image data. The input data is the question based on the voice command, and the output is related image data. For example, "photos of the sea from last summer" is searched for.
[1337] Step 8:
[1338] The server sends the relevant images to the user's device. The server encrypts the image data it finds and sends it to the user's device. The input data is the image data retrieved from the database, and the output is the image data sent to the device.
[1339] Step 9:
[1340] The user's device displays the image. The device displays the received image to the user. The input data is the image data received from the server, and the output is the image displayed on the device's display. For example, a photo of a summer beach is displayed on the screen of a smartphone.
[1341] Step 10:
[1342] The server performs emotion analysis using an emotion engine and associates it with the image. The server analyzes the voice and facial expression data to recognize the user's emotion. The image data with the emotion tag is saved for future searches. The input data is the voice and facial expression data at the time of shooting, and the output is image data with the emotion tag. For example, if the user is feeling "joy," that emotion tag is associated with the image.
[1343] (Application example 2)
[1344] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".
[1345] In modern commerce, when consumers visit a store and select products, they need a way to easily save specific scenes and favorite products for later reference. Furthermore, there is no system that can classify and search saved information based on consumers' emotions. This leaves users lacking a way to efficiently manage their shopping experience and receive emotion-based feedback.
[1346] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1347] In this invention, the server includes means for accepting shooting instructions by voice instructions or hand gestures, means for transmitting the captured images to the cloud, means for retrieving and displaying the images stored in the cloud based on associated questions, means for analyzing the user's emotions, and means for organizing and searching for images based on the emotion data.
[1348] This allows users to easily save specific scenes or products and later categorize and search for images based on their emotions, resulting in a more fulfilling shopping experience.
[1349] A "capture command" refers to a voice command or hand gesture made by a user to record a particular scene or experience.
[1350] "Voice instruction" refers to a method in which a user instructs a system to perform a specific operation using voice.
[1351] "Hand gestures" refers to the way in which a user uses hand movements to instruct a system to perform certain actions.
[1352] The term "captured image" refers to a still image captured by a camera based on a user's instruction to capture a picture.
[1353] "Cloud" refers to a technological infrastructure that stores and manages data on remote servers provided via the Internet.
[1354] "Communication means" refers to the technical means and devices for transmitting and receiving data.
[1355] "User's emotions" refers to psychological reactions and sensations analyzed from the user's voice and facial expressions.
[1356] "Emotion data" refers to data obtained by analyzing information regarding a user's emotions.
[1357] An "emotion engine" is a technology that automatically identifies and analyzes a user's emotions from their voice, facial expressions, etc., and makes the results available within the system.
[1358] "Server" refers to a computer system on a network that provides and manages information and data.
[1359] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, an embodiment of the present invention will be described.
[1360] This invention uses a smartphone, smart glasses, or other smart devices as a terminal. The terminal has a built-in camera and microphone, and an interface that accepts voice instructions and hand gestures. This allows users to easily capture and save specific scenes and experiences.
[1361] The device uses voice recognition technology to analyze the user's voice instructions. The "SpeechRecognition" library is used for voice recognition. For example, when the user says "save it," the device will start the camera and take a picture of the scene. It also uses the "OpenCV" library to recognize hand gestures, so it is possible to give a shooting command using specific gestures.
[1362] The captured images are sent to the cloud using the device's communication function. The transmission to the cloud is done using the "Requests" library, and the captured images are stored on the cloud server.
[1363] Images stored on the cloud are retrieved and displayed in response to a user's question. A question is accepted by the device and sent to the cloud server. The cloud server analyzes the question, retrieves the associated image, and sends it to the device.
[1364] In addition, the system analyzes the voice data emitted by the user when taking a photo and uses an emotion engine to recognize the user's emotions. The emotion engine receives the voice data as input and analyzes the emotions. The emotion analysis results are associated with the image and stored in the cloud.
[1365] Image sorting and retrieval based on emotions is performed by the emotion engine on the cloud server. For example, it is possible to search for only images of scenes that the user felt were "fun."
[1366] Examples:
[1367] When a user is shopping in a brick-and-mortar store, they can say "Save this" when they see a product they like. The device that receives this command activates the camera and takes a picture of the scene. This picture is immediately sent to the cloud and saved. If the user then asks, "Show me some fun products," the cloud server uses the emotion engine to search for images associated with "fun" and sends them to the device, allowing the user to check the product again.
[1368] Example prompt:
[1369] "I want to create an application that allows users to easily save their favorite products or specific displays they find in physical stores. The user can take a picture by voice command or gesture, and the captured image will be sent to the cloud. Also, the application will analyze the user's emotions based on the voice data, and later create a product list based on the emotions. Please generate a program for this application."
[1370] The flow of the specific process in the application example 2 will be described with reference to FIG.
[1371] Step 1:
[1372] The device accepts voice commands or hand gestures. The user issues a voice command such as "save" or makes a specific hand gesture. The device receives voice or gestures as input using a microphone or camera, and analyzes them using voice recognition technology (SpeechRecognition library) or gesture recognition technology (OpenCV library). As a result of the analysis, it determines whether there is an instruction to take a photo.
[1373] Input: User's voice commands or hand gestures
[1374] Output: Whether or not shooting is instructed
[1375] Step 2:
[1376] The device starts up the camera and takes a picture of the specified scene. The captured image data is generated and temporarily stored in the device's memory.
[1377] Input:Shooting instructions
[1378] Output: Captured image data
[1379] Step 3:
[1380] The device sends the captured image data to the cloud. The captured image data is sent to the cloud server using the communication function (Requests library). The cloud server stores the received image data in a database.
[1381] Input: Photographed image data
[1382] Output: Image data stored in the cloud
[1383] Step 4:
[1384] The user asks a question about an image stored in the cloud. The user enters the question into the smart device, which then sends the question to the cloud server.
[1385] Input: User question
[1386] Output: A query request to the cloud server
[1387] Step 5:
[1388] The cloud server acquires information related to the stored images, searches the database for the corresponding image data based on the user's question, and transmits the searched image data to the terminal.
[1389] Input: A query request to the cloud server
[1390] Output: Retrieved image data
[1391] Step 6:
[1392] The terminal displays the searched image data. The terminal that receives the image data sent from the cloud server displays it on the screen and provides it to the user.
[1393] Input: Retrieved image data
[1394] Output: Display the image to the user
[1395] Step 7:
[1396] The cloud server uses an emotion engine to analyze the user's voice data and recognize the user's emotions. The emotion engine receives the voice data as input and executes an emotion analysis process. As a result of the analysis, it generates user emotion data and stores it in association with the image data.
[1397] Input: User's voice data
[1398] Output: Emotion data and associated image data
[1399] Step 8:
[1400] The cloud server organizes and searches image data based on the user's emotion data. For example, if a user asks, "Show me some fun products," the cloud server uses the emotion engine to search for images associated with "fun" and sends them to the device.
[1401] Input: emotion data and user questions
[1402] Output: Image data of sentiment-based search results
[1403] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1404] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1405] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the robot 414.
[1406] The emotion identification model 59 as an emotion engine may determine the emotion of the user according to a specific mapping. Specifically, the emotion identification model 59 may determine the emotion of the user according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the emotion of the robot, and the identification processing unit 290 may perform identification processing using the emotion of the robot.
[1407] FIG. 9 is a diagram showing an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive emotions are arranged. The more outside the concentric circles, the more emotions that represent states and actions that arise from a state of mind are arranged. Emotions are a concept that includes emotions and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions that occur in the brain are arranged. On the right side of the concentric circles, emotions that are generally induced by situational judgment are arranged. On the upper and lower sides of the concentric circles, emotions that are generally generated from reactions that occur in the brain and are induced by situational judgment are arranged. In addition, on the upper side of the concentric circles, emotions of "pleasure" are arranged, and on the lower side, emotions of "discomfort" are arranged. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1408] These emotions are distributed in the three o'clock direction of emotion map 400 and usually fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1409] The inside of emotion map 400 represents what is going on inside one's mind, and the outside of emotion map 400 represents behavior, so the further out on emotion map 400 you go, the more visible the emotions become (the more they are expressed in behavior).
[1410] Here, human emotions are based on various balances such as posture and blood sugar level, and when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. Emotions can also be created for robots, cars, motorcycles, etc., based on various balances such as posture and battery level, so that when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. The emotion map may be generated, for example, based on the emotion map of Dr. Mitsuyoshi (Research on speech emotion recognition and emotion brain physiological signal analysis system, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). On the left half of the emotion map, emotions belonging to an area called "reaction" where sensation is dominant are lined up. On the right half of the emotion map, emotions belonging to an area called "situation" where situation recognition is dominant are lined up.
[1411] The emotion map defines two emotions that promote learning. The first is the negative emotion around the middle of "repentance" or "remorse" on the situation side. In other words, this is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the positive emotion around "desire" on the response side. In other words, this is when the robot has positive feelings such as "I want more" or "I want to know more."
[1412] The emotion identification model 59 inputs the user input to a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the emotion of the user. This neural network is pre-trained based on multiple learning data that are combinations of the user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in Fig. 10. Fig. 10 shows an example in which multiple emotions, "relief," "calm," and "encouraging," have similar emotion values.
[1413] Although the system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, the system according to the present disclosure is not necessarily implemented in a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program that runs on a personal computer, or an application that runs on a smartphone or the like. The method according to the present disclosure may be provided to a user in the form of SaaS (Software as a Service).
[1414] In the above embodiment, an example is given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to input data.
[1415] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable computer-readable non-transitory storage medium such as a Universal Serial Bus (USB) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1416] In addition, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 upon request from the data processing device 12.
[1417] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1418] As the hardware resource for executing the specific process, various processors as shown below can be used. An example of the processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing the specific process by executing software, i.e., a program. Another example of the processor is a dedicated electric circuit, which is a processor having a circuit configuration designed exclusively for executing the specific process, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), or an Application Specific Integrated Circuit (ASIC). Each processor has a built-in or connected memory, and each processor executes the specific process by using the memory.
[1419] The hardware resource that executes the specific process may be one of these various processors, or may be a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.
[1420] As an example of a configuration using one processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a configuration using a processor that realizes the functions of the entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1421] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements. The specific processes described above are merely examples. It goes without saying that unnecessary steps may be deleted, new steps may be added, or the order of processes may be changed without departing from the spirit of the invention.
[1422] The above description and illustrations are detailed descriptions of the parts related to the technology of the present disclosure, and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, function, action, and effect is an example of the configuration, function, action, and effect of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above description and illustrations, within the scope of the gist of the technology of the present disclosure. In addition, in order to avoid confusion and to facilitate understanding of the parts related to the technology of the present disclosure, the above description and illustrations omit explanations of technical common sense that do not require explanation in order to enable the implementation of the technology of the present disclosure.
[1423] All publications, patent applications, and standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or standard was specifically and individually indicated to be incorporated by reference.
[1424] The following supplementary notes are further disclosed regarding the above embodiment.
[1425] (Appendix 1) a means for receiving a shooting instruction by voice instruction or hand gesture; A means for transmitting the captured image to a cloud; means for retrieving and displaying the cloud stored images based on the associated questions; A system including:
[1426] (Appendix 2) The system according to claim 1, wherein the means for accepting shooting instructions by voice instructions analyzes the voice instructions using voice recognition technology.
[1427] (Appendix 3) The system described in Appendix 1 or Appendix 2, characterized in that the means for transmitting the captured image to the cloud transmits the image to a server using network communication technology.
[1428] (Appendix 4) 5. The system according to any one of claims 1 to 4, further comprising an emotion engine for recognizing user emotions.
[1429] (Appendix 5) The system described in Appendix 4, characterized in that the emotion engine utilizes technology to recognize user emotions such as voice and facial expressions.
[1430] (Appendix 6) The system according to claim 4 or 5, wherein the emotion engine analyzes the user's emotion at the time of shooting and associates it with the captured image, thereby automatically classifying and searching for images based on emotion.
[1431] "Example 1" (Claim 1) a means for receiving a shooting instruction by voice instruction or hand gesture; means for analyzing the received photographing instruction, activating a camera, and photographing the scene; A means for transmitting the captured image to a cloud server via a network; A means for storing the received images in a database; means for receiving a question from a user about a stored image, and retrieving and displaying related images; A system including:
[1432] (Claim 2) 2. The system of claim 1, further comprising: a voice recognition technique for analyzing the voice instructions.
[1433] (Claim 3) The system according to claim 1, characterized in that the captured image is transmitted to a cloud server using network communication technology.
[1434] "Application example 1" (Claim 1) a means for receiving a shooting instruction by voice instruction or hand gesture; A means for transmitting the captured image to a cloud; A means for retrieving images stored in the cloud based on the associated questions and displaying them on a display device within the vehicle; A system including:
[1435] (Claim 2) 2. The system according to claim 1, wherein the means for accepting the shooting instruction by voice instruction analyzes the voice instruction by utilizing a voice recognition technology.
[1436] (Claim 3) The system according to claim 1, characterized in that the means for transmitting the captured image to the cloud transmits the image to the server using network communication technology.
[1437] "Example 2 of combining emotion engines" (Claim 1) A means for allowing a user to give instructions for photographing by voice or hand gestures; a means for activating a camera in the terminal based on the voice instruction or gesture to capture a scene; A means for transmitting images captured by the terminal to a cloud server; A means for the server to store the images in a database; A means for the server to receive a related query, retrieve related images from a database, and transmit the images to the terminal; A means for the server to analyze the user's emotion using an emotion engine and associate the emotion with the image; A system including:
[1438] (Claim 2) 2. The system according to claim 1, characterized in that the terminal uses voice recognition technology to analyze the voice instructions.
[1439] (Claim 3) The system according to claim 1, characterized in that the terminal transmits the captured images to a cloud server using network communication technology.
[1440] "Application example 2 when combining emotion engines" (Claim 1) a means for receiving a shooting instruction by voice instruction or hand gesture; A means for transmitting the captured image to a cloud; means for retrieving and displaying the cloud stored images based on the associated questions; A means for analyzing user emotions; A means to organize and search images based on emotion data; A system including:
[1441] (Claim 2) 2. The system according to claim 1, wherein the means for accepting the shooting instruction by voice instruction analyzes the voice instruction by utilizing a voice recognition technology.
[1442] (Claim 3) The system according to claim 1, characterized in that the means for transmitting the captured image to the cloud transmits the image to the server using network communication technology.
[1443] (Claim 4) 2. The system according to claim 1, wherein the means for analyzing the user's emotions utilizes a technology for analyzing emotions using voice data.
[1444] (Claim 5) 2. The system according to claim 1, wherein the means for organizing and searching for images based on emotion data utilizes an emotion engine to automatically classify and search for images. [Explanation of symbols]
[1445] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for receiving a shooting instruction by voice instruction or hand gesture; A means for analyzing the received photographing instruction and activating the camera to photograph; A means for storing the captured images in a database; means for accepting queries from a user regarding the stored images; means for retrieving an image related to the question from the database and transmitting the image to the user's terminal; A system including:
2. The method further includes a means for recognizing an emotion of the user at the time of shooting using an emotion engine, The system according to claim 1 , wherein the means for storing in the database associates the recognized emotion with the image and stores the association in the database.
3. The means for accepting a question accepts a question including the emotion at the time of shooting, The system according to claim 2 , wherein the means for transmitting to the terminal retrieves an image related to the accepted question including the emotion from the database and transmits the image to the terminal.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A