system
A system using mobile devices and generative AI for image recognition allows visually impaired individuals to independently locate products in stores by converting image analysis into voice guidance, addressing the challenge of product location in supermarkets.
Patent Information
- Application Number
- JP2024161798
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-09-19
- Filing Date
- 2024-09-19
- Publication Date
- 2025-10-22
- Estimated Expiration
- 2044-09-19
AI Technical Summary
Visually impaired individuals face difficulties in locating specific products in stores such as supermarkets, relying heavily on assistance from others due to inadequate image recognition technologies lacking accuracy and real-time capabilities.
A system utilizing a mobile device to capture images, upload them to a server, generate a URL, link it to a generative AI for image recognition, and communicate location information via voice, enabling independent product location.
Enables visually impaired individuals to independently determine the location of specific products by converting image analysis into voice guidance, enhancing their shopping independence.
Smart Images

Figure 0007758820000001 
Figure 0007758820000002 
Figure 0007758820000003
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] There is a problem that visually impaired people have difficulty in locating specific products in stores such as supermarkets. [Means for solving the problem]
[0005] To solve this problem, we provide a system that includes an image capture means, a means for uploading captured images, a means for linking the URL of the uploaded image to a generative AI, and a means for the generative AI to recognize the image and communicate the location information of specific objects to the user. Specifically, a specific area inside a supermarket is captured and the image is linked to the generative AI. The generative AI then communicates the location information of specific objects in the image to the user in the form of "in front," "to the right, in front," "not in front," etc. This enables visually impaired people to understand the location of specific products. [Brief explanation of the drawings]
[0006] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 2 is a sequence diagram showing a flow of processing in the data processing system according to the first embodiment of the first form example. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1 of Embodiment 1. [Figure 13] FIG. 10 is a sequence diagram showing a processing flow of a data processing system in a second embodiment of the second form example. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 of Embodiment Example 2. [Figure 15] FIG. 10 is a sequence diagram showing the flow of processing in a data processing system according to a third embodiment of the third embodiment. [Figure 16] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 3 of Embodiment 3. [Figure 17] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the first embodiment of the first form example when an emotion engine is combined. [Figure 18] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1 of Form Example 1 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0007] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0008] First, the terms used in the following description will be explained.
[0009] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (TENSOR PROCESSING UNIT (registered trademark)).
[0010] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0011] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0012] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0013] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0014] [First embodiment]
[0015] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0016] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0017] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0018] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0019] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0020] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0021] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0022] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0023] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0024] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0025] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0026] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.
[0027] "Example 1"
[0028] One embodiment of the present invention provides a system that enables visually impaired people to determine the location of specific products in a supermarket or other store. In this system, a user takes a photo of a specific area in the store using a device such as a smartphone. The captured image is automatically uploaded by the system, and a URL for the image is generated. The generated URL is linked to a generative AI. The generative AI analyzes the location information of a specific object in the image, such as a "daikon radish," and communicates this information to the user in the form of "in front," "to the right, in front," or "not in front." This enables visually impaired people to determine the location of specific products.
[0029] "Example 2"
[0030] One embodiment of the present invention provides a system that enables visually impaired people to determine the location of specific products in a supermarket or other store. In this system, a user takes a photo of a specific area in the store using a device such as a smartphone. The captured image is automatically uploaded by the system, and a URL for the image is generated. The generated URL is linked to a generative AI. The generative AI analyzes the location information of a specific object in the image, such as a "daikon radish," and communicates this information to the user in the form of "in front," "to the right, in front," or "not in front." This enables visually impaired people to determine the location of specific products.
[0031] The processing flow of each embodiment will be described below.
[0032] "Example 1"
[0033] Step 1: A visually impaired person uses a device such as a smartphone to take a photo of a specific area inside a store such as a supermarket.
[0034] Step 2: The system will automatically upload the captured image and generate a URL for it.
[0035] Step 3: The generated URL is linked to the generative AI.
[0036] Step 4: The generative AI analyzes the location information of a specific object in the image, such as a radish.
[0037] Step 5: The generative AI communicates the location of the specific object to the user in the form of "in front of you, to the right, in front of you, not in front of you", etc.
[0038] "Example 2"
[0039] Step 1: A visually impaired person uses a device such as a smartphone to take a photo of a specific area inside a store such as a supermarket.
[0040] Step 2: The system will automatically upload the captured image and generate a URL for it.
[0041] Step 3: The generated URL is linked to the generative AI.
[0042] Step 4: The generative AI analyzes the location information of a specific object in the image, such as a radish.
[0043] Step 5: The generative AI communicates the location of the specific object to the user in the form of "in front of you, to the right, in front of you, not in front of you", etc.
[0044] Example 1
[0045] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0046] It is extremely difficult for visually impaired people to locate specific products in supermarkets and other commercial facilities. Conventional methods require the visually impaired to rely on others for assistance, making it difficult for them to shop independently. Furthermore, existing technologies lack the accuracy and real-time capabilities of image recognition, and are unable to provide sufficient support for visually impaired people.
[0047] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0048] In this invention, the server includes a means for a user to take a picture of a specific area using a mobile device, a means for uploading the taken image to the server, a means for generating a URL for the uploaded image, a means for linking the generated URL to a generative AI model, a means for the generative AI model to analyze location information of a specific object in the image, and a means for communicating the analyzed location information to the user by voice, thereby enabling visually impaired people to independently determine the location of a specific product.
[0049] "User" refers to any user, including visually impaired people, who use the system to locate a particular product.
[0050] "Mobile device" refers to a portable electronic device such as a smartphone or tablet.
[0051] "Specific area" refers to a specific location or section within a commercial establishment.
[0052] "Means for taking a photograph" refers to a method for capturing an image using the camera function built into the mobile terminal.
[0053] "Server" refers to a computer system for storing, processing, and managing data.
[0054] "Means for uploading" refers to a method for transmitting data from a mobile device to a server.
[0055] "URL" refers to an address used to identify a resource on the Internet.
[0056] "Means of generating" refers to the method of creating a URL that indicates the location where the uploaded image is saved.
[0057] A "generative AI model" refers to an artificial intelligence algorithm used for image recognition and data analysis.
[0058] "Means of collaboration" refers to the method of sending data from the server to the generative AI model.
[0059] "Means of analysis" refers to how the generative AI model identifies the location information of specific objects within an image.
[0060] "Means of communicating by voice" refers to a method of communicating the analysis results to the user by voice.
[0061] This invention is a system that enables visually impaired people to locate specific products in a commercial facility such as a supermarket. The system works by having the user take a picture of a specific area using a mobile device and uploading the image to a server.
[0062] Hardware and software used
[0063] Hardware:
[0064] Mobile devices (e.g. smartphones, tablets)
[0065] Server (e.g. cloud server, on-premise server)
[0066] software:
[0067] Mobile application for uploading images
[0068] Image management systems (e.g., cloud storage services)
[0069] API integration system (e.g. REST API)
[0070] Generative AI models (e.g., image recognition algorithms)
[0071] Voice assistants (e.g. text-to-speech software)
[0072] System Operation
[0073] The user takes a photo with their mobile device
[0074] Users can open the camera app on their mobile device and take a picture of a specific area in a shopping mall. For example, if a user is looking for daikon radishes in the vegetable section, they can take a picture of the entire vegetable section.
[0075] The device uploads the image to the server.
[0076] The captured images are automatically uploaded from the mobile device to the server, and a notification is displayed when the upload is complete.
[0077] The server generates the image URL
[0078] The server generates a URL for the uploaded image, which is then stored in the database.
[0079] The server links the URL to the generated AI model
[0080] The server connects the generated URL to the AI model, and the URL is sent via API.
[0081] A generative AI model analyzes the image
[0082] The generative AI model analyzes the location of specific objects (e.g., radishes) in the image, and the analysis results are output in text format.
[0083] Generative AI model communicates location information to user
[0084] Based on the analysis results, the generative AI model communicates the location information to the user via voice, and the voice assistant is activated to read the information aloud.
[0085] Specific examples
[0086] Let's say a user is looking for a "daikon radish" in the vegetable section of a supermarket. The user takes a photo of the entire vegetable section with their smartphone and uploads the image to the system. The system generates a URL for the image and connects it to the generative AI model. The generative AI model analyzes the image and tells the user by voice, "The daikon radish is in the front right."
[0087] Prompt Sentence Examples
[0088] "Please tell me the location of the radish in this image. Please tell the user in the form of 'in front, to the right, in front, not in front'."
[0089] This system allows visually impaired people to independently locate specific products.
[0090] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0091] Step 1:
[0092] Users take photos of the store interior using their mobile devices
[0093] Input: A user opens the camera app on their mobile device and takes a picture of a specific area.
[0094] Specific Actions: The user launches the camera app on their mobile device and presses the camera button to capture an image.
[0095] Output: The captured image data is saved on the mobile device.
[0096] Step 2:
[0097] The device uploads the image to the server.
[0098] Input: Captured image data
[0099] What it does: Your device will automatically upload the image to the server, and once the upload is complete, a notification will appear on your device.
[0100] Output: Image data is saved on the server.
[0101] Step 3:
[0102] The server generates the image URL
[0103] Input: Image data stored on the server
[0104] Specific operation: The server generates a URL indicating the location where the image is saved and saves that URL in the database.
[0105] Output: URL of the generated image (e.g. https: / / example.com / image123.jpg)
[0106] Step 4:
[0107] The server links the URL to the generated AI model
[0108] Input: URL of the generated image
[0109] Specific operation: The server sends the image URL to the generative AI model via API, and records the completion of the integration in a log.
[0110] Output: The image URL is sent to the generative AI model.
[0111] Step 5:
[0112] A generative AI model analyzes the image
[0113] Input: The URL of the image sent to the generative AI model
[0114] How it works: The generative AI model downloads an image and analyzes the location of a specific object (e.g., a radish). The analysis results are output in text format.
[0115] Output: Analysis result (e.g. "The radish is in the front right").
[0116] Step 6:
[0117] Generative AI model communicates location information to user
[0118] Input: Text data of analysis results
[0119] How it works: The generative AI model sends the analysis results to the voice assistant, which then reads the information aloud.
[0120] Output: The location information is spoken to the user (e.g., "The radish is in front of you to the right").
[0121] (Application example 1)
[0122] Next, a description will be given of Application Example 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0123] It is very difficult for visually impaired people to find specific products in commercial facilities such as supermarkets. With conventional methods, visually impaired people need the help of others, making it difficult for them to shop independently. To solve this problem, a system that allows visually impaired people to find specific products by themselves is needed.
[0124] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and communicate location information of the specific object to the user, and a means for providing audio feedback of the location information. This enables visually impaired people to find specific products in commercial facilities using their smartphones.
[0125] An "image capture means" is a device that allows a user to capture an image of a particular area.
[0126] "Means for uploading captured images" is a function for sending captured images to a cloud or server.
[0127] "Means for linking the URL of uploaded images to generative AI" is a function for providing the URL of uploaded images to generative AI.
[0128] "Generative AI" is artificial intelligence that performs image recognition and analyzes the location information of specific objects.
[0129] "Means of communicating the location information of specific objects to users" is a function that informs users of the location information analyzed by generative AI.
[0130] "Means for providing audio feedback of location information" is a function for conveying analyzed location information to the user via audio.
[0131] A "commercial facility" is a place where consumers can purchase goods, such as a supermarket or department store.
[0132] A system for implementing this invention is for assisting visually impaired people in finding specific products in commercial facilities. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to recognize the image and communicate the location information of the specific object to the user, and a means for providing audio feedback of the location information.
[0133] Users use their smartphone camera to take pictures of specific areas within a commercial facility. The images are then uploaded to the cloud via a smartphone application. A URL is automatically generated for the uploaded image, and this URL is linked to the generative AI.
[0134] The generative AI performs image recognition based on the provided image URL and analyzes the location information of a specific object, such as a "daikon radish." As a result of the analysis, the generative AI generates location information in the form of "in front," "to the right, in front," "not in front," etc. This location information is fed back to the user via audio via a smartphone application.
[0135] Specifically, the system works as follows:
[0136] 1. Hardware: Use your smartphone camera to capture the image.
[0137] 2. Software: Use OpenCV to capture images and requests library to upload images to cloud.
[0138] 3. Data processing: Once the image is uploaded to the cloud, a URL is generated.
[0139] 4. Data calculation: Send the image URL and item name to the generative AI, which analyzes the location information of the specific product.
[0140] 5. Feedback: The acquired location information is converted into audio using gTTS (Google (registered trademark) Text-to-Speech) and provided as feedback to the user.
[0141] For example, if the user is looking for "daikon radish," the prompt text might look like this:
[0142] Example prompt sentence:
[0143] Could you please tell me the location of the radish in the image?
[0144] By sending this prompt to the generative AI, the AI will analyze the location of the radish in the image and return location information in the form of "in front, to the right, in front, not in front," etc. This allows visually impaired people to find specific products on their own.
[0145] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0146] Step 1:
[0147] A user takes a picture of a specific area in a commercial facility using the camera on their smartphone. The input is the image captured by the user through the camera. The output is an image file stored on the smartphone.
[0148] Step 2:
[0149] The device uploads the captured image to the cloud. The input is the image file obtained in step 1. The output is the URL of the image stored on the cloud. Specifically, the device uses the requests library to send the image file to the cloud server.
[0150] Step 3:
[0151] The device connects the URL of the uploaded image to the generative AI. The input is the URL of the image generated in step 2. The output is the URL and prompt sent to the generative AI. Specifically, the device sends the generative AI a prompt saying, "Please tell me the location of a specific object in the image."
[0152] Step 4:
[0153] The server uses generative AI to perform image recognition and analyze the location information of specific objects. The input is the URL of the image sent in step 3 and the prompt text. The output is the location information of the specific object. Specifically, the generative AI analyzes the specific object in the image and generates location information such as "in front," "to the right and in front," or "not in front."
[0154] Step 5:
[0155] The device provides voice feedback of the location information obtained from the generative AI. The input is the location information obtained in step 4. The output is voice feedback provided to the user. Specifically, the device converts the location information into voice using gTTS (Google Text-to-Speech) and transmits it to the user through the smartphone speaker.
[0156] These steps allow a visually impaired person to locate a particular product within a commercial establishment.
[0157] Example 2
[0158] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0159] It is difficult for visually impaired people to locate specific products in commercial facilities such as supermarkets. With conventional methods, it is difficult for visually impaired people to find products on their own and they need to get help from others. To solve this problem, a system that allows visually impaired people to locate specific products on their own is needed.
[0160] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0161] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI model, a means for the generative AI model to perform image recognition and communicate location information of a specific object to the user, a means for transmitting the generated location information to the user's terminal, and a means for the user to check the location information on the terminal. This enables visually impaired people to independently determine the location of specific products.
[0162] An "image capturing means" is a device that allows a user to capture an image of a specific area within a commercial facility.
[0163] "Means for uploading captured images" is a function for sending image data from the user's terminal to the server.
[0164] "Means for linking the URL of uploaded images to the generative AI model" is a function that generates a URL for the image data received by the server and provides that URL to the generative AI model.
[0165] A "generative AI model" is an artificial intelligence model that uses image recognition technology to analyze the location information of specific objects within a provided image.
[0166] "Means of communicating location information of specific objects to users" is a function for communicating location information analyzed by the generative AI model to users.
[0167] "Means for transmitting the generated location information to the user's terminal" refers to a function that enables the server to transmit the location information received from the generating AI model to the user's terminal.
[0168] "Means for users to check location information on their devices" refers to a function that allows users to check received location information by voice or text using their own devices.
[0169] This invention is a system that enables visually impaired people to determine the location of specific products in a commercial facility. The system allows users to take pictures of specific areas in a store using a device such as a smartphone, and analyzes the images to provide location information for specific objects.
[0170] A user launches the camera app on their smartphone and takes a picture of a specific area in a shopping mall. For example, if the user is looking for a "daikon radish," they take a picture of the vegetable section. The captured image is automatically uploaded to the server through the smartphone's application. At this time, the device uses an Internet connection to send the image data to the server. Specifically, the image data is sent using an HTTP POST request.
[0171] The server stores the received image data and generates a URL for that image. The generated URL is linked to the generative AI model. Specifically, an API request is sent to the generative AI model, providing a prompt message including the image URL. The generative AI model analyzes the image based on the provided image URL. For example, it uses Google Cloud Vision API or Amazon Rekognition to identify the location of the "daikon radish" in the image. As a result of the analysis, location information such as "The daikon radish is in the front right" is generated.
[0172] The server sends the location information received from the generative AI model to the user's smartphone. Specifically, it sends data containing the location information as an HTTP response. The user then checks the received location information through an application on their smartphone. The application then communicates the location information to the user via voice or text. For example, it may provide a voice prompt saying, "The radish is in front of you to the right."
[0173] As a concrete example, consider the case where a user is looking for a "daikon radish" in a supermarket. The user takes a photo of the area inside the store with their smartphone and uploads the image to the system. The server generates a URL for the image and connects it to the generative AI model. The generative AI model analyzes the image and generates location information such as "The daikon radish is in the front right." The server sends this information to the user's smartphone, and the user confirms the location information by voice.
[0174] Example prompt sentence:
[0175] "Please tell me the location of the radish in this image."
[0176] This system allows visually impaired people to locate specific products.
[0177] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0178] Step 1:
[0179] The user takes a photo of a specific area inside the store using their smartphone.
[0180] Specifically, a user launches a camera app on their smartphone and takes a picture of a specific area, such as a vegetable section. The input is the image taken by the user, and the output is the image data stored on the smartphone.
[0181] Step 2:
[0182] The device uploads the captured image to the server.
[0183] Specifically, the device uses an Internet connection to send image data to a server. The input is image data stored on the smartphone, and the output is image data uploaded to the server. The image data is sent using an HTTP POST request.
[0184] Step 3:
[0185] The server generates a URL for the image and connects it to the generative AI model.
[0186] Specifically, the server saves the received image data and generates a URL for the image. The input is the image data uploaded to the server, and the output is the generated image URL. The generated URL sends an API request to the generative AI model and provides a prompt containing the image URL.
[0187] Step 4:
[0188] A generative AI model analyzes the image and generates location information for specific objects.
[0189] Specifically, the generative AI model analyzes an image based on the provided image URL. The input is the image URL and a prompt, and the output is the location information of a specific object. For example, the location of a "daikon radish" in the image can be identified using the Google Cloud Vision API or Amazon Rekognition. The analysis results in location information such as "The daikon radish is in the front right."
[0190] Step 5:
[0191] The server sends the generated location information to the user's smartphone.
[0192] Specifically, the server sends the location information received from the generative AI model to the user's smartphone. The input is the location information from the generative AI model, and the output is the location information sent to the user's smartphone. Data including the location information is sent as an HTTP response.
[0193] Step 6:
[0194] The user checks the location on their smartphone.
[0195] Specifically, the user checks the location information received through a smartphone application. The input is the location information sent from the server, and the output is the location information checked by the user. The application then communicates the location information to the user via voice or text. For example, it may provide a voice prompt saying, "The radish is in front of you to the right."
[0196] (Application example 2)
[0197] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0198] It is difficult for visually impaired people to locate specific products in stores such as supermarkets. Conventional methods have made it difficult for visually impaired people to find products on their own and have required the help of others. For this reason, there is a demand for a system that allows visually impaired people to shop independently.
[0199] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0200] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and communicate location information of a specific object to the user, a means for communicating the analysis result to the user by voice, a voice synthesis means for generating voice, and a voice playback means for playing voice, thereby enabling visually impaired people to locate specific products and shop independently.
[0201] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area inside the store.
[0202] "Means for uploading captured images" refers to devices or functions for sending captured images to a cloud or server.
[0203] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function for passing the URL of an uploaded image to the generative AI.
[0204] "Generative AI" is artificial intelligence that performs image recognition and analyzes the location information of specific objects.
[0205] "Means for communicating location information of a specific object to a user" refers to a device or function for informing a user of analyzed location information.
[0206] "Means for communicating analysis results to users via voice" refers to devices or functions that communicate the results of analysis by generative AI to users via voice.
[0207] "Speech synthesis means for generating speech" refers to a device or function for converting text information into speech.
[0208] The "audio playback means for playing back audio" refers to a device or function for allowing the user to hear the generated audio.
[0209] A system for carrying out this invention assists visually impaired people in locating specific products in a store such as a supermarket. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and communicate the location information of the specific object to the user, a means for communicating the analysis results to the user by voice, a voice synthesis means for generating voice, and a voice playback means for playing back voice.
[0210] Hardware and software used
[0211] Hardware: Smartphone (camera, microphone, speaker)
[0212] software:
[0213] OpenCV (Image Capture)
[0214] requests (HTTP requests)
[0215] gTTS (Google Text-to-Speech)
[0216] mpg321 (audio playback)
[0217] System Operation
[0218] 1. Image capture:
[0219] The user takes a photo of a specific area in the store using their smartphone camera, and the image is captured using OpenCV and saved locally.
[0220] 2. Image upload:
[0221] Upload the captured image to the cloud by using the requests library to send the image as a POST request to the specified URL.
[0222] 3. Image Analysis:
[0223] The URL of the uploaded image is linked to the generative AI. The generative AI analyzes the location information of specific objects in the image. The analysis results are returned in the form of, for example, "In front, in front to the right, not in front."
[0224] 4. Providing Feedback:
[0225] The analysis results are communicated to the user via voice. gTTS is used to convert the text information into voice and play it back in mpg321.
[0226] Specific examples
[0227] If a user is looking for a "daikon radish," they can take a photo of the inside of the store with their smartphone camera. The image is uploaded to the cloud, and generative AI analyzes the location of the "daikon radish." If the analysis returns "It's in front of you on the right," the smartphone will relay that information to the user via voice.
[0228] Prompt Sentence Examples
[0229] Analyze the location of the "radish" in the image. Return the result in the form of "In front, in front right, not in front", etc.
[0230] The system allows visually impaired people to locate specific products and shop independently.
[0231] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0232] Step 1:
[0233] A user uses a smartphone camera to take a picture of a specific area in a store. The input is real-time video from the camera, and the output is a still image. Specifically, a user launches the camera app, frames a specific area, and presses the shutter button to capture the image.
[0234] Step 2:
[0235] The device saves the captured image to local storage. The input is the image data captured in step 1, and the output is an image file saved in local storage. Specifically, the image is saved in JPEG format using the OpenCV library.
[0236] Step 3:
[0237] Uploads an image saved on the device to a cloud server. The input is an image file saved in local storage, and the output is an image URL on the cloud server. Specifically, the requests library is used to send the image file to the specified URL as a POST request.
[0238] Step 4:
[0239] The server connects the URL of the uploaded image to the generative AI. The input is the image URL on the cloud server, and the output is the URL data passed to the generative AI. Specifically, the server sends the URL to the API endpoint of the generative AI.
[0240] Step 5:
[0241] The generative AI analyzes the location information of a specific object in an image. The input is the image URL, and the output is the location information of the specific object. Specifically, the generative AI uses an image recognition algorithm to analyze the location of a specific object (e.g., a "radish") in the image and returns a result in the form of "In front, in front to the right, not in front of you," etc.
[0242] Step 6:
[0243] The server generates data to communicate the analysis results to the user via voice. The input is location information from the generative AI, and the output is text data for speech synthesis. Specifically, the analysis results are converted into text format and passed to the speech synthesis API.
[0244] Step 7:
[0245] The device generates the audio and plays it back to the user. The input is text data for speech synthesis, and the output is audio that the user hears. Specifically, the gTTS library is used to convert the text to an audio file and play it back in mpg321.
[0246] The above steps enable a visually impaired person to locate specific products and shop independently.
[0247] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0248] "Example 1"
[0249] In one embodiment of the present invention, the generative AI includes an emotion engine that recognizes a user's emotions. This emotion engine recognizes emotions from the user's tone of voice. For example, if a user says, "I don't know where the radish is," with a sense of frustration in their voice, the emotion engine recognizes that frustration. The generative AI then communicates location information for a specific object based on the recognized emotion. Specifically, if the generative AI recognizes that the user is feeling frustrated, it communicates location information more politely, such as, "Don't worry, the radish is three meters in front of you to your right." This makes it possible to provide a service that responds to the user's emotions.
[0250] "Example 2"
[0251] In one embodiment of the present invention, the generative AI includes an emotion engine that recognizes a user's emotions. This emotion engine recognizes emotions from the user's tone of voice. For example, if a user says, "I don't know where the radish is," with a sense of frustration in their voice, the emotion engine recognizes that frustration. The generative AI then communicates location information for a specific object based on the recognized emotion. Specifically, if the generative AI recognizes that the user is feeling frustrated, it communicates location information more politely, such as, "Don't worry, the radish is three meters in front of you to your right." This makes it possible to provide a service that responds to the user's emotions.
[0252] The processing flow of each embodiment will be described below.
[0253] "Example 1"
[0254] Step 1: The user says with frustration in their voice, "I don't know where the radish is."
[0255] Step 2: The emotion engine recognizes frustration from the user's tone of voice. Step 3: Based on the emotion recognized by the generative AI, the location information of the radish is communicated. Specifically, the location information is communicated more politely, such as "Don't worry, the radish is three meters in front of you to your right."
[0256] "Example 2"
[0257] Step 1: The user says with frustration in their voice, "I don't know where the radish is."
[0258] Step 2: The emotion engine recognizes frustration from the user's tone of voice. Step 3: Based on the emotion recognized by the generative AI, the location information of the radish is communicated. Specifically, the location information is communicated more politely, such as "Don't worry, the radish is three meters in front of you to your right."
[0259] Example 1
[0260] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0261] It is difficult for visually impaired people to locate specific products in supermarkets and other commercial facilities. It is also difficult to provide appropriate feedback based on the user's emotions. This often causes stress for visually impaired people when shopping. Therefore, there is a need for a system that allows visually impaired people to easily locate specific products and provides feedback based on the user's emotions.
[0262] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0263] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and communicate location information of a specific object to the user, an emotion engine that recognizes the user's emotions, and a means for providing feedback based on the emotions recognized by the emotion engine. This enables visually impaired people to easily determine the location of specific products and provides appropriate feedback according to the user's emotions.
[0264] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area within a commercial facility.
[0265] "Means for uploading captured images" refers to a device or function that allows a user to send captured images to a server.
[0266] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function that generates a URL for the image received by the server and sends that URL to the generative AI.
[0267] "Generative AI" is an artificial intelligence system that performs image recognition and analyzes the location information of specific objects.
[0268] "Means for communicating the location information of a specific object to the user" refers to devices or functions for notifying the user of the location information analyzed by generative AI.
[0269] An "emotion engine that recognizes user emotions" is a device or function that analyzes the tone of a user's voice and other emotional expressions to recognize the user's emotions.
[0270] "Means for providing feedback based on emotions recognized by the emotion engine" refers to devices or functions that enable the generative AI to provide appropriate feedback in response to the user's emotions recognized by the emotion engine.
[0271] MODE FOR CARRYING OUT THE INVENTION
[0272] This invention is a system that allows visually impaired people to determine the location of specific products in a commercial facility. The system allows users to take pictures of specific areas in a store using a device such as a smartphone, and by analyzing the images, provides location information for specific objects. It also includes a function to recognize the user's emotions and provide feedback according to the emotions.
[0273] Hardware and software used
[0274] 1. Smartphone (device)
[0275] Camera application: Used by users to take pictures of specific areas in the store.
[0276] Upload function: Used to send captured images to the server.
[0277] 2. Server
[0278] Image saving function: Saves received images and generates their URLs.
[0279] URL generation function: Generates a URL for the saved image and connects it to the generative AI.
[0280] 3. Generative AI
[0281] Image recognition software: Uses software such as Google Cloud Vision API or Amazon Rekognition to analyze the location of specific objects in an image.
[0282] Feedback function: Used to notify the user of the analysis results.
[0283] 4. Emotion Engine
[0284] Voice analysis software: such as IBM Watson® Tone Analyzer, is used to recognize emotions from the user's tone of voice.
[0285] Emotion-based feedback: Generative AI provides appropriate feedback based on the emotions it recognizes.
[0286] Specific examples
[0287] Consider a scenario in which a user takes a photo of a vegetable section in a shopping mall with their smartphone and says, "I don't know where the radishes are." In this case, the system operates as follows.
[0288] 1. The user launches the camera app on their smartphone and takes a picture of the vegetable section. When the user presses the capture button, the image is saved on the device.
[0289] 2. The device automatically uploads the image to the server using the server endpoint configured in the application (e.g., https: / / example.com / upload).
[0290] 3. The server receives and saves the image, generates a URL for the saved image (e.g., https: / / example.com / images / 12345.jpg), and sends the URL to the generative AI API.
[0291] 4. The generative AI receives the image URL and uses the Google Cloud Vision API to analyze the location of the "radish" in the image. As a result, it obtains location information such as "The radish is in the front right."
[0292] 5. The generative AI notifies the user of the analysis results using the smartphone's voice synthesis function. The user's smartphone will then say aloud, "The radish is in front of you to the right."
[0293] 6. The emotion engine analyzes the user's tone of voice and recognizes frustration. The generative AI then conveys the location more politely, such as "Don't worry, the radish is three meters in front of you to your right."
[0294] Prompt Sentence Examples
[0295] "Users can upload images they take with their smartphones and then analyze and tell us the location of specific items. If users get frustrated, we'll give them a more polite way to find the location."
[0296] This system allows visually impaired people to easily locate specific products within a commercial facility and provides appropriate feedback based on the user's emotions.
[0297] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0298] Program processing flow
[0299] Step 1:
[0300] A user takes a photo of a specific area in a commercial facility using their smartphone.
[0301] Input: The user launches the camera app and presses the capture button.
[0302] Data processing: The smartphone camera captures image data and stores it on the device.
[0303] Output: Captured image data.
[0304] Specific operation: The user launches the camera app on their smartphone and takes a picture of the vegetable section. When the user presses the capture button, the image is saved on the device.
[0305] Step 2:
[0306] The device uploads the captured image to the server.
[0307] Input: Image data stored on the device.
[0308] Data processing: The device sends image data to the server using an HTTP POST request.
[0309] Output: Image data uploaded to the server.
[0310] Specific behavior: The device automatically uploads the image to the server using the server endpoint configured in the application (e.g., https: / / example.com / upload).
[0311] Step 3:
[0312] The server generates a URL for the image and connects it to the generative AI.
[0313] Input: Image data uploaded to the server.
[0314] Data processing: The server saves the image and generates a URL for it. The generated URL is sent to the generative AI API.
[0315] Output: The generated image URL.
[0316] Specific operation: The server receives and saves the image. It generates a URL for the saved image (e.g., https: / / example.com / images / 12345.jpg) and sends that URL to the generative AI API.
[0317] Step 4:
[0318] Generative AI analyzes the location of specific objects within an image.
[0319] Input: Image URL sent to the generative AI.
[0320] Data processing: Generative AI uses image recognition software (e.g., Google Cloud Vision API) to analyze the location of a specific object (e.g., a "daikon radish") in an image.
[0321] Output: Location information of a specific object.
[0322] Specific operation: The generative AI receives the image URL and uses the Google Cloud Vision API to analyze the location of the "radish" in the image. As a result of the analysis, it obtains location information such as "The radish is in the front right."
[0323] Step 5:
[0324] The generative AI communicates the analysis results to the user.
[0325] Input: Location information of a specific object.
[0326] Data processing: The generative AI notifies the user of the analysis results via voice or text message.
[0327] Output: The location information communicated to the user.
[0328] Specific operation: The generative AI notifies the user of the analysis results using the smartphone's voice synthesis function. The user's smartphone will then say aloud, "The radish is in front of you to the right."
[0329] Step 6:
[0330] The emotion engine recognizes the user's emotions, and the generative AI provides feedback based on the emotions.
[0331] Input: The user's tone of voice.
[0332] Data processing: The emotion engine uses voice analysis software (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions. The generative AI provides feedback based on the recognized emotions.
[0333] Output: Emotion-based feedback.
[0334] How it works: The emotion engine analyzes the user's tone of voice and recognizes frustration. The generative AI then conveys the location information more politely, such as, "Don't worry, the radish is three meters in front of you to your right."
[0335] (Application example 1)
[0336] Next, a description will be given of Application Example 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0337] It is extremely difficult for visually impaired people to locate specific products in commercial facilities. Conventional systems make it difficult for visually impaired people to find products on their own, often causing them great stress when shopping. Furthermore, there is no system that provides feedback based on the user's emotions, making it impossible to reduce user frustration. A new system is needed to solve these issues.
[0338] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0339] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and communicate location information of a specific object to the user, an emotion recognition means for recognizing the user's emotion, and a means for adjusting feedback based on the emotion recognition result. This enables visually impaired people to locate specific products in a commercial facility and further reduces stress when shopping by providing feedback according to the user's emotion.
[0340] An "image capturing means" is a device or function that allows a user to capture an image of a specific area within a commercial facility.
[0341] "Means for uploading captured images" refers to a device or function for transferring captured image data to cloud storage or a server.
[0342] "Means for linking the URL of an uploaded image to a generative AI" refers to a device or function for providing the URL of an uploaded image to a generative AI.
[0343] "Generative AI" is an artificial intelligence system that analyzes uploaded images and generates location information for specific objects.
[0344] "Means for communicating the location information of a specific object to the user" refers to a device or function that notifies the user of the location information analyzed by the generative AI via voice or text.
[0345] The "emotion recognition means for recognizing the user's emotions" is a device or function for analyzing the emotions of the user from the tone of voice and facial expressions.
[0346] The "means for adjusting feedback based on emotion recognition results" refers to a device or function for adjusting the content and tone of feedback provided to the user based on the results of analysis by the emotion recognition means.
[0347] A system for implementing this invention assists visually impaired people in locating specific products within a commercial facility. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to recognize the image and communicate the location information of the specific object to the user, an emotion recognition means for recognizing the user's emotion, and a means for adjusting feedback based on the emotion recognition result.
[0348] Hardware and software used
[0349] Hardware: Smartphone (camera, microphone, speaker)
[0350] Software: Cloud storage services (e.g., AWS® S3), generative AI models (e.g., OpenAI® GPT-4®), emotion recognition engines (e.g., IBM Watson Tone Analyzer)
[0351] System Operation Overview
[0352] 1. Taking and uploading images
[0353] Users use their smartphone cameras to take pictures of specific areas within a commercial facility, and the images are uploaded from the smartphone to a cloud storage service (e.g., AWS S3).
[0354] 2. Image Analysis
[0355] The URL of the uploaded image is linked to a generative AI model (e.g., OpenAI GPT-4), which analyzes the location of specific objects (e.g., radishes) in the image and generates a result.
[0356] 3. Audio Feedback
[0357] The generated location information is then communicated to the user via voice via their smartphone, for example, in the form of a notification such as, "The radish is three meters in front of you and to your right."
[0358] 4. Emotion recognition
[0359] The tone of voice a user uses to speak into their smartphone is analyzed by an emotion recognition engine (e.g., IBM Watson Tone Analyzer), which analyzes the user's emotions and identifies feelings such as frustration or relief.
[0360] 5. Feedback adjustment
[0361] Based on the emotion recognition results, the generative AI will adjust the content and tone of the feedback. For example, if the user is feeling frustrated, it will provide more polite feedback, such as, "Don't worry, the radish is three meters in front of you to your right."
[0362] Specific examples
[0363] User: "I don't know where the radish is."
[0364] System: "No problem, the radish is three meters in front of you and to your right."
[0365] Prompt Sentence Examples
[0366] Analyze the image at https: / / your-bucket-name.s3.amazonaws.com / image.jpg and find the location of the radish.
[0367] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0368] Step 1:
[0369] A user takes a picture of a specific area in a commercial facility using the camera on their smartphone.
[0370] Input: Image of a specific area within a commercial facility
[0371] Output: Captured image data
[0372] Specific operation: The user launches the smartphone's camera app and takes a picture of a specific area. The captured image is saved in the smartphone.
[0373] Step 2:
[0374] The device uploads the captured image to a cloud storage service.
[0375] Input: Captured image data
[0376] Output: Image URL on cloud storage
[0377] Specific operation: The smartphone uploads image data to a cloud storage service such as AWS S3 and obtains the image URL.
[0378] Step 3:
[0379] The device connects the URL of the uploaded image to the generative AI.
[0380] Input: Image URL on cloud storage
[0381] Output: Send URL to generative AI
[0382] Specific operation: The smartphone sends the image URL it obtained to a generative AI (e.g., OpenAI GPT-4).
[0383] Step 4:
[0384] The server uses generative AI to analyze the location information of specific objects within the image.
[0385] Input: Image URL
[0386] Output: Location of a specific object
[0387] Specific operation: The generative AI analyzes the image URL and generates the location information of a specific object (e.g., radish). The analysis results are output in text format.
[0388] Step 5:
[0389] The device communicates location information from the generative AI to the user via voice.
[0390] Input: Location of a specific object
[0391] Output: Audio feedback
[0392] Specific operation: The smartphone converts the location information received from the generative AI into voice using a speech synthesis engine and conveys it to the user.
[0393] Step 6:
[0394] The device analyzes the tone of the user's voice using an emotion recognition engine.
[0395] Input: User's tone of voice
[0396] Output: Emotion recognition result
[0397] How it works: The smartphone records the user's voice and sends it to an emotion recognition engine such as IBM Watson Tone Analyzer for analysis. The analysis results include the type and intensity of the emotion.
[0398] Step 7:
[0399] The server adjusts the feedback based on the emotion recognition results.
[0400] Input: Emotion recognition results, location information of specific objects
[0401] Output: Regulated Feedback
[0402] What it does: Generative AI takes emotion recognition results into account and adjusts the content and tone of feedback. For example, if the user is feeling frustrated, it generates more polite feedback.
[0403] Step 8:
[0404] The device then audibly conveys the adjusted feedback to the user.
[0405] Input: Calibrated feedback
[0406] Output: Audio feedback
[0407] Specific operation: The smartphone converts the adjusted feedback into voice using a speech synthesis engine and conveys it to the user.
[0408] Example 2
[0409] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0410] It is difficult for visually impaired people to locate specific products in supermarkets and other commercial facilities. Furthermore, since information is not provided in response to the user's emotions, this can increase the user's frustration. To solve this problem, a system is needed that allows visually impaired people to easily locate specific products and provides information in response to the user's emotions.
[0411] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and transmit location information of a specific object to the user, an emotion recognition means for recognizing the user's emotion, and a means for the generative AI to adjust the information based on the emotion recognition means and transmit it again to the user. This enables visually impaired people to easily grasp the location of specific products and makes it possible to provide information according to the user's emotion.
[0412] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area within a commercial facility.
[0413] "Means for uploading captured images" refers to a device or function that allows a user to send captured images to a server.
[0414] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function that generates a URL for the image received by the server and sends that URL to the generative AI.
[0415] "Generative AI" is an artificial intelligence system that analyzes received images and generates location information for specific objects.
[0416] "Means for communicating the location information of a specific object to the user" refers to devices or functions that notify the user of the location information analyzed by generative AI via voice or text.
[0417] The "emotion recognition means for recognizing the user's emotions" refers to a device or function for analyzing and recognizing the emotions of the user from the tone of voice and facial expressions.
[0418] "Means for the generative AI to adjust information based on the emotion recognition means and re-communicate it to the user" refers to devices or functions that allow the generative AI to adjust information based on the user's emotions recognized by the emotion recognition means and re-notify the user.
[0419] This invention is a system that enables visually impaired people to locate specific products in a commercial facility. This system operates by allowing users to take pictures of specific areas in a store using a device such as a smartphone and upload the images to a server. The specific hardware and software configurations, as well as the data processing and calculation methods, are described below.
[0420] Hardware and software used
[0421] Hardware: Smartphones (e.g., iPhone, Android devices)
[0422] Software: Image upload applications, generative AI (e.g., OpenAI's GPT-4), emotion engines (e.g., Affectiva)
[0423] System Operation Overview
[0424] 1. User takes a photo with their smartphone:
[0425] A user starts the camera app on their smartphone and takes a picture of a specific area in a shopping mall (for example, the vegetable section). The user wants to know where the radishes are.
[0426] 2. The device uploads the captured image to the server:
[0427] The device (smartphone) automatically uploads the captured images to the server. Uploading is done through a dedicated application. This application has the function of compressing the images and sending them to the server.
[0428] 3. The server generates a URL for the image and sends it to the generative AI:
[0429] The server saves the received image and generates a URL for the image. The generated URL is sent to the generative AI. The server then sends this URL to the generative AI's API endpoint.
[0430] 4. Generative AI analyzes specific objects in images:
[0431] The generative AI analyzes the image based on the received URL. Specifically, it uses an image recognition algorithm to identify the location of the "daikon radish." For example, it detects that the radish is in the front right of the image.
[0432] 5. Generative AI communicates analysis results to users:
[0433] The generative AI generates the analysis results in natural language and communicates them to the user. For example, it generates a message such as "The radish is in the front right." This message is sent to the device via the server, and the device application notifies the user by voice or text.
[0434] 6. Emotion engine recognizes user emotions:
[0435] If a user says with frustration, "I don't know where the radish is," the device's microphone captures the voice, and the emotion engine analyzes this voice data to recognize the user's emotion.
[0436] 7. Generative AI adjusts information based on emotions and re-communicates it to the user:
[0437] If the emotion engine recognizes the user's frustration, the generative AI will use that information to tailor the message, for example, generating a more polite message like, "Don't worry, the radish is three meters in front of you to your right." This message is also sent via the server to the device and notified to the user.
[0438] Examples and prompts
[0439] Examples:
[0440] A user takes a photo of the vegetable section with their smartphone and says, "I don't know where the radishes are." The device uploads the image to a server and captures audio. The server generates a URL for the image and sends it to a generative AI. The generative AI analyzes the image and generates location information. It adjusts the message based on information from the emotion engine. Finally, the device notifies the user, "Don't worry, the radishes are three meters in front of you to your right."
[0441] Example prompt sentence:
[0442] "Please tell me where the radish is in this image. The user is frustrated."
[0443] This system not only makes it easier for visually impaired people to locate specific products, but also provides services that respond to the user's emotions.
[0444] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0445] Step 1:
[0446] A user takes a photo of a specific area in a store using their smartphone. The user then launches the smartphone's camera app and takes a photo of a specific area in the shopping mall (e.g., the vegetable section). The input is the image taken by the user, and the output is the image data stored in the smartphone.
[0447] Step 2:
[0448] The device uploads the captured images to the server. The device (smartphone) automatically uploads the captured images to the server. The upload is done through a dedicated application. This application has the function of compressing the images and sending them to the server. The input is the image data stored in the smartphone, and the output is the image data uploaded to the server.
[0449] Step 3:
[0450] The server generates a URL for the image and sends it to the generative AI. The server saves the received image and generates a URL for that image. The generated URL is sent to the generative AI. The server sends this URL to the generative AI's API endpoint. The input is the image data uploaded to the server, and the output is the URL of the generated image.
[0451] Step 4:
[0452] The generative AI analyzes a specific object in an image. The generative AI analyzes the image based on the received URL. Specifically, it uses an image recognition algorithm to identify the location of the "radish." For example, it detects that the radish is in the front right of the image. The input is the URL of the image sent to the generative AI, and the output is the location information of the specific object that was analyzed.
[0453] Step 5:
[0454] The generative AI communicates the analysis results to the user. The generative AI generates the analysis results in natural language and communicates them to the user. For example, it generates a message such as "The radish is in the front right." This message is sent to the device via the server, and the device application notifies the user by voice or text. The input is the location information of the specific object that was analyzed, and the output is the message that is communicated to the user.
[0455] Step 6:
[0456] The emotion engine recognizes the user's emotions. If a user says with frustration, "I don't know where the radish is," the device's microphone captures the voice. The emotion engine analyzes this voice data and recognizes the user's emotions. The input is the user's voice data, and the output is the recognized user's emotional information.
[0457] Step 7:
[0458] The generative AI adjusts the information based on the emotion and transmits it back to the user. If the emotion engine recognizes the user's frustration, the generative AI adjusts the message based on that information. For example, it generates a more polite message such as, "Don't worry, the radish is three meters in front of you to your right." This message is also sent via the server to the device and notified to the user. The input is the recognized user's emotional information, and the output is the adjusted message.
[0459] (Application example 2)
[0460] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0461] It is extremely difficult for visually impaired people to find specific products in commercial facilities. Conventional systems do not provide sufficient information to help visually impaired people find specific products, and do not provide feedback based on the user's emotions. Therefore, there is a need to reduce the stress and frustration that visually impaired people experience when trying to find products.
[0462] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0463] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and transmit location information of a specific object to the user, a means for analyzing the user's voice and recognizing emotions, and a means for adjusting and transmitting the location information of the specific object based on the recognized emotion. This enables visually impaired people to receive appropriate feedback according to their emotions when finding a specific product in a commercial facility.
[0464] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area within a commercial facility.
[0465] The "means for uploading captured images" refers to a device or function for transmitting captured images to a server via the Internet.
[0466] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function for providing the internet address of an uploaded image to the generative AI.
[0467] "Generative AI" is an artificial intelligence system that analyzes uploaded images and extracts location information for specific objects.
[0468] "Means for communicating the location information of a specific object to the user" refers to devices or functions that inform the user of the location information analyzed by generative AI via voice or text.
[0469] "Means for analyzing the user's voice and recognizing emotions" refers to a device or function for analyzing the tone and content of the user's voice to determine their emotional state.
[0470] "Means for adjusting and transmitting location information of a specific object based on recognized emotions" refers to a device or function for transmitting location information of a specific object in an appropriate manner depending on the user's emotional state.
[0471] A system for implementing this invention is one that makes it easier for visually impaired people to find specific products in commercial facilities. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and transmit location information of the specific object to the user, a means for analyzing the user's voice and recognizing emotions, and a means for adjusting and transmitting the location information of the specific object based on the recognized emotion.
[0472] Hardware and software used
[0473] Hardware: Smartphone (camera, microphone)
[0474] Software: OpenCV (image capture), requests (image upload and analysis), gTTS (voice feedback), speech_recognition (voice recognition), emotion_recognition (emotion recognition)
[0475] Data processing and calculation
[0476] Image capture and upload
[0477] Users can take photos of specific areas within a commercial facility using their smartphone camera. The images are automatically uploaded to the cloud, and the URLs of the uploaded images are linked to the generative AI.
[0478] Image analysis
[0479] The generative AI on the server analyzes the uploaded image and extracts the location information of a specific object (e.g., a product). This location information is communicated to the user in the form of "in front," "to the right, in front," "not in front," etc.
[0480] Audio Feedback
[0481] The location information analyzed by the generative AI is provided to the user as voice feedback, which is generated using gTTS and played through the smartphone speaker.
[0482] emotion recognition
[0483] The smartphone's microphone is used to analyze the user's voice and recognize emotions. The voice data is converted to text using speech_recognition, and then emotion_recognition is used to analyze emotions.
[0484] Emotion-based feedback regulation
[0485] The generative AI will adjust and communicate the location of a particular object based on the perceived emotion: for example, if the user is feeling frustrated, the generative AI will communicate the location more politely.
[0486] Specific examples
[0487] If a user says, "I don't know where the radish is," the app will recognize frustration from the tone of voice and respond with a voice prompt: "Don't worry, the radish is three meters in front of you to your right."
[0488] Prompt Sentence Examples
[0489] If a user takes a photo of the store with their smartphone camera and says, "I don't know where the radish is," the app analyzes the image, recognizes frustration from the user's tone of voice, and provides a voice prompt saying, "Don't worry, the radish is three meters in front of you to your right."
[0490] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0491] Step 1:
[0492] A user takes a picture of a specific area in a commercial facility using the camera on their smartphone. The input is the image data captured by the camera, and the output is an image file stored on the smartphone.
[0493] Step 2:
[0494] The device uploads the captured image to the cloud server. The input is the image file acquired in step 1, and the output is the image file saved on the cloud server and its URL.
[0495] Step 3:
[0496] The server connects the URL of the uploaded image to the generative AI. The input is the URL of the image file on the cloud server, and the output is the URL passed to the generative AI.
[0497] Step 4:
[0498] The generative AI performs image recognition and extracts the location information of a specific object. The input is the URL of the image linked in step 3, and the output is the location information of the specific object (e.g., "In front, to the right, in front, not in front").
[0499] Step 5:
[0500] The server provides the user with the location information obtained from the generative AI as voice feedback. The input is the location information obtained in step 4, and the output is voice data that is transmitted to the user as voice feedback.
[0501] Step 6:
[0502] The device records the user's voice and uploads it to the cloud server. The input is the user's voice data, and the output is an audio file stored on the cloud server.
[0503] Step 7:
[0504] The server analyzes the voice data and recognizes the user's emotions. The input is the voice file uploaded in step 6, and the output is the recognized emotion information (e.g., frustration, joy, etc.).
[0505] Step 8:
[0506] The server adjusts and transmits the location information of a specific object based on the recognized emotion. The input is the emotion information recognized in step 7 and the location information obtained in step 4, and the output is the adjusted location information (e.g., "It's okay, the radish is 3 meters in front of you to your right").
[0507] Step 9:
[0508] The server provides the adjusted location information to the user as voice feedback. The input is the adjusted location information from step 8, and the output is the voice data that is transmitted to the user as voice feedback.
[0509] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0510] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0511] Another example of generative AI is Gemini (registered trademark) (Internet search engine). <url: https: gemini.google.com ?hl="ja">) are listed.
[0512] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0513] [Second embodiment]
[0514] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0515] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0516] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0517] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0518] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0519] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0520] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0521] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0522] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0523] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0524] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0525] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.
[0526] "Example 1"
[0527] One embodiment of the present invention provides a system that enables visually impaired people to determine the location of specific products in a supermarket or other store. In this system, a user takes a photo of a specific area in the store using a device such as a smartphone. The captured image is automatically uploaded by the system, and a URL for the image is generated. The generated URL is linked to a generative AI. The generative AI analyzes the location information of a specific object in the image, such as a "daikon radish," and communicates this information to the user in the form of "in front," "to the right, in front," or "not in front." This enables visually impaired people to determine the location of specific products.
[0528] "Example 2"
[0529] One embodiment of the present invention provides a system that enables visually impaired people to determine the location of specific products in a supermarket or other store. In this system, a user takes a photo of a specific area in the store using a device such as a smartphone. The captured image is automatically uploaded by the system, and a URL for the image is generated. The generated URL is linked to a generative AI. The generative AI analyzes the location information of a specific object in the image, such as a "daikon radish," and communicates this information to the user in the form of "in front," "to the right, in front," or "not in front." This enables visually impaired people to determine the location of specific products.
[0530] The processing flow of each embodiment will be described below.
[0531] "Example 1"
[0532] Step 1: A visually impaired person uses a device such as a smartphone to take a photo of a specific area inside a store such as a supermarket.
[0533] Step 2: The system will automatically upload the captured image and generate a URL for it.
[0534] Step 3: The generated URL is linked to the generative AI.
[0535] Step 4: The generative AI analyzes the location information of a specific object in the image, such as a radish.
[0536] Step 5: The generative AI communicates the location of the specific object to the user in the form of "in front of you, to the right, in front of you, not in front of you", etc.
[0537] "Example 2"
[0538] Step 1: A visually impaired person uses a device such as a smartphone to take a photo of a specific area inside a store such as a supermarket.
[0539] Step 2: The system will automatically upload the captured image and generate a URL for it.
[0540] Step 3: The generated URL is linked to the generative AI.
[0541] Step 4: The generative AI analyzes the location information of a specific object in the image, such as a radish.
[0542] Step 5: The generative AI communicates the location of the specific object to the user in the form of "in front of you, to the right, in front of you, not in front of you", etc.
[0543] Example 1
[0544] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0545] It is extremely difficult for visually impaired people to locate specific products in supermarkets and other commercial facilities. Conventional methods require the visually impaired to rely on others for assistance, making it difficult for them to shop independently. Furthermore, existing technologies lack the accuracy and real-time capabilities of image recognition, and are unable to provide sufficient support for visually impaired people.
[0546] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0547] In this invention, the server includes a means for a user to take a picture of a specific area using a mobile device, a means for uploading the taken image to the server, a means for generating a URL for the uploaded image, a means for linking the generated URL to a generative AI model, a means for the generative AI model to analyze location information of a specific object in the image, and a means for communicating the analyzed location information to the user by voice, thereby enabling visually impaired people to independently determine the location of a specific product.
[0548] "User" refers to any user, including visually impaired people, who use the system to locate a particular product.
[0549] "Mobile device" refers to a portable electronic device such as a smartphone or tablet.
[0550] "Specific area" refers to a specific location or section within a commercial establishment.
[0551] "Means for taking a photograph" refers to a method for capturing an image using the camera function built into the mobile terminal.
[0552] "Server" refers to a computer system for storing, processing, and managing data.
[0553] "Means for uploading" refers to the method for transmitting data from the mobile device to the server.
[0554] "URL" refers to an address used to identify a resource on the Internet.
[0555] "Means of generating" refers to the method of creating a URL that indicates the storage location of the uploaded image.
[0556] A "generative AI model" refers to an artificial intelligence algorithm used for image recognition and data analysis.
[0557] "Means of collaboration" refers to the method of sending data from the server to the generative AI model.
[0558] "Means of analysis" refers to how the generative AI model identifies the location information of specific objects within an image.
[0559] "Means of communicating by voice" refers to a method of communicating the analysis results to the user by voice.
[0560] This invention is a system that enables visually impaired people to locate specific products in a commercial facility such as a supermarket. The system works by having the user take a picture of a specific area using a mobile device and uploading the image to a server.
[0561] Hardware and software used
[0562] Hardware:
[0563] Mobile devices (e.g. smartphones, tablets)
[0564] Server (e.g. cloud server, on-premise server)
[0565] software:
[0566] Mobile application for uploading images
[0567] Image management systems (e.g., cloud storage services)
[0568] API integration system (e.g. REST API)
[0569] Generative AI models (e.g., image recognition algorithms)
[0570] Voice assistants (e.g. text-to-speech software)
[0571] System Operation
[0572] The user takes a photo with their mobile device
[0573] Users can open the camera app on their mobile device and take a picture of a specific area in a shopping mall. For example, if a user is looking for daikon radishes in the vegetable section, they can take a picture of the entire vegetable section.
[0574] The device uploads the image to the server.
[0575] The captured images are automatically uploaded from the mobile device to the server, and a notification is displayed when the upload is complete.
[0576] The server generates the image URL
[0577] The server generates a URL for the uploaded image, which is then stored in the database.
[0578] The server links the URL to the generated AI model
[0579] The server connects the generated URL to the AI model, and the URL is sent via API.
[0580] A generative AI model analyzes the image
[0581] The generative AI model analyzes the location of specific objects (e.g., radishes) in the image, and the analysis results are output in text format.
[0582] Generative AI model communicates location information to user
[0583] Based on the analysis results, the generative AI model communicates the location information to the user via voice, and the voice assistant is activated to read the information aloud.
[0584] Specific examples
[0585] Let's say a user is looking for a "daikon radish" in the vegetable section of a supermarket. The user takes a photo of the entire vegetable section with their smartphone and uploads the image to the system. The system generates a URL for the image and connects it to the generative AI model. The generative AI model analyzes the image and tells the user by voice, "The daikon radish is in the front right."
[0586] Prompt Sentence Examples
[0587] "Please tell me the location of the radish in this image. Please tell the user in the form of 'in front, to the right, in front, not in front'."
[0588] This system allows visually impaired people to independently locate specific products.
[0589] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0590] Step 1:
[0591] Users take photos of the store interior using their mobile devices
[0592] Input: A user opens the camera app on their mobile device and takes a picture of a specific area.
[0593] Specific Actions: The user launches the camera app on their mobile device and presses the camera button to capture an image.
[0594] Output: The captured image data is saved on the mobile device.
[0595] Step 2:
[0596] The device uploads the image to the server.
[0597] Input: Captured image data
[0598] What it does: Your device will automatically upload the image to the server, and once the upload is complete, a notification will appear on your device.
[0599] Output: Image data is saved on the server.
[0600] Step 3:
[0601] The server generates the image URL
[0602] Input: Image data stored on the server
[0603] Specific operation: The server generates a URL indicating the location where the image is saved and saves that URL in the database.
[0604] Output: URL of the generated image (e.g. https: / / example.com / image123.jpg)
[0605] Step 4:
[0606] The server links the URL to the generated AI model
[0607] Input: URL of the generated image
[0608] Specific operation: The server sends the image URL to the generative AI model via API, and records the completion of the integration in a log.
[0609] Output: The image URL is sent to the generative AI model.
[0610] Step 5:
[0611] A generative AI model analyzes the image
[0612] Input: The URL of the image sent to the generative AI model
[0613] How it works: The generative AI model downloads an image and analyzes the location of a specific object (e.g., a radish). The analysis results are output in text format.
[0614] Output: Analysis result (e.g. "The radish is in the front right").
[0615] Step 6:
[0616] Generative AI model communicates location information to user
[0617] Input: Text data of analysis results
[0618] How it works: The generative AI model sends the analysis results to the voice assistant, which then reads the information aloud.
[0619] Output: The location information is spoken to the user (e.g., "The radish is in front of you to the right").
[0620] (Application example 1)
[0621] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0622] It is very difficult for visually impaired people to find specific products in commercial facilities such as supermarkets. With conventional methods, visually impaired people need the help of others, making it difficult for them to shop independently. To solve this problem, a system that allows visually impaired people to find specific products by themselves is needed.
[0623] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and communicate location information of the specific object to the user, and a means for providing audio feedback of the location information. This enables visually impaired people to find specific products in commercial facilities using their smartphones.
[0624] An "image capture means" is a device that allows a user to capture an image of a particular area.
[0625] "Means for uploading captured images" is a function for sending captured images to a cloud or server.
[0626] "Means for linking the URL of uploaded images to generative AI" is a function for providing the URL of uploaded images to generative AI.
[0627] "Generative AI" is artificial intelligence that performs image recognition and analyzes the location information of specific objects.
[0628] "Means of communicating the location information of specific objects to users" is a function that informs users of the location information analyzed by generative AI.
[0629] "Means for providing audio feedback of location information" is a function for conveying analyzed location information to the user via audio.
[0630] A "commercial facility" is a place where consumers can purchase goods, such as a supermarket or department store.
[0631] A system for implementing this invention is for assisting visually impaired people in finding specific products in commercial facilities. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to recognize the image and communicate the location information of the specific object to the user, and a means for providing audio feedback of the location information.
[0632] Users use their smartphone camera to take pictures of specific areas within a commercial facility. The images are then uploaded to the cloud via a smartphone application. A URL is automatically generated for the uploaded image, and this URL is linked to the generative AI.
[0633] The generative AI performs image recognition based on the provided image URL and analyzes the location information of a specific object, such as a "daikon radish." As a result of the analysis, the generative AI generates location information in the form of "in front," "to the right, in front," "not in front," etc. This location information is fed back to the user via audio via a smartphone application.
[0634] Specifically, the system works as follows:
[0635] 1. Hardware: Use your smartphone camera to capture the image.
[0636] 2. Software: Use OpenCV to capture images and requests library to upload images to cloud.
[0637] 3. Data processing: Once the image is uploaded to the cloud, a URL is generated.
[0638] 4. Data calculation: Send the image URL and item name to the generative AI, which analyzes the location information of the specific product.
[0639] 5. Feedback: The acquired location information is converted into audio using gTTS (Google Text-to-Speech) and provided as feedback to the user.
[0640] For example, if the user is looking for "daikon radish," the prompt text might look like this:
[0641] Example prompt sentence:
[0642] Could you please tell me the location of the radish in the image?
[0643] By sending this prompt to the generative AI, the AI will analyze the location of the radish in the image and return location information in the form of "in front, to the right, in front, not in front," etc. This allows visually impaired people to find specific products on their own.
[0644] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0645] Step 1:
[0646] A user takes a picture of a specific area in a commercial facility using the camera on their smartphone. The input is the image captured by the user through the camera. The output is an image file stored on the smartphone.
[0647] Step 2:
[0648] The device uploads the captured image to the cloud. The input is the image file obtained in step 1. The output is the URL of the image stored on the cloud. Specifically, the device uses the requests library to send the image file to the cloud server.
[0649] Step 3:
[0650] The device connects the URL of the uploaded image to the generative AI. The input is the URL of the image generated in step 2. The output is the URL and prompt sent to the generative AI. Specifically, the device sends the generative AI a prompt saying, "Please tell me the location of a specific object in the image."
[0651] Step 4:
[0652] The server uses generative AI to perform image recognition and analyze the location information of specific objects. The input is the URL of the image sent in step 3 and the prompt text. The output is the location information of the specific object. Specifically, the generative AI analyzes the specific object in the image and generates location information such as "in front," "to the right and in front," or "not in front."
[0653] Step 5:
[0654] The device provides voice feedback of the location information obtained from the generative AI. The input is the location information obtained in step 4. The output is voice feedback provided to the user. Specifically, the device converts the location information into voice using gTTS (Google Text-to-Speech) and transmits it to the user through the smartphone speaker.
[0655] These steps allow a visually impaired person to locate a particular product within a commercial establishment.
[0656] Example 2
[0657] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0658] It is difficult for visually impaired people to locate specific products in commercial facilities such as supermarkets. With conventional methods, it is difficult for visually impaired people to find products on their own and they need to get help from others. To solve this problem, a system that allows visually impaired people to locate specific products on their own is needed.
[0659] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0660] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI model, a means for the generative AI model to perform image recognition and communicate location information of a specific object to the user, a means for transmitting the generated location information to the user's terminal, and a means for the user to check the location information on the terminal. This enables visually impaired people to independently determine the location of specific products.
[0661] An "image capturing means" is a device that allows a user to capture an image of a specific area within a commercial facility.
[0662] "Means for uploading captured images" is a function for sending image data from the user's terminal to the server.
[0663] "Means for linking the URL of uploaded images to the generative AI model" is a function that generates a URL for the image data received by the server and provides that URL to the generative AI model.
[0664] A "generative AI model" is an artificial intelligence model that uses image recognition technology to analyze the location information of specific objects within a provided image.
[0665] "Means of communicating location information of specific objects to users" is a function for communicating location information analyzed by the generative AI model to users.
[0666] "Means for transmitting the generated location information to the user's terminal" refers to a function that enables the server to transmit the location information received from the generating AI model to the user's terminal.
[0667] "Means for users to check location information on their devices" refers to a function that allows users to check received location information by voice or text using their own devices.
[0668] This invention is a system that enables visually impaired people to determine the location of specific products in a commercial facility. The system allows users to take pictures of specific areas in a store using a device such as a smartphone, and analyzes the images to provide location information for specific objects.
[0669] A user launches the camera app on their smartphone and takes a picture of a specific area in a shopping mall. For example, if the user is looking for a "daikon radish," they take a picture of the vegetable section. The captured image is automatically uploaded to the server through the smartphone's application. At this time, the device uses an Internet connection to send the image data to the server. Specifically, the image data is sent using an HTTP POST request.
[0670] The server stores the received image data and generates a URL for that image. The generated URL is linked to the generative AI model. Specifically, an API request is sent to the generative AI model, providing a prompt message including the image URL. The generative AI model analyzes the image based on the provided image URL. For example, it uses Google Cloud Vision API or Amazon Rekognition to identify the location of the "daikon radish" in the image. As a result of the analysis, location information such as "The daikon radish is in the front right" is generated.
[0671] The server sends the location information received from the generative AI model to the user's smartphone. Specifically, it sends data containing the location information as an HTTP response. The user then checks the received location information through an application on their smartphone. The application then communicates the location information to the user via voice or text. For example, it may provide a voice prompt saying, "The radish is in front of you to the right."
[0672] As a concrete example, consider the case where a user is looking for a "daikon radish" in a supermarket. The user takes a photo of the area inside the store with their smartphone and uploads the image to the system. The server generates a URL for the image and connects it to the generative AI model. The generative AI model analyzes the image and generates location information such as "The daikon radish is in the front right." The server sends this information to the user's smartphone, and the user confirms the location information by voice.
[0673] Example prompt sentence:
[0674] "Please tell me the location of the radish in this image."
[0675] This system allows visually impaired people to locate specific products.
[0676] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0677] Step 1:
[0678] The user takes a photo of a specific area inside the store using their smartphone.
[0679] Specifically, a user launches a camera app on their smartphone and takes a picture of a specific area, such as a vegetable section. The input is the image taken by the user, and the output is the image data stored on the smartphone.
[0680] Step 2:
[0681] The device uploads the captured image to the server.
[0682] Specifically, the device uses an Internet connection to send image data to a server. The input is image data stored on the smartphone, and the output is image data uploaded to the server. The image data is sent using an HTTP POST request.
[0683] Step 3:
[0684] The server generates a URL for the image and connects it to the generative AI model.
[0685] Specifically, the server saves the received image data and generates a URL for the image. The input is the image data uploaded to the server, and the output is the generated image URL. The generated URL sends an API request to the generative AI model and provides a prompt containing the image URL.
[0686] Step 4:
[0687] A generative AI model analyzes the image and generates location information for specific objects.
[0688] Specifically, the generative AI model analyzes an image based on the provided image URL. The input is the image URL and a prompt, and the output is the location information of a specific object. For example, the location of a "daikon radish" in the image can be identified using the Google Cloud Vision API or Amazon Rekognition. The analysis results in location information such as "The daikon radish is in the front right."
[0689] Step 5:
[0690] The server sends the generated location information to the user's smartphone.
[0691] Specifically, the server sends the location information received from the generative AI model to the user's smartphone. The input is the location information from the generative AI model, and the output is the location information sent to the user's smartphone. Data including the location information is sent as an HTTP response.
[0692] Step 6:
[0693] The user checks the location on their smartphone.
[0694] Specifically, the user checks the location information received through a smartphone application. The input is the location information sent from the server, and the output is the location information checked by the user. The application then communicates the location information to the user via voice or text. For example, it may provide a voice prompt saying, "The radish is in front of you to the right."
[0695] (Application example 2)
[0696] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0697] It is difficult for visually impaired people to locate specific products in stores such as supermarkets. Conventional methods have made it difficult for visually impaired people to find products on their own and have required the help of others. For this reason, there is a demand for a system that allows visually impaired people to shop independently.
[0698] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0699] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and communicate location information of a specific object to the user, a means for communicating the analysis result to the user by voice, a voice synthesis means for generating voice, and a voice playback means for playing voice, thereby enabling visually impaired people to locate specific products and shop independently.
[0700] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area inside the store.
[0701] "Means for uploading captured images" refers to devices or functions for sending captured images to a cloud or server.
[0702] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function for passing the URL of an uploaded image to the generative AI.
[0703] "Generative AI" is artificial intelligence that performs image recognition and analyzes the location information of specific objects.
[0704] "Means for communicating location information of a specific object to a user" refers to a device or function for informing a user of analyzed location information.
[0705] "Means for communicating analysis results to users via voice" refers to devices or functions that communicate the results of analysis by generative AI to users via voice.
[0706] "Speech synthesis means for generating speech" refers to a device or function for converting text information into speech.
[0707] The "audio playback means for playing back audio" refers to a device or function for allowing the user to hear the generated audio.
[0708] A system for carrying out this invention assists visually impaired people in locating specific products in a store such as a supermarket. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and communicate the location information of the specific object to the user, a means for communicating the analysis results to the user by voice, a voice synthesis means for generating voice, and a voice playback means for playing back voice.
[0709] Hardware and software used
[0710] Hardware: Smartphone (camera, microphone, speaker)
[0711] software:
[0712] OpenCV (Image Capture)
[0713] requests (HTTP requests)
[0714] gTTS (Google Text-to-Speech)
[0715] mpg321 (audio playback)
[0716] System Operation
[0717] 1. Image capture:
[0718] The user takes a photo of a specific area in the store using their smartphone camera, and the image is captured using OpenCV and saved locally.
[0719] 2. Image upload:
[0720] Upload the captured image to the cloud by using the requests library to send the image as a POST request to the specified URL.
[0721] 3. Image Analysis:
[0722] The URL of the uploaded image is linked to the generative AI. The generative AI analyzes the location information of specific objects in the image. The analysis results are returned in the form of, for example, "In front, in front to the right, not in front."
[0723] 4. Providing Feedback:
[0724] The analysis results are communicated to the user via voice. gTTS is used to convert the text information into voice and play it back in mpg321.
[0725] Specific examples
[0726] If a user is looking for a "daikon radish," they can take a photo of the inside of the store with their smartphone camera. The image is uploaded to the cloud, and generative AI analyzes the location of the "daikon radish." If the analysis returns "It's in front of you on the right," the smartphone will relay that information to the user via voice.
[0727] Prompt Sentence Examples
[0728] Analyze the location of the "radish" in the image. Return the result in the form of "In front, in front right, not in front", etc.
[0729] The system allows visually impaired people to locate specific products and shop independently.
[0730] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0731] Step 1:
[0732] A user uses a smartphone camera to take a picture of a specific area in a store. The input is real-time video from the camera, and the output is a still image. Specifically, a user launches the camera app, frames a specific area, and presses the shutter button to capture the image.
[0733] Step 2:
[0734] The device saves the captured image to local storage. The input is the image data captured in step 1, and the output is an image file saved in local storage. Specifically, the image is saved in JPEG format using the OpenCV library.
[0735] Step 3:
[0736] Uploads an image saved on the device to a cloud server. The input is an image file saved in local storage, and the output is an image URL on the cloud server. Specifically, the requests library is used to send the image file to the specified URL as a POST request.
[0737] Step 4:
[0738] The server connects the URL of the uploaded image to the generative AI. The input is the image URL on the cloud server, and the output is the URL data passed to the generative AI. Specifically, the server sends the URL to the API endpoint of the generative AI.
[0739] Step 5:
[0740] The generative AI analyzes the location information of a specific object in an image. The input is the image URL, and the output is the location information of the specific object. Specifically, the generative AI uses an image recognition algorithm to analyze the location of a specific object (e.g., a "radish") in the image and returns a result in the form of "In front, in front to the right, not in front of you," etc.
[0741] Step 6:
[0742] The server generates data to communicate the analysis results to the user via voice. The input is location information from the generative AI, and the output is text data for speech synthesis. Specifically, the analysis results are converted into text format and passed to the speech synthesis API.
[0743] Step 7:
[0744] The device generates the audio and plays it back to the user. The input is text data for speech synthesis, and the output is audio that the user hears. Specifically, the gTTS library is used to convert the text to an audio file and play it back in mpg321.
[0745] The above steps enable a visually impaired person to locate specific products and shop independently.
[0746] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0747] "Example 1"
[0748] In one embodiment of the present invention, the generative AI includes an emotion engine that recognizes a user's emotions. This emotion engine recognizes emotions from the user's tone of voice. For example, if a user says, "I don't know where the radish is," with a sense of frustration in their voice, the emotion engine recognizes that frustration. The generative AI then communicates location information for a specific object based on the recognized emotion. Specifically, if the generative AI recognizes that the user is feeling frustrated, it communicates location information more politely, such as, "Don't worry, the radish is three meters in front of you to your right." This makes it possible to provide a service that responds to the user's emotions.
[0749] "Example 2"
[0750] In one embodiment of the present invention, the generative AI includes an emotion engine that recognizes a user's emotions. This emotion engine recognizes emotions from the user's tone of voice. For example, if a user says, "I don't know where the radish is," with a sense of frustration in their voice, the emotion engine recognizes that frustration. The generative AI then communicates location information for a specific object based on the recognized emotion. Specifically, if the generative AI recognizes that the user is feeling frustrated, it communicates location information more politely, such as, "Don't worry, the radish is three meters in front of you to your right." This makes it possible to provide a service that responds to the user's emotions.
[0751] The processing flow of each embodiment will be described below.
[0752] "Example 1"
[0753] Step 1: The user says with frustration in their voice, "I don't know where the radish is."
[0754] Step 2: The emotion engine recognizes frustration from the user's tone of voice. Step 3: Based on the emotion recognized by the generative AI, the location information of the radish is communicated. Specifically, the location information is communicated more politely, such as "Don't worry, the radish is three meters in front of you to your right."
[0755] "Example 2"
[0756] Step 1: The user says with frustration in their voice, "I don't know where the radish is."
[0757] Step 2: The emotion engine recognizes frustration from the user's tone of voice. Step 3: Based on the emotion recognized by the generative AI, the location information of the radish is communicated. Specifically, the location information is communicated more politely, such as "Don't worry, the radish is three meters in front of you to your right."
[0758] Example 1
[0759] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0760] It is difficult for visually impaired people to locate specific products in supermarkets and other commercial facilities. It is also difficult to provide appropriate feedback based on the user's emotions. This often causes stress for visually impaired people when shopping. Therefore, there is a need for a system that allows visually impaired people to easily locate specific products and provides feedback based on the user's emotions.
[0761] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0762] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and communicate location information of a specific object to the user, an emotion engine that recognizes the user's emotions, and a means for providing feedback based on the emotions recognized by the emotion engine. This enables visually impaired people to easily determine the location of specific products and provides appropriate feedback according to the user's emotions.
[0763] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area within a commercial facility.
[0764] "Means for uploading captured images" refers to a device or function that allows a user to send captured images to a server.
[0765] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function that generates a URL for the image received by the server and sends that URL to the generative AI.
[0766] "Generative AI" is an artificial intelligence system that performs image recognition and analyzes the location information of specific objects.
[0767] "Means for communicating the location information of a specific object to the user" refers to devices or functions for notifying the user of the location information analyzed by generative AI.
[0768] An "emotion engine that recognizes user emotions" is a device or function that analyzes the tone of a user's voice and other emotional expressions to recognize the user's emotions.
[0769] "Means for providing feedback based on emotions recognized by the emotion engine" refers to devices or functions that enable the generative AI to provide appropriate feedback in response to the user's emotions recognized by the emotion engine.
[0770] MODE FOR CARRYING OUT THE INVENTION
[0771] This invention is a system that allows visually impaired people to determine the location of specific products in a commercial facility. The system allows users to take pictures of specific areas in a store using a device such as a smartphone, and by analyzing the images, provides location information for specific objects. It also includes a function to recognize the user's emotions and provide feedback according to the emotions.
[0772] Hardware and software used
[0773] 1. Smartphone (device)
[0774] Camera application: Used by users to take pictures of specific areas in the store.
[0775] Upload function: Used to send captured images to the server.
[0776] 2. Server
[0777] Image saving function: Saves received images and generates their URLs.
[0778] URL generation function: Generates a URL for the saved image and connects it to the generative AI.
[0779] 3. Generative AI
[0780] Image recognition software: Uses software such as Google Cloud Vision API or Amazon Rekognition to analyze the location of specific objects in an image.
[0781] Feedback function: Used to notify the user of the analysis results.
[0782] 4. Emotion Engine
[0783] Voice analysis software: Using software such as IBM Watson Tone Analyzer, it recognizes emotions from the tone of a user's voice.
[0784] Emotion-based feedback: Generative AI provides appropriate feedback based on the emotions it recognizes.
[0785] Specific examples
[0786] Consider a scenario in which a user takes a photo of a vegetable section in a shopping mall with their smartphone and says, "I don't know where the radishes are." In this case, the system operates as follows.
[0787] 1. The user launches the camera app on their smartphone and takes a picture of the vegetable section. When the user presses the capture button, the image is saved on the device.
[0788] 2. The device automatically uploads the image to the server using the server endpoint configured in the application (e.g., https: / / example.com / upload).
[0789] 3. The server receives and saves the image, generates a URL for the saved image (e.g., https: / / example.com / images / 12345.jpg), and sends the URL to the generative AI API.
[0790] 4. The generative AI receives the image URL and uses the Google Cloud Vision API to analyze the location of the "radish" in the image. As a result, it obtains location information such as "The radish is in the front right."
[0791] 5. The generative AI notifies the user of the analysis results using the smartphone's voice synthesis function. The user's smartphone will then say aloud, "The radish is in front of you to the right."
[0792] 6. The emotion engine analyzes the user's tone of voice and recognizes frustration. The generative AI then conveys the location more politely, such as "Don't worry, the radish is three meters in front of you to your right."
[0793] Prompt Sentence Examples
[0794] "Users can upload images they take with their smartphones and then analyze and tell us the location of specific items. If users get frustrated, we'll give them a more polite way to find the location."
[0795] This system allows visually impaired people to easily locate specific products within a commercial facility and provides appropriate feedback based on the user's emotions.
[0796] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0797] Program processing flow
[0798] Step 1:
[0799] A user takes a photo of a specific area in a commercial facility using their smartphone.
[0800] Input: The user launches the camera app and presses the capture button.
[0801] Data processing: The smartphone camera captures image data and stores it on the device.
[0802] Output: Captured image data.
[0803] Specific operation: The user launches the camera app on their smartphone and takes a picture of the vegetable section. When the user presses the capture button, the image is saved on the device.
[0804] Step 2:
[0805] The device uploads the captured image to the server.
[0806] Input: Image data stored on the device.
[0807] Data processing: The device sends image data to the server using an HTTP POST request.
[0808] Output: Image data uploaded to the server.
[0809] Specific behavior: The device automatically uploads the image to the server using the server endpoint configured in the application (e.g., https: / / example.com / upload).
[0810] Step 3:
[0811] The server generates a URL for the image and connects it to the generative AI.
[0812] Input: Image data uploaded to the server.
[0813] Data processing: The server saves the image and generates a URL for it. The generated URL is sent to the generative AI API.
[0814] Output: The generated image URL.
[0815] Specific operation: The server receives and saves the image. It generates a URL for the saved image (e.g., https: / / example.com / images / 12345.jpg) and sends that URL to the generative AI API.
[0816] Step 4:
[0817] Generative AI analyzes the location of specific objects within an image.
[0818] Input: Image URL sent to the generative AI.
[0819] Data processing: Generative AI uses image recognition software (e.g., Google Cloud Vision API) to analyze the location of a specific object (e.g., a "daikon radish") in an image.
[0820] Output: Location information of a specific object.
[0821] Specific operation: The generative AI receives the image URL and uses the Google Cloud Vision API to analyze the location of the "radish" in the image. As a result of the analysis, it obtains location information such as "The radish is in the front right."
[0822] Step 5:
[0823] The generative AI communicates the analysis results to the user.
[0824] Input: Location information of a specific object.
[0825] Data processing: The generative AI notifies the user of the analysis results via voice or text message.
[0826] Output: The location information communicated to the user.
[0827] Specific operation: The generative AI notifies the user of the analysis results using the smartphone's voice synthesis function. The user's smartphone will then say aloud, "The radish is in front of you to the right."
[0828] Step 6:
[0829] The emotion engine recognizes the user's emotions, and the generative AI provides feedback based on the emotions.
[0830] Input: The user's tone of voice.
[0831] Data processing: The emotion engine uses voice analysis software (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions. The generative AI provides feedback based on the recognized emotions.
[0832] Output: Emotion-based feedback.
[0833] How it works: The emotion engine analyzes the user's tone of voice and recognizes frustration. The generative AI then conveys the location information more politely, such as, "Don't worry, the radish is three meters in front of you to your right."
[0834] (Application example 1)
[0835] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0836] It is extremely difficult for visually impaired people to locate specific products in commercial facilities. Conventional systems make it difficult for visually impaired people to find products on their own, often causing them great stress when shopping. Furthermore, there is no system that provides feedback based on the user's emotions, making it impossible to reduce user frustration. A new system is needed to solve these issues.
[0837] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0838] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and communicate location information of a specific object to the user, an emotion recognition means for recognizing the user's emotion, and a means for adjusting feedback based on the emotion recognition result. This enables visually impaired people to locate specific products in a commercial facility and further reduces stress when shopping by providing feedback according to the user's emotion.
[0839] An "image capturing means" is a device or function that allows a user to capture an image of a specific area within a commercial facility.
[0840] "Means for uploading captured images" refers to a device or function for transferring captured image data to cloud storage or a server.
[0841] "Means for linking the URL of an uploaded image to a generative AI" refers to a device or function for providing the URL of an uploaded image to a generative AI.
[0842] "Generative AI" is an artificial intelligence system that analyzes uploaded images and generates location information for specific objects.
[0843] "Means for communicating the location information of a specific object to the user" refers to a device or function that notifies the user of the location information analyzed by the generative AI via voice or text.
[0844] The "emotion recognition means for recognizing the user's emotions" is a device or function for analyzing the emotions of the user from the tone of voice and facial expressions.
[0845] The "means for adjusting feedback based on emotion recognition results" refers to a device or function for adjusting the content and tone of feedback provided to the user based on the results of analysis by the emotion recognition means.
[0846] A system for implementing this invention assists visually impaired people in locating specific products within a commercial facility. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to recognize the image and communicate the location information of the specific object to the user, an emotion recognition means for recognizing the user's emotion, and a means for adjusting feedback based on the emotion recognition result.
[0847] Hardware and software used
[0848] Hardware: Smartphone (camera, microphone, speaker)
[0849] Software: Cloud storage services (e.g., AWS S3), generative AI models (e.g., OpenAI GPT-4), emotion recognition engines (e.g., IBM Watson Tone Analyzer)
[0850] System Operation Overview
[0851] 1. Taking and uploading images
[0852] Users use their smartphone cameras to take pictures of specific areas within a commercial facility, and the images are uploaded from the smartphone to a cloud storage service (e.g., AWS S3).
[0853] 2. Image Analysis
[0854] The URL of the uploaded image is linked to a generative AI model (e.g., OpenAI GPT-4), which analyzes the location of specific objects (e.g., radishes) in the image and generates a result.
[0855] 3. Audio Feedback
[0856] The generated location information is then communicated to the user via voice via their smartphone, for example, in the form of a notification such as, "The radish is three meters in front of you and to your right."
[0857] 4. Emotion recognition
[0858] The tone of voice a user uses to speak into their smartphone is analyzed by an emotion recognition engine (e.g., IBM Watson Tone Analyzer), which analyzes the user's emotions and identifies feelings such as frustration or relief.
[0859] 5. Feedback adjustment
[0860] Based on the emotion recognition results, the generative AI will adjust the content and tone of the feedback. For example, if the user is feeling frustrated, it will provide more polite feedback, such as, "Don't worry, the radish is three meters in front of you to your right."
[0861] Specific examples
[0862] User: "I don't know where the radish is."
[0863] System: "No problem, the radish is three meters in front of you and to your right."
[0864] Prompt Sentence Examples
[0865] Analyze the image at https: / / your-bucket-name.s3.amazonaws.com / image.jpg and find the location of the radish.
[0866] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0867] Step 1:
[0868] A user takes a picture of a specific area in a commercial facility using the camera on their smartphone.
[0869] Input: Image of a specific area within a commercial facility
[0870] Output: Captured image data
[0871] Specific operation: The user launches the smartphone's camera app and takes a picture of a specific area. The captured image is saved in the smartphone.
[0872] Step 2:
[0873] The device uploads the captured image to a cloud storage service.
[0874] Input: Captured image data
[0875] Output: Image URL on cloud storage
[0876] Specific operation: The smartphone uploads image data to a cloud storage service such as AWS S3 and obtains the image URL.
[0877] Step 3:
[0878] The device connects the URL of the uploaded image to the generative AI.
[0879] Input: Image URL on cloud storage
[0880] Output: Send URL to generative AI
[0881] Specific operation: The smartphone sends the image URL it obtained to a generative AI (e.g., OpenAI GPT-4).
[0882] Step 4:
[0883] The server uses generative AI to analyze the location information of specific objects within the image.
[0884] Input: Image URL
[0885] Output: Location of a specific object
[0886] Specific operation: The generative AI analyzes the image URL and generates the location information of a specific object (e.g., radish). The analysis results are output in text format.
[0887] Step 5:
[0888] The device communicates location information from the generative AI to the user via voice.
[0889] Input: Location of a specific object
[0890] Output: Audio feedback
[0891] Specific operation: The smartphone converts the location information received from the generative AI into voice using a speech synthesis engine and conveys it to the user.
[0892] Step 6:
[0893] The device analyzes the tone of the user's voice using an emotion recognition engine.
[0894] Input: User's tone of voice
[0895] Output: Emotion recognition result
[0896] How it works: The smartphone records the user's voice and sends it to an emotion recognition engine such as IBM Watson Tone Analyzer for analysis. The analysis results include the type and intensity of the emotion.
[0897] Step 7:
[0898] The server adjusts the feedback based on the emotion recognition results.
[0899] Input: Emotion recognition results, location information of specific objects
[0900] Output: Regulated Feedback
[0901] What it does: Generative AI takes emotion recognition results into account and adjusts the content and tone of feedback. For example, if the user is feeling frustrated, it generates more polite feedback.
[0902] Step 8:
[0903] The device then audibly conveys the adjusted feedback to the user.
[0904] Input: Calibrated feedback
[0905] Output: Audio feedback
[0906] Specific operation: The smartphone converts the adjusted feedback into voice using a speech synthesis engine and conveys it to the user.
[0907] Example 2
[0908] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0909] It is difficult for visually impaired people to locate specific products in supermarkets and other commercial facilities. Furthermore, since information is not provided in response to the user's emotions, this can increase the user's frustration. To solve this problem, a system is needed that allows visually impaired people to easily locate specific products and provides information in response to the user's emotions.
[0910] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and transmit location information of a specific object to the user, an emotion recognition means for recognizing the user's emotion, and a means for the generative AI to adjust the information based on the emotion recognition means and transmit it again to the user. This enables visually impaired people to easily grasp the location of specific products and makes it possible to provide information according to the user's emotion.
[0911] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area within a commercial facility.
[0912] "Means for uploading captured images" refers to a device or function that allows a user to send captured images to a server.
[0913] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function that generates a URL for the image received by the server and sends that URL to the generative AI.
[0914] "Generative AI" is an artificial intelligence system that analyzes received images and generates location information for specific objects.
[0915] "Means for communicating the location information of a specific object to the user" refers to devices or functions that notify the user of the location information analyzed by generative AI via voice or text.
[0916] The "emotion recognition means for recognizing the user's emotions" refers to a device or function for analyzing and recognizing the emotions of the user from the tone of voice and facial expressions.
[0917] "Means for the generative AI to adjust information based on the emotion recognition means and re-communicate it to the user" refers to devices or functions that allow the generative AI to adjust information based on the user's emotions recognized by the emotion recognition means and re-notify the user.
[0918] This invention is a system that enables visually impaired people to locate specific products in a commercial facility. This system operates by allowing users to take pictures of specific areas in a store using a device such as a smartphone and upload the images to a server. The specific hardware and software configurations, as well as the data processing and calculation methods, are described below.
[0919] Hardware and software used
[0920] Hardware: Smartphone (e.g. iPhone, Android device)
[0921] Software: Image upload applications, generative AI (e.g., OpenAI's GPT-4), emotion engines (e.g., Affectiva)
[0922] System Operation Overview
[0923] 1. User takes a photo with their smartphone:
[0924] A user starts the camera app on their smartphone and takes a picture of a specific area in a shopping mall (for example, the vegetable section). The user wants to know where the radishes are.
[0925] 2. The device uploads the captured image to the server:
[0926] The device (smartphone) automatically uploads the captured images to the server. Uploading is done through a dedicated application. This application has the function of compressing the images and sending them to the server.
[0927] 3. The server generates a URL for the image and sends it to the generative AI:
[0928] The server saves the received image and generates a URL for the image. The generated URL is sent to the generative AI. The server then sends this URL to the generative AI's API endpoint.
[0929] 4. Generative AI analyzes specific objects in images:
[0930] The generative AI analyzes the image based on the received URL. Specifically, it uses an image recognition algorithm to identify the location of the "daikon radish." For example, it detects that the radish is in the front right of the image.
[0931] 5. Generative AI communicates analysis results to users:
[0932] The generative AI generates the analysis results in natural language and communicates them to the user. For example, it generates a message such as "The radish is in the front right." This message is sent to the device via the server, and the device application notifies the user by voice or text.
[0933] 6. Emotion engine recognizes user emotions:
[0934] If a user says with frustration, "I don't know where the radish is," the device's microphone captures the voice, and the emotion engine analyzes this voice data to recognize the user's emotion.
[0935] 7. Generative AI adjusts information based on emotions and re-communicates it to the user:
[0936] If the emotion engine recognizes the user's frustration, the generative AI will use that information to tailor the message, for example, generating a more polite message like, "Don't worry, the radish is three meters in front of you to your right." This message is also sent via the server to the device and notified to the user.
[0937] Examples and prompts
[0938] Examples:
[0939] A user takes a photo of the vegetable section with their smartphone and says, "I don't know where the radishes are." The device uploads the image to a server and captures audio. The server generates a URL for the image and sends it to a generative AI. The generative AI analyzes the image and generates location information. It adjusts the message based on information from the emotion engine. Finally, the device notifies the user, "Don't worry, the radishes are three meters in front of you to your right."
[0940] Example prompt sentence:
[0941] "Please tell me where the radish is in this image. The user is frustrated."
[0942] This system not only makes it easier for visually impaired people to locate specific products, but also provides services that respond to the user's emotions.
[0943] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0944] Step 1:
[0945] A user takes a photo of a specific area in a store using their smartphone. The user then launches the smartphone's camera app and takes a photo of a specific area in the shopping mall (e.g., the vegetable section). The input is the image taken by the user, and the output is the image data stored in the smartphone.
[0946] Step 2:
[0947] The device uploads the captured images to the server. The device (smartphone) automatically uploads the captured images to the server. The upload is done through a dedicated application. This application has the function of compressing the images and sending them to the server. The input is the image data stored in the smartphone, and the output is the image data uploaded to the server.
[0948] Step 3:
[0949] The server generates a URL for the image and sends it to the generative AI. The server saves the received image and generates a URL for that image. The generated URL is sent to the generative AI. The server sends this URL to the generative AI's API endpoint. The input is the image data uploaded to the server, and the output is the URL of the generated image.
[0950] Step 4:
[0951] The generative AI analyzes a specific object in an image. The generative AI analyzes the image based on the received URL. Specifically, it uses an image recognition algorithm to identify the location of the "radish." For example, it detects that the radish is in the front right of the image. The input is the URL of the image sent to the generative AI, and the output is the location information of the specific object that was analyzed.
[0952] Step 5:
[0953] The generative AI communicates the analysis results to the user. The generative AI generates the analysis results in natural language and communicates them to the user. For example, it generates a message such as "The radish is in the front right." This message is sent to the device via the server, and the device application notifies the user by voice or text. The input is the location information of the specific object that was analyzed, and the output is the message that is communicated to the user.
[0954] Step 6:
[0955] The emotion engine recognizes the user's emotions. If a user says with frustration, "I don't know where the radish is," the device's microphone captures the voice. The emotion engine analyzes this voice data and recognizes the user's emotions. The input is the user's voice data, and the output is the recognized user's emotional information.
[0956] Step 7:
[0957] The generative AI adjusts the information based on the emotion and transmits it back to the user. If the emotion engine recognizes the user's frustration, the generative AI adjusts the message based on that information. For example, it generates a more polite message such as, "Don't worry, the radish is three meters in front of you to your right." This message is also sent via the server to the device and notified to the user. The input is the recognized user's emotional information, and the output is the adjusted message.
[0958] (Application example 2)
[0959] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0960] It is extremely difficult for visually impaired people to find specific products in commercial facilities. Conventional systems do not provide sufficient information to help visually impaired people find specific products, and do not provide feedback based on the user's emotions. Therefore, there is a need to reduce the stress and frustration that visually impaired people experience when trying to find products.
[0961] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0962] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and transmit location information of a specific object to the user, a means for analyzing the user's voice and recognizing emotions, and a means for adjusting and transmitting the location information of the specific object based on the recognized emotion. This enables visually impaired people to receive appropriate feedback according to their emotions when finding a specific product in a commercial facility.
[0963] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area within a commercial facility.
[0964] The "means for uploading captured images" refers to a device or function for transmitting captured images to a server via the Internet.
[0965] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function for providing the internet address of an uploaded image to the generative AI.
[0966] "Generative AI" is an artificial intelligence system that analyzes uploaded images and extracts location information for specific objects.
[0967] "Means for communicating the location information of a specific object to the user" refers to devices or functions that inform the user of the location information analyzed by generative AI via voice or text.
[0968] "Means for analyzing the user's voice and recognizing emotions" refers to a device or function for analyzing the tone and content of the user's voice to determine their emotional state.
[0969] "Means for adjusting and transmitting location information of a specific object based on recognized emotions" refers to a device or function for transmitting location information of a specific object in an appropriate manner depending on the user's emotional state.
[0970] A system for implementing this invention is one that makes it easier for visually impaired people to find specific products in commercial facilities. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and transmit location information of the specific object to the user, a means for analyzing the user's voice and recognizing emotions, and a means for adjusting and transmitting the location information of the specific object based on the recognized emotion.
[0971] Hardware and software used
[0972] Hardware: Smartphone (camera, microphone)
[0973] Software: OpenCV (image capture), requests (image upload and analysis), gTTS (voice feedback), speech_recognition (voice recognition), emotion_recognition (emotion recognition)
[0974] Data processing and calculation
[0975] Image capture and upload
[0976] Users can take photos of specific areas within a commercial facility using their smartphone camera. The images are automatically uploaded to the cloud, and the URLs of the uploaded images are linked to the generative AI.
[0977] Image analysis
[0978] The generative AI on the server analyzes the uploaded image and extracts the location information of a specific object (e.g., a product). This location information is communicated to the user in the form of "in front," "to the right, in front," "not in front," etc.
[0979] Audio Feedback
[0980] The location information analyzed by the generative AI is provided to the user as voice feedback, which is generated using gTTS and played through the smartphone speaker.
[0981] emotion recognition
[0982] The smartphone's microphone is used to analyze the user's voice and recognize emotions. The voice data is converted to text using speech_recognition, and then emotion_recognition is used to analyze emotions.
[0983] Emotion-based feedback regulation
[0984] The generative AI will adjust and communicate the location of a particular object based on the perceived emotion: for example, if the user is feeling frustrated, the generative AI will communicate the location more politely.
[0985] Specific examples
[0986] If a user says, "I don't know where the radish is," the app will recognize frustration from the tone of voice and respond with a voice prompt: "Don't worry, the radish is three meters in front of you to your right."
[0987] Prompt Sentence Examples
[0988] If a user takes a photo of the store with their smartphone camera and says, "I don't know where the radish is," the app analyzes the image, recognizes frustration from the user's tone of voice, and provides a voice prompt saying, "Don't worry, the radish is three meters in front of you to your right."
[0989] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0990] Step 1:
[0991] A user takes a picture of a specific area in a commercial facility using the camera on their smartphone. The input is the image data captured by the camera, and the output is an image file stored on the smartphone.
[0992] Step 2:
[0993] The device uploads the captured image to the cloud server. The input is the image file acquired in step 1, and the output is the image file saved on the cloud server and its URL.
[0994] Step 3:
[0995] The server connects the URL of the uploaded image to the generative AI. The input is the URL of the image file on the cloud server, and the output is the URL passed to the generative AI.
[0996] Step 4:
[0997] The generative AI performs image recognition and extracts the location information of a specific object. The input is the URL of the image linked in step 3, and the output is the location information of the specific object (e.g., "In front, to the right, in front, not in front").
[0998] Step 5:
[0999] The server provides the user with the location information obtained from the generative AI as voice feedback. The input is the location information obtained in step 4, and the output is voice data that is transmitted to the user as voice feedback.
[1000] Step 6:
[1001] The device records the user's voice and uploads it to the cloud server. The input is the user's voice data, and the output is an audio file stored on the cloud server.
[1002] Step 7:
[1003] The server analyzes the voice data and recognizes the user's emotions. The input is the voice file uploaded in step 6, and the output is the recognized emotion information (e.g., frustration, joy, etc.).
[1004] Step 8:
[1005] The server adjusts and transmits the location information of a specific object based on the recognized emotion. The input is the emotion information recognized in step 7 and the location information obtained in step 4, and the output is the adjusted location information (e.g., "It's okay, the radish is 3 meters in front of you to your right").
[1006] Step 9:
[1007] The server provides the adjusted location information to the user as voice feedback. The input is the adjusted location information from step 8, and the output is the voice data that is transmitted to the user as voice feedback.
[1008] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1009] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1010] Another example of generative AI is Gemini (internet search engine). <url: https: gemini.google.com ?hl="ja">) are listed.
[1011] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1012] [Third embodiment]
[1013] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1014] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1015] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1016] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1017] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1018] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1019] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1020] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1021] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1022] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1023] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1024] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.
[1025] "Example 1"
[1026] One embodiment of the present invention provides a system that enables visually impaired people to determine the location of specific products in a supermarket or other store. In this system, a user takes a photo of a specific area in the store using a device such as a smartphone. The captured image is automatically uploaded by the system, and a URL for the image is generated. The generated URL is linked to a generative AI. The generative AI analyzes the location information of a specific object in the image, such as a "daikon radish," and communicates this information to the user in the form of "in front," "to the right, in front," or "not in front." This enables visually impaired people to determine the location of specific products.
[1027] "Example 2"
[1028] One embodiment of the present invention provides a system that enables visually impaired people to determine the location of specific products in a supermarket or other store. In this system, a user takes a photo of a specific area in the store using a device such as a smartphone. The captured image is automatically uploaded by the system, and a URL for the image is generated. The generated URL is linked to a generative AI. The generative AI analyzes the location information of a specific object in the image, such as a "daikon radish," and communicates this information to the user in the form of "in front," "to the right, in front," or "not in front." This enables visually impaired people to determine the location of specific products.
[1029] The processing flow of each embodiment will be described below.
[1030] "Example 1"
[1031] Step 1: A visually impaired person uses a device such as a smartphone to take a photo of a specific area inside a store such as a supermarket.
[1032] Step 2: The system will automatically upload the captured image and generate a URL for it.
[1033] Step 3: The generated URL is linked to the generative AI.
[1034] Step 4: The generative AI analyzes the location information of a specific object in the image, such as a radish.
[1035] Step 5: The generative AI communicates the location of the specific object to the user in the form of "in front of you, to the right, in front of you, not in front of you", etc.
[1036] "Example 2"
[1037] Step 1: A visually impaired person uses a device such as a smartphone to take a photo of a specific area inside a store such as a supermarket.
[1038] Step 2: The system will automatically upload the captured image and generate a URL for it.
[1039] Step 3: The generated URL is linked to the generative AI.
[1040] Step 4: The generative AI analyzes the location information of a specific object in the image, such as a radish.
[1041] Step 5: The generative AI communicates the location of the specific object to the user in the form of "in front of you, to the right, in front of you, not in front of you", etc.
[1042] Example 1
[1043] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1044] It is extremely difficult for visually impaired people to locate specific products in supermarkets and other commercial facilities. Conventional methods require the visually impaired to rely on others for assistance, making it difficult for them to shop independently. Furthermore, existing technologies lack the accuracy and real-time capabilities of image recognition, and are unable to provide sufficient support for visually impaired people.
[1045] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1046] In this invention, the server includes a means for a user to take a picture of a specific area using a mobile device, a means for uploading the taken image to the server, a means for generating a URL for the uploaded image, a means for linking the generated URL to a generative AI model, a means for the generative AI model to analyze location information of a specific object in the image, and a means for communicating the analyzed location information to the user by voice, thereby enabling visually impaired people to independently determine the location of a specific product.
[1047] "User" refers to any user, including visually impaired people, who use the system to locate a particular product.
[1048] "Mobile device" refers to a portable electronic device such as a smartphone or tablet.
[1049] "Specific area" refers to a specific location or section within a commercial establishment.
[1050] "Means for taking a photograph" refers to a method for capturing an image using the camera function built into the mobile terminal.
[1051] "Server" refers to a computer system for storing, processing, and managing data.
[1052] "Means for uploading" refers to a method for transmitting data from a mobile device to a server.
[1053] "URL" refers to an address used to identify a resource on the Internet.
[1054] "Means of generating" refers to the method of creating a URL that indicates the location where the uploaded image is saved.
[1055] A "generative AI model" refers to an artificial intelligence algorithm used for image recognition and data analysis.
[1056] "Means of collaboration" refers to the method of sending data from the server to the generative AI model.
[1057] "Means of analysis" refers to how the generative AI model identifies the location information of specific objects within an image.
[1058] "Means of communicating by voice" refers to a method of communicating the analysis results to the user by voice.
[1059] This invention is a system that enables visually impaired people to locate specific products in a commercial facility such as a supermarket. The system works by having the user take a picture of a specific area using a mobile device and uploading the image to a server.
[1060] Hardware and software used
[1061] Hardware:
[1062] Mobile devices (e.g. smartphones, tablets)
[1063] Server (e.g. cloud server, on-premise server)
[1064] software:
[1065] Mobile application for uploading images
[1066] Image management systems (e.g., cloud storage services)
[1067] API integration system (e.g. REST API)
[1068] Generative AI models (e.g., image recognition algorithms)
[1069] Voice assistants (e.g. text-to-speech software)
[1070] System Operation
[1071] The user takes a photo with their mobile device
[1072] Users can open the camera app on their mobile device and take a picture of a specific area in a shopping mall. For example, if a user is looking for daikon radishes in the vegetable section, they can take a picture of the entire vegetable section.
[1073] The device uploads the image to the server.
[1074] The captured images are automatically uploaded from the mobile device to the server, and a notification is displayed when the upload is complete.
[1075] The server generates the image URL
[1076] The server generates a URL for the uploaded image, which is then stored in the database.
[1077] The server links the URL to the generated AI model
[1078] The server connects the generated URL to the AI model, and the URL is sent via API.
[1079] A generative AI model analyzes the image
[1080] The generative AI model analyzes the location of specific objects (e.g., radishes) in the image, and the analysis results are output in text format.
[1081] Generative AI model communicates location information to user
[1082] Based on the analysis results, the generative AI model communicates the location information to the user via voice, and the voice assistant is activated to read the information aloud.
[1083] Specific examples
[1084] Let's say a user is looking for a "daikon radish" in the vegetable section of a supermarket. The user takes a photo of the entire vegetable section with their smartphone and uploads the image to the system. The system generates a URL for the image and connects it to the generative AI model. The generative AI model analyzes the image and tells the user by voice, "The daikon radish is in the front right."
[1085] Prompt Sentence Examples
[1086] "Please tell me the location of the radish in this image. Please tell the user in the form of 'in front, to the right, in front, not in front'."
[1087] This system allows visually impaired people to independently locate specific products.
[1088] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1089] Step 1:
[1090] Users take photos of the store interior using their mobile devices
[1091] Input: A user opens the camera app on their mobile device and takes a picture of a specific area.
[1092] Specific Actions: The user launches the camera app on their mobile device and presses the camera button to capture an image.
[1093] Output: The captured image data is saved on the mobile device.
[1094] Step 2:
[1095] The device uploads the image to the server.
[1096] Input: Captured image data
[1097] What it does: Your device will automatically upload the image to the server, and once the upload is complete, a notification will appear on your device.
[1098] Output: Image data is saved on the server.
[1099] Step 3:
[1100] The server generates the image URL
[1101] Input: Image data stored on the server
[1102] Specific operation: The server generates a URL indicating the location where the image is saved and saves that URL in the database.
[1103] Output: URL of the generated image (e.g. https: / / example.com / image123.jpg)
[1104] Step 4:
[1105] The server links the URL to the generated AI model
[1106] Input: URL of the generated image
[1107] Specific operation: The server sends the image URL to the generative AI model via API, and records the completion of the integration in a log.
[1108] Output: The image URL is sent to the generative AI model.
[1109] Step 5:
[1110] A generative AI model analyzes the image
[1111] Input: The URL of the image sent to the generative AI model
[1112] How it works: The generative AI model downloads an image and analyzes the location of a specific object (e.g., a radish). The analysis results are output in text format.
[1113] Output: Analysis result (e.g. "The radish is in the front right").
[1114] Step 6:
[1115] Generative AI model communicates location information to user
[1116] Input: Text data of analysis results
[1117] How it works: The generative AI model sends the analysis results to the voice assistant, which then reads the information aloud.
[1118] Output: The location information is spoken to the user (e.g., "The radish is in front of you to the right").
[1119] (Application example 1)
[1120] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1121] It is very difficult for visually impaired people to find specific products in commercial facilities such as supermarkets. With conventional methods, visually impaired people need the help of others, making it difficult for them to shop independently. To solve this problem, a system that allows visually impaired people to find specific products by themselves is needed.
[1122] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and communicate location information of the specific object to the user, and a means for providing audio feedback of the location information. This enables visually impaired people to find specific products in commercial facilities using their smartphones.
[1123] An "image capture means" is a device that allows a user to capture an image of a particular area.
[1124] "Means for uploading captured images" is a function for sending captured images to a cloud or server.
[1125] "Means for linking the URL of uploaded images to generative AI" is a function for providing the URL of uploaded images to generative AI.
[1126] "Generative AI" is artificial intelligence that performs image recognition and analyzes the location information of specific objects.
[1127] "Means of communicating the location information of specific objects to users" is a function that informs users of the location information analyzed by generative AI.
[1128] "Means for providing audio feedback of location information" is a function for conveying analyzed location information to the user via audio.
[1129] A "commercial facility" is a place where consumers can purchase goods, such as a supermarket or department store.
[1130] A system for implementing this invention is for assisting visually impaired people in finding specific products in commercial facilities. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to recognize the image and communicate the location information of the specific object to the user, and a means for providing audio feedback of the location information.
[1131] Users use their smartphone camera to take pictures of specific areas within a commercial facility. The images are then uploaded to the cloud via a smartphone application. A URL is automatically generated for the uploaded image, and this URL is linked to the generative AI.
[1132] The generative AI performs image recognition based on the provided image URL and analyzes the location information of a specific object, such as a "daikon radish." As a result of the analysis, the generative AI generates location information in the form of "in front," "to the right, in front," "not in front," etc. This location information is fed back to the user via audio via a smartphone application.
[1133] Specifically, the system works as follows:
[1134] 1. Hardware: Use your smartphone camera to capture the image.
[1135] 2. Software: Use OpenCV to capture images and requests library to upload images to cloud.
[1136] 3. Data processing: Once the image is uploaded to the cloud, a URL is generated.
[1137] 4. Data calculation: Send the image URL and item name to the generative AI, which analyzes the location information of the specific product.
[1138] 5. Feedback: The acquired location information is converted into audio using gTTS (Google Text-to-Speech) and provided as feedback to the user.
[1139] For example, if the user is looking for "daikon radish," the prompt text might look like this:
[1140] Example prompt sentence:
[1141] Could you please tell me the location of the radish in the image?
[1142] By sending this prompt to the generative AI, the AI will analyze the location of the radish in the image and return location information in the form of "in front, to the right, in front, not in front," etc. This allows visually impaired people to find specific products on their own.
[1143] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1144] Step 1:
[1145] A user takes a picture of a specific area in a commercial facility using the camera on their smartphone. The input is the image captured by the user through the camera. The output is an image file stored on the smartphone.
[1146] Step 2:
[1147] The device uploads the captured image to the cloud. The input is the image file obtained in step 1. The output is the URL of the image stored on the cloud. Specifically, the device uses the requests library to send the image file to the cloud server.
[1148] Step 3:
[1149] The device connects the URL of the uploaded image to the generative AI. The input is the URL of the image generated in step 2. The output is the URL and prompt sent to the generative AI. Specifically, the device sends the generative AI a prompt saying, "Please tell me the location of a specific object in the image."
[1150] Step 4:
[1151] The server uses generative AI to perform image recognition and analyze the location information of specific objects. The input is the URL of the image sent in step 3 and the prompt text. The output is the location information of the specific object. Specifically, the generative AI analyzes the specific object in the image and generates location information such as "in front," "to the right and in front," or "not in front."
[1152] Step 5:
[1153] The device provides voice feedback of the location information obtained from the generative AI. The input is the location information obtained in step 4. The output is voice feedback provided to the user. Specifically, the device converts the location information into voice using gTTS (Google Text-to-Speech) and transmits it to the user through the smartphone speaker.
[1154] These steps allow a visually impaired person to locate a particular product within a commercial establishment.
[1155] Example 2
[1156] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1157] It is difficult for visually impaired people to locate specific products in commercial facilities such as supermarkets. With conventional methods, it is difficult for visually impaired people to find products on their own and they need to get help from others. To solve this problem, a system that allows visually impaired people to locate specific products on their own is needed.
[1158] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1159] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI model, a means for the generative AI model to perform image recognition and communicate location information of a specific object to the user, a means for transmitting the generated location information to the user's terminal, and a means for the user to check the location information on the terminal. This enables visually impaired people to independently determine the location of specific products.
[1160] An "image capturing means" is a device that allows a user to capture an image of a specific area within a commercial facility.
[1161] "Means for uploading captured images" is a function for sending image data from the user's terminal to the server.
[1162] "Means for linking the URL of uploaded images to the generative AI model" is a function that generates a URL for the image data received by the server and provides that URL to the generative AI model.
[1163] A "generative AI model" is an artificial intelligence model that uses image recognition technology to analyze the location information of specific objects within a provided image.
[1164] "Means of communicating location information of specific objects to users" is a function for communicating location information analyzed by the generative AI model to users.
[1165] "Means for transmitting the generated location information to the user's terminal" refers to a function that enables the server to transmit the location information received from the generating AI model to the user's terminal.
[1166] "Means for users to check location information on their devices" refers to a function that allows users to check received location information by voice or text using their own devices.
[1167] This invention is a system that enables visually impaired people to determine the location of specific products in a commercial facility. The system allows users to take a photo of a specific area in a store using a device such as a smartphone, and analyzes the image to provide location information for the specific object.
[1168] A user launches the camera app on their smartphone and takes a picture of a specific area in a shopping mall. For example, if the user is looking for a "daikon radish," they take a picture of the vegetable section. The captured image is automatically uploaded to the server through the smartphone's application. At this time, the device uses an Internet connection to send the image data to the server. Specifically, the image data is sent using an HTTP POST request.
[1169] The server stores the received image data and generates a URL for that image. The generated URL is linked to the generative AI model. Specifically, an API request is sent to the generative AI model, providing a prompt message including the image URL. The generative AI model analyzes the image based on the provided image URL. For example, it uses Google Cloud Vision API or Amazon Rekognition to identify the location of the "daikon radish" in the image. As a result of the analysis, location information such as "The daikon radish is in the front right" is generated.
[1170] The server sends the location information received from the generative AI model to the user's smartphone. Specifically, it sends data containing the location information as an HTTP response. The user then checks the received location information through an application on their smartphone. The application then communicates the location information to the user via voice or text. For example, it may provide a voice prompt saying, "The radish is in front of you to the right."
[1171] As a concrete example, consider the case where a user is looking for a "daikon radish" in a supermarket. The user takes a photo of the area inside the store with their smartphone and uploads the image to the system. The server generates a URL for the image and connects it to the generative AI model. The generative AI model analyzes the image and generates location information such as "The daikon radish is in the front right." The server sends this information to the user's smartphone, and the user confirms the location information by voice.
[1172] Example prompt sentence:
[1173] "Please tell me the location of the radish in this image."
[1174] This system allows visually impaired people to locate specific products.
[1175] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1176] Step 1:
[1177] The user takes a photo of a specific area inside the store using their smartphone.
[1178] Specifically, a user launches a camera app on their smartphone and takes a picture of a specific area, such as a vegetable section. The input is the image taken by the user, and the output is the image data stored on the smartphone.
[1179] Step 2:
[1180] The device uploads the captured image to the server.
[1181] Specifically, the device uses an Internet connection to send image data to a server. The input is image data stored on the smartphone, and the output is image data uploaded to the server. The image data is sent using an HTTP POST request.
[1182] Step 3:
[1183] The server generates a URL for the image and connects it to the generative AI model.
[1184] Specifically, the server saves the received image data and generates a URL for the image. The input is the image data uploaded to the server, and the output is the generated image URL. The generated URL sends an API request to the generative AI model and provides a prompt containing the image URL.
[1185] Step 4:
[1186] A generative AI model analyzes the image and generates location information for specific objects.
[1187] Specifically, the generative AI model analyzes an image based on the provided image URL. The input is the image URL and a prompt, and the output is the location information of a specific object. For example, the location of a "daikon radish" in the image can be identified using the Google Cloud Vision API or Amazon Rekognition. The analysis results in location information such as "The daikon radish is in the front right."
[1188] Step 5:
[1189] The server sends the generated location information to the user's smartphone.
[1190] Specifically, the server sends the location information received from the generative AI model to the user's smartphone. The input is the location information from the generative AI model, and the output is the location information sent to the user's smartphone. Data including the location information is sent as an HTTP response.
[1191] Step 6:
[1192] The user checks the location on their smartphone.
[1193] Specifically, the user checks the location information received through a smartphone application. The input is the location information sent from the server, and the output is the location information checked by the user. The application then communicates the location information to the user via voice or text. For example, it may provide a voice prompt saying, "The radish is in front of you to the right."
[1194] (Application example 2)
[1195] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1196] It is difficult for visually impaired people to locate specific products in stores such as supermarkets. With conventional methods, it is difficult for visually impaired people to find products on their own and they have to ask for help from others. For this reason, there is a need for a system that allows visually impaired people to shop independently.
[1197] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1198] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and communicate location information of a specific object to the user, a means for communicating the analysis result to the user by voice, a voice synthesis means for generating voice, and a voice playback means for playing voice, thereby enabling visually impaired people to locate specific products and shop independently.
[1199] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area inside the store.
[1200] "Means for uploading captured images" refers to devices or functions for sending captured images to a cloud or server.
[1201] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function for passing the URL of an uploaded image to the generative AI.
[1202] "Generative AI" is artificial intelligence that performs image recognition and analyzes the location information of specific objects.
[1203] "Means for communicating location information of a specific object to a user" refers to a device or function for informing a user of analyzed location information.
[1204] "Means for communicating analysis results to users via voice" refers to devices or functions that communicate the results of analysis by generative AI to users via voice.
[1205] "Speech synthesis means for generating speech" refers to a device or function for converting text information into speech.
[1206] The "audio playback means for playing back audio" refers to a device or function for allowing the user to hear the generated audio.
[1207] A system for carrying out this invention assists visually impaired people in locating specific products in a store such as a supermarket. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and communicate the location information of the specific object to the user, a means for communicating the analysis results to the user by voice, a voice synthesis means for generating voice, and a voice playback means for playing back voice.
[1208] Hardware and software used
[1209] Hardware: Smartphone (camera, microphone, speaker)
[1210] software:
[1211] OpenCV (Image Capture)
[1212] requests (HTTP requests)
[1213] gTTS (Google Text-to-Speech)
[1214] mpg321 (audio playback)
[1215] System Operation
[1216] 1. Image capture:
[1217] The user takes a photo of a specific area in the store using their smartphone camera, and the image is captured using OpenCV and saved locally.
[1218] 2. Image upload:
[1219] Upload the captured image to the cloud by using the requests library to send the image as a POST request to the specified URL.
[1220] 3. Image Analysis:
[1221] The URL of the uploaded image is linked to the generative AI. The generative AI analyzes the location information of specific objects in the image. The analysis results are returned in the form of, for example, "In front, in front to the right, not in front."
[1222] 4. Providing Feedback:
[1223] The analysis results are communicated to the user via voice. gTTS is used to convert the text information into voice and play it back in mpg321.
[1224] Specific examples
[1225] If a user is looking for a "daikon radish," they can take a photo of the inside of the store with their smartphone camera. The image is uploaded to the cloud, and generative AI analyzes the location of the "daikon radish." If the analysis returns "It's in front of you on the right," the smartphone will relay that information to the user via voice.
[1226] Prompt Sentence Examples
[1227] Analyze the location of the "radish" in the image. Return the result in the form of "In front, in front right, not in front", etc.
[1228] The system allows visually impaired people to locate specific products and shop independently.
[1229] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1230] Step 1:
[1231] A user uses a smartphone camera to take a picture of a specific area in a store. The input is real-time video from the camera, and the output is a still image. Specifically, a user launches the camera app, frames a specific area, and presses the shutter button to capture the image.
[1232] Step 2:
[1233] The device saves the captured image to local storage. The input is the image data captured in step 1, and the output is an image file saved in local storage. Specifically, the image is saved in JPEG format using the OpenCV library.
[1234] Step 3:
[1235] Uploads an image saved on the device to a cloud server. The input is an image file saved in local storage, and the output is an image URL on the cloud server. Specifically, the requests library is used to send the image file to the specified URL as a POST request.
[1236] Step 4:
[1237] The server connects the URL of the uploaded image to the generative AI. The input is the image URL on the cloud server, and the output is the URL data passed to the generative AI. Specifically, the server sends the URL to the API endpoint of the generative AI.
[1238] Step 5:
[1239] The generative AI analyzes the location information of a specific object in an image. The input is the image URL, and the output is the location information of the specific object. Specifically, the generative AI uses an image recognition algorithm to analyze the location of a specific object (e.g., a "radish") in the image and returns a result in the form of "In front, in front to the right, not in front of you," etc.
[1240] Step 6:
[1241] The server generates data to communicate the analysis results to the user via voice. The input is location information from the generative AI, and the output is text data for speech synthesis. Specifically, the analysis results are converted into text format and passed to the speech synthesis API.
[1242] Step 7:
[1243] The device generates the audio and plays it back to the user. The input is text data for speech synthesis, and the output is audio that the user hears. Specifically, the gTTS library is used to convert the text to an audio file and play it back in mpg321.
[1244] The above steps enable a visually impaired person to locate specific products and shop independently.
[1245] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1246] "Example 1"
[1247] In one embodiment of the present invention, the generative AI includes an emotion engine that recognizes a user's emotions. This emotion engine recognizes emotions from the user's tone of voice. For example, if a user says, "I don't know where the radish is," with a sense of frustration in their voice, the emotion engine recognizes that frustration. The generative AI then communicates location information for a specific object based on the recognized emotion. Specifically, if the generative AI recognizes that the user is feeling frustrated, it communicates location information more politely, such as, "Don't worry, the radish is three meters in front of you to your right." This makes it possible to provide a service that responds to the user's emotions.
[1248] "Example 2"
[1249] In one embodiment of the present invention, the generative AI includes an emotion engine that recognizes a user's emotions. This emotion engine recognizes emotions from the user's tone of voice. For example, if a user says, "I don't know where the radish is," with a sense of frustration in their voice, the emotion engine recognizes that frustration. The generative AI then communicates location information for a specific object based on the recognized emotion. Specifically, if the generative AI recognizes that the user is feeling frustrated, it communicates location information more politely, such as, "Don't worry, the radish is three meters in front of you to your right." This makes it possible to provide a service that responds to the user's emotions.
[1250] The processing flow of each embodiment will be described below.
[1251] "Example 1"
[1252] Step 1: The user says with frustration in their voice, "I don't know where the radish is."
[1253] Step 2: The emotion engine recognizes frustration from the user's tone of voice. Step 3: Based on the emotion recognized by the generative AI, the location information of the radish is communicated. Specifically, the location information is communicated more politely, such as "Don't worry, the radish is three meters in front of you to your right."
[1254] "Example 2"
[1255] Step 1: The user says with frustration in their voice, "I don't know where the radish is."
[1256] Step 2: The emotion engine recognizes frustration from the user's tone of voice. Step 3: Based on the emotion recognized by the generative AI, the location information of the radish is communicated. Specifically, the location information is communicated more politely, such as "Don't worry, the radish is three meters in front of you to your right."
[1257] Example 1
[1258] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1259] It is difficult for visually impaired people to locate specific products in supermarkets and other commercial facilities. It is also difficult to provide appropriate feedback based on the user's emotions. This often causes stress for visually impaired people when shopping. Therefore, there is a need for a system that allows visually impaired people to easily locate specific products and provides feedback based on the user's emotions.
[1260] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1261] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and communicate location information of a specific object to the user, an emotion engine that recognizes the user's emotions, and a means for providing feedback based on the emotions recognized by the emotion engine. This enables visually impaired people to easily determine the location of specific products and provides appropriate feedback according to the user's emotions.
[1262] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area within a commercial facility.
[1263] "Means for uploading captured images" refers to a device or function that allows a user to send captured images to a server.
[1264] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function that generates a URL for the image received by the server and sends that URL to the generative AI.
[1265] "Generative AI" is an artificial intelligence system that performs image recognition and analyzes the location information of specific objects.
[1266] "Means for communicating the location information of a specific object to the user" refers to devices or functions for notifying the user of the location information analyzed by generative AI.
[1267] An "emotion engine that recognizes user emotions" is a device or function that analyzes the tone of a user's voice and other emotional expressions to recognize the user's emotions.
[1268] "Means for providing feedback based on emotions recognized by the emotion engine" refers to devices or functions that enable the generative AI to provide appropriate feedback in response to the user's emotions recognized by the emotion engine.
[1269] MODE FOR CARRYING OUT THE INVENTION
[1270] This invention is a system that allows visually impaired people to determine the location of specific products in a commercial facility. The system allows users to take pictures of specific areas in a store using a device such as a smartphone, and by analyzing the images, provides location information for specific objects. It also includes a function to recognize the user's emotions and provide feedback according to the emotions.
[1271] Hardware and software used
[1272] 1. Smartphone (device)
[1273] Camera application: Used by users to take pictures of specific areas in the store.
[1274] Upload function: Used to send captured images to the server.
[1275] 2. Server
[1276] Image saving function: Saves received images and generates their URLs.
[1277] URL generation function: Generates a URL for the saved image and connects it to the generative AI.
[1278] 3. Generative AI
[1279] Image recognition software: Uses software such as Google Cloud Vision API or Amazon Rekognition to analyze the location of specific objects in an image.
[1280] Feedback function: Used to notify the user of the analysis results.
[1281] 4. Emotion Engine
[1282] Voice analysis software: Using software such as IBM Watson Tone Analyzer, it recognizes emotions from the tone of a user's voice.
[1283] Emotion-based feedback: Generative AI provides appropriate feedback based on the emotions it recognizes.
[1284] Specific examples
[1285] Consider a scenario in which a user takes a photo of a vegetable section in a shopping mall with their smartphone and says, "I don't know where the radishes are." In this case, the system operates as follows.
[1286] 1. The user launches the camera app on their smartphone and takes a picture of the vegetable section. When the user presses the capture button, the image is saved on the device.
[1287] 2. The device automatically uploads the image to the server using the server endpoint configured in the application (e.g., https: / / example.com / upload).
[1288] 3. The server receives and saves the image, generates a URL for the saved image (e.g., https: / / example.com / images / 12345.jpg), and sends the URL to the generative AI API.
[1289] 4. The generative AI receives the image URL and uses the Google Cloud Vision API to analyze the location of the "radish" in the image. As a result, it obtains location information such as "The radish is in the front right."
[1290] 5. The generative AI notifies the user of the analysis results using the smartphone's voice synthesis function. The user's smartphone will then say aloud, "The radish is in front of you to the right."
[1291] 6. The emotion engine analyzes the user's tone of voice and recognizes frustration. The generative AI then conveys the location more politely, such as "Don't worry, the radish is three meters in front of you to your right."
[1292] Prompt Sentence Examples
[1293] "Users can upload images they take with their smartphones and then analyze and tell us the location of specific items. If users get frustrated, we'll give them a more polite way to find the location."
[1294] This system allows visually impaired people to easily locate specific products within a commercial facility and provides appropriate feedback based on the user's emotions.
[1295] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1296] Program processing flow
[1297] Step 1:
[1298] A user takes a photo of a specific area in a commercial facility using their smartphone.
[1299] Input: The user launches the camera app and presses the capture button.
[1300] Data processing: The smartphone camera captures image data and stores it on the device.
[1301] Output: Captured image data.
[1302] Specific operation: The user launches the camera app on their smartphone and takes a picture of the vegetable section. When the user presses the capture button, the image is saved on the device.
[1303] Step 2:
[1304] The device uploads the captured image to the server.
[1305] Input: Image data stored on the device.
[1306] Data processing: The device sends image data to the server using an HTTP POST request.
[1307] Output: Image data uploaded to the server.
[1308] Specific behavior: The device automatically uploads the image to the server using the server endpoint configured in the application (e.g., https: / / example.com / upload).
[1309] Step 3:
[1310] The server generates a URL for the image and connects it to the generative AI.
[1311] Input: Image data uploaded to the server.
[1312] Data processing: The server saves the image and generates a URL for it. The generated URL is sent to the generative AI API.
[1313] Output: The generated image URL.
[1314] Specific operation: The server receives and saves the image. It generates a URL for the saved image (e.g., https: / / example.com / images / 12345.jpg) and sends that URL to the generative AI API.
[1315] Step 4:
[1316] Generative AI analyzes the location of specific objects within an image.
[1317] Input: Image URL sent to the generative AI.
[1318] Data processing: Generative AI uses image recognition software (e.g., Google Cloud Vision API) to analyze the location of a specific object (e.g., a "daikon radish") in an image.
[1319] Output: Location information of a specific object.
[1320] Specific operation: The generative AI receives the image URL and uses the Google Cloud Vision API to analyze the location of the "radish" in the image. As a result of the analysis, it obtains location information such as "The radish is in the front right."
[1321] Step 5:
[1322] The generative AI communicates the analysis results to the user.
[1323] Input: Location information of a specific object.
[1324] Data processing: The generative AI notifies the user of the analysis results via voice or text message.
[1325] Output: The location information communicated to the user.
[1326] Specific operation: The generative AI notifies the user of the analysis results using the smartphone's voice synthesis function. The user's smartphone will then say aloud, "The radish is in front of you to the right."
[1327] Step 6:
[1328] The emotion engine recognizes the user's emotions, and the generative AI provides feedback based on the emotions.
[1329] Input: The user's tone of voice.
[1330] Data processing: The emotion engine uses voice analysis software (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions. The generative AI provides feedback based on the recognized emotions.
[1331] Output: Emotion-based feedback.
[1332] How it works: The emotion engine analyzes the user's tone of voice and recognizes frustration. The generative AI then conveys the location information more politely, such as, "Don't worry, the radish is three meters in front of you to your right."
[1333] (Application example 1)
[1334] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1335] It is extremely difficult for visually impaired people to locate specific products in commercial facilities. Conventional systems make it difficult for visually impaired people to find products on their own, often causing them great stress when shopping. Furthermore, there is no system that provides feedback based on the user's emotions, making it impossible to reduce user frustration. A new system is needed to solve these issues.
[1336] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1337] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and communicate location information of a specific object to the user, an emotion recognition means for recognizing the user's emotion, and a means for adjusting feedback based on the emotion recognition result. This enables visually impaired people to locate specific products in a commercial facility and further reduces stress when shopping by providing feedback according to the user's emotion.
[1338] An "image capturing means" is a device or function that allows a user to capture an image of a specific area within a commercial facility.
[1339] "Means for uploading captured images" refers to a device or function for transferring captured image data to cloud storage or a server.
[1340] "Means for linking the URL of an uploaded image to a generative AI" refers to a device or function for providing the URL of an uploaded image to a generative AI.
[1341] "Generative AI" is an artificial intelligence system that analyzes uploaded images and generates location information for specific objects.
[1342] "Means for communicating the location information of a specific object to the user" refers to a device or function that notifies the user of the location information analyzed by the generative AI via voice or text.
[1343] The "emotion recognition means for recognizing the user's emotions" is a device or function for analyzing the emotions of the user from the tone of voice and facial expressions.
[1344] The "means for adjusting feedback based on emotion recognition results" refers to a device or function for adjusting the content and tone of feedback provided to the user based on the results of analysis by the emotion recognition means.
[1345] A system for implementing this invention assists visually impaired people in locating specific products within a commercial facility. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to recognize the image and communicate the location information of the specific object to the user, an emotion recognition means for recognizing the user's emotion, and a means for adjusting feedback based on the emotion recognition result.
[1346] Hardware and software used
[1347] Hardware: Smartphone (camera, microphone, speaker)
[1348] Software: Cloud storage services (e.g., AWS S3), generative AI models (e.g., OpenAI GPT-4), emotion recognition engines (e.g., IBM Watson Tone Analyzer)
[1349] System Operation Overview
[1350] 1. Taking and uploading images
[1351] Users use their smartphone cameras to take pictures of specific areas within a commercial facility, and the images are uploaded from the smartphone to a cloud storage service (e.g., AWS S3).
[1352] 2. Image Analysis
[1353] The URL of the uploaded image is linked to a generative AI model (e.g., OpenAI GPT-4), which analyzes the location of specific objects (e.g., radishes) in the image and generates a result.
[1354] 3. Audio Feedback
[1355] The generated location information is then communicated to the user via voice via their smartphone, for example, in the form of a notification such as, "The radish is three meters in front of you and to your right."
[1356] 4. Emotion recognition
[1357] The tone of voice a user uses to speak into their smartphone is analyzed by an emotion recognition engine (e.g., IBM Watson Tone Analyzer), which analyzes the user's emotions and identifies feelings such as frustration or relief.
[1358] 5. Feedback adjustment
[1359] Based on the emotion recognition results, the generative AI will adjust the content and tone of the feedback. For example, if the user is feeling frustrated, it will provide more polite feedback, such as, "Don't worry, the radish is three meters in front of you to your right."
[1360] Specific examples
[1361] User: "I don't know where the radish is."
[1362] System: "No problem, the radish is three meters in front of you and to your right."
[1363] Prompt Sentence Examples
[1364] Analyze the image at https: / / your-bucket-name.s3.amazonaws.com / image.jpg and find the location of the radish.
[1365] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1366] Step 1:
[1367] A user takes a picture of a specific area in a commercial facility using the camera on their smartphone.
[1368] Input: Image of a specific area within a commercial facility
[1369] Output: Captured image data
[1370] Specific operation: The user launches the smartphone's camera app and takes a picture of a specific area. The captured image is saved in the smartphone.
[1371] Step 2:
[1372] The device uploads the captured image to a cloud storage service.
[1373] Input: Captured image data
[1374] Output: Image URL on cloud storage
[1375] Specific operation: The smartphone uploads image data to a cloud storage service such as AWS S3 and obtains the image URL.
[1376] Step 3:
[1377] The device connects the URL of the uploaded image to the generative AI.
[1378] Input: Image URL on cloud storage
[1379] Output: Send URL to generative AI
[1380] Specific operation: The smartphone sends the image URL it obtained to a generative AI (e.g., OpenAI GPT-4).
[1381] Step 4:
[1382] The server uses generative AI to analyze the location information of specific objects within the image.
[1383] Input: Image URL
[1384] Output: Location of a specific object
[1385] Specific operation: The generative AI analyzes the image URL and generates the location information of a specific object (e.g., radish). The analysis results are output in text format.
[1386] Step 5:
[1387] The device communicates location information from the generative AI to the user via voice.
[1388] Input: Location of a specific object
[1389] Output: Audio feedback
[1390] Specific operation: The smartphone converts the location information received from the generative AI into voice using a speech synthesis engine and conveys it to the user.
[1391] Step 6:
[1392] The device analyzes the tone of the user's voice using an emotion recognition engine.
[1393] Input: User's tone of voice
[1394] Output: Emotion recognition result
[1395] How it works: The smartphone records the user's voice and sends it to an emotion recognition engine such as IBM Watson Tone Analyzer for analysis. The analysis results include the type and intensity of the emotion.
[1396] Step 7:
[1397] The server adjusts the feedback based on the emotion recognition results.
[1398] Input: Emotion recognition results, location information of specific objects
[1399] Output: Regulated Feedback
[1400] What it does: Generative AI takes emotion recognition results into account and adjusts the content and tone of feedback. For example, if the user is feeling frustrated, it generates more polite feedback.
[1401] Step 8:
[1402] The device then audibly conveys the adjusted feedback to the user.
[1403] Input: Calibrated feedback
[1404] Output: Audio feedback
[1405] Specific operation: The smartphone converts the adjusted feedback into voice using a speech synthesis engine and conveys it to the user.
[1406] Example 2
[1407] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1408] It is difficult for visually impaired people to locate specific products in supermarkets and other commercial facilities. Furthermore, since information is not provided in response to the user's emotions, this can increase the user's frustration. To solve this problem, a system is needed that allows visually impaired people to easily locate specific products and provides information in response to the user's emotions.
[1409] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and transmit location information of a specific object to the user, an emotion recognition means for recognizing the user's emotion, and a means for the generative AI to adjust the information based on the emotion recognition means and transmit it again to the user. This enables visually impaired people to easily grasp the location of specific products and makes it possible to provide information according to the user's emotion.
[1410] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area within a commercial facility.
[1411] "Means for uploading captured images" refers to a device or function that allows a user to send captured images to a server.
[1412] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function that generates a URL for the image received by the server and sends that URL to the generative AI.
[1413] "Generative AI" is an artificial intelligence system that analyzes received images and generates location information for specific objects.
[1414] "Means for communicating the location information of a specific object to the user" refers to devices or functions that notify the user of the location information analyzed by generative AI via voice or text.
[1415] The "emotion recognition means for recognizing the user's emotions" refers to a device or function for analyzing and recognizing the emotions of the user from the tone of voice and facial expressions.
[1416] "Means for the generative AI to adjust information based on the emotion recognition means and re-communicate it to the user" refers to devices or functions that allow the generative AI to adjust information based on the user's emotions recognized by the emotion recognition means and re-notify the user.
[1417] This invention is a system that enables visually impaired people to locate specific products in a commercial facility. This system operates by allowing users to take pictures of specific areas in a store using a device such as a smartphone and upload the images to a server. The specific hardware and software configurations, as well as the data processing and calculation methods, are described below.
[1418] Hardware and software used
[1419] Hardware: Smartphone (e.g. iPhone, Android device)
[1420] Software: Image upload applications, generative AI (e.g., OpenAI's GPT-4), emotion engines (e.g., Affectiva)
[1421] System Operation Overview
[1422] 1. User takes a photo with their smartphone:
[1423] A user starts the camera app on their smartphone and takes a picture of a specific area in a shopping mall (for example, the vegetable section). The user wants to know where the radishes are.
[1424] 2. The device uploads the captured image to the server:
[1425] The device (smartphone) automatically uploads the captured images to the server. Uploading is done through a dedicated application. This application has the function of compressing the images and sending them to the server.
[1426] 3. The server generates a URL for the image and sends it to the generative AI:
[1427] The server saves the received image and generates a URL for the image. The generated URL is sent to the generative AI. The server then sends this URL to the generative AI's API endpoint.
[1428] 4. Generative AI analyzes specific objects in images:
[1429] The generative AI analyzes the image based on the received URL. Specifically, it uses an image recognition algorithm to identify the location of the "daikon radish." For example, it detects that the radish is in the front right of the image.
[1430] 5. Generative AI communicates analysis results to users:
[1431] The generative AI generates the analysis results in natural language and communicates them to the user. For example, it generates a message such as "The radish is in the front right." This message is sent to the device via the server, and the device application notifies the user by voice or text.
[1432] 6. Emotion engine recognizes user emotions:
[1433] If a user says with frustration, "I don't know where the radish is," the device's microphone captures the voice, and the emotion engine analyzes this voice data to recognize the user's emotion.
[1434] 7. Generative AI adjusts information based on emotions and re-communicates it to the user:
[1435] If the emotion engine recognizes the user's frustration, the generative AI will use that information to tailor the message, for example, generating a more polite message like, "Don't worry, the radish is three meters in front of you to your right." This message is also sent via the server to the device and notified to the user.
[1436] Examples and prompts
[1437] Examples:
[1438] A user takes a photo of the vegetable section with their smartphone and says, "I don't know where the radishes are." The device uploads the image to a server and captures audio. The server generates a URL for the image and sends it to a generative AI. The generative AI analyzes the image and generates location information. It adjusts the message based on information from the emotion engine. Finally, the device notifies the user, "Don't worry, the radishes are three meters in front of you to your right."
[1439] Example prompt sentence:
[1440] "Please tell me where the radish is in this image. The user is frustrated."
[1441] This system not only makes it easier for visually impaired people to locate specific products, but also provides services that respond to the user's emotions.
[1442] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1443] Step 1:
[1444] A user takes a photo of a specific area in a store using their smartphone. The user then launches the smartphone's camera app and takes a photo of a specific area in the shopping mall (e.g., the vegetable section). The input is the image taken by the user, and the output is the image data stored in the smartphone.
[1445] Step 2:
[1446] The device uploads the captured images to the server. The device (smartphone) automatically uploads the captured images to the server. The upload is done through a dedicated application. This application has the function of compressing the images and sending them to the server. The input is the image data stored in the smartphone, and the output is the image data uploaded to the server.
[1447] Step 3:
[1448] The server generates a URL for the image and sends it to the generative AI. The server saves the received image and generates a URL for that image. The generated URL is sent to the generative AI. The server sends this URL to the generative AI's API endpoint. The input is the image data uploaded to the server, and the output is the URL of the generated image.
[1449] Step 4:
[1450] The generative AI analyzes a specific object in an image. The generative AI analyzes the image based on the received URL. Specifically, it uses an image recognition algorithm to identify the location of the "radish." For example, it detects that the radish is in the front right of the image. The input is the URL of the image sent to the generative AI, and the output is the location information of the specific object that was analyzed.
[1451] Step 5:
[1452] The generative AI communicates the analysis results to the user. The generative AI generates the analysis results in natural language and communicates them to the user. For example, it generates a message such as "The radish is in the front right." This message is sent to the device via the server, and the device application notifies the user by voice or text. The input is the location information of the specific object that was analyzed, and the output is the message that is communicated to the user.
[1453] Step 6:
[1454] The emotion engine recognizes the user's emotions. If a user says with frustration, "I don't know where the radish is," the device's microphone captures the voice. The emotion engine analyzes this voice data and recognizes the user's emotions. The input is the user's voice data, and the output is the recognized user's emotional information.
[1455] Step 7:
[1456] The generative AI adjusts the information based on the emotion and transmits it back to the user. If the emotion engine recognizes the user's frustration, the generative AI adjusts the message based on that information. For example, it generates a more polite message such as, "Don't worry, the radish is three meters in front of you to your right." This message is also sent via the server to the device and notified to the user. The input is the recognized user's emotional information, and the output is the adjusted message.
[1457] (Application example 2)
[1458] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1459] It is extremely difficult for visually impaired people to find specific products in commercial facilities. Conventional systems do not provide sufficient information to help visually impaired people find specific products, and do not provide feedback based on the user's emotions. Therefore, there is a need to reduce the stress and frustration that visually impaired people experience when trying to find products.
[1460] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1461] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and transmit location information of a specific object to the user, a means for analyzing the user's voice and recognizing emotions, and a means for adjusting and transmitting the location information of the specific object based on the recognized emotion. This enables visually impaired people to receive appropriate feedback according to their emotions when finding a specific product in a commercial facility.
[1462] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area within a commercial facility.
[1463] The "means for uploading captured images" refers to a device or function for transmitting captured images to a server via the Internet.
[1464] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function for providing the internet address of an uploaded image to the generative AI.
[1465] "Generative AI" is an artificial intelligence system that analyzes uploaded images and extracts location information for specific objects.
[1466] "Means for communicating the location information of a specific object to the user" refers to devices or functions that inform the user of the location information analyzed by generative AI via voice or text.
[1467] "Means for analyzing the user's voice and recognizing emotions" refers to a device or function for analyzing the tone and content of the user's voice to determine their emotional state.
[1468] "Means for adjusting and transmitting location information of a specific object based on recognized emotions" refers to a device or function for transmitting location information of a specific object in an appropriate manner depending on the user's emotional state.
[1469] A system for implementing this invention is one that makes it easier for visually impaired people to find specific products in commercial facilities. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and transmit location information of the specific object to the user, a means for analyzing the user's voice and recognizing emotions, and a means for adjusting and transmitting the location information of the specific object based on the recognized emotion.
[1470] Hardware and software used
[1471] Hardware: Smartphone (camera, microphone)
[1472] Software: OpenCV (image capture), requests (image upload and analysis), gTTS (voice feedback), speech_recognition (voice recognition), emotion_recognition (emotion recognition)
[1473] Data processing and calculation
[1474] Image capture and upload
[1475] Users can take photos of specific areas within a commercial facility using their smartphone camera. The images are automatically uploaded to the cloud, and the URLs of the uploaded images are linked to the generative AI.
[1476] Image analysis
[1477] The generative AI on the server analyzes the uploaded image and extracts the location information of a specific object (e.g., a product). This location information is communicated to the user in the form of "in front," "to the right, in front," "not in front," etc.
[1478] Audio Feedback
[1479] The location information analyzed by the generative AI is provided to the user as voice feedback, which is generated using gTTS and played through the smartphone speaker.
[1480] emotion recognition
[1481] The smartphone's microphone is used to analyze the user's voice and recognize emotions. The voice data is converted to text using speech_recognition, and then emotion_recognition is used to analyze emotions.
[1482] Emotion-based feedback regulation
[1483] The generative AI will adjust and communicate the location of a particular object based on the perceived emotion: for example, if the user is feeling frustrated, the generative AI will communicate the location more politely.
[1484] Specific examples
[1485] If a user says, "I don't know where the radish is," the app will recognize frustration from the tone of voice and respond with a voice prompt: "Don't worry, the radish is three meters in front of you to your right."
[1486] Prompt Sentence Examples
[1487] If a user takes a photo of the store with their smartphone camera and says, "I don't know where the radish is," the app analyzes the image, recognizes frustration from the user's tone of voice, and provides a voice prompt saying, "Don't worry, the radish is three meters in front of you to your right."
[1488] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1489] Step 1:
[1490] A user takes a picture of a specific area in a commercial facility using the camera on their smartphone. The input is the image data captured by the camera, and the output is an image file stored on the smartphone.
[1491] Step 2:
[1492] The device uploads the captured image to the cloud server. The input is the image file acquired in step 1, and the output is the image file saved on the cloud server and its URL.
[1493] Step 3:
[1494] The server connects the URL of the uploaded image to the generative AI. The input is the URL of the image file on the cloud server, and the output is the URL passed to the generative AI.
[1495] Step 4:
[1496] The generative AI performs image recognition and extracts the location information of a specific object. The input is the URL of the image linked in step 3, and the output is the location information of the specific object (e.g., "In front, to the right, in front, not in front").
[1497] Step 5:
[1498] The server provides the user with the location information obtained from the generative AI as voice feedback. The input is the location information obtained in step 4, and the output is voice data that is transmitted to the user as voice feedback.
[1499] Step 6:
[1500] The device records the user's voice and uploads it to the cloud server. The input is the user's voice data, and the output is an audio file stored on the cloud server.
[1501] Step 7:
[1502] The server analyzes the voice data and recognizes the user's emotions. The input is the voice file uploaded in step 6, and the output is the recognized emotion information (e.g., frustration, joy, etc.).
[1503] Step 8:
[1504] The server adjusts and transmits the location information of a specific object based on the recognized emotion. The input is the emotion information recognized in step 7 and the location information obtained in step 4, and the output is the adjusted location information (e.g., "It's okay, the radish is 3 meters in front of you to your right").
[1505] Step 9:
[1506] The server provides the adjusted location information to the user as voice feedback. The input is the adjusted location information from step 8, and the output is the voice data that is transmitted to the user as voice feedback.
[1507] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1508] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1509] Another example of generative AI is Gemini (internet search engine). <url: https: gemini.google.com ?hl="ja">) are listed.
[1510] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1511] [Fourth embodiment]
[1512] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1513] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1514] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1515] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1516] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1517] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1518] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1519] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1520] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1521] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1522] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1523] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1524] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.
[1525] "Example 1"
[1526] One embodiment of the present invention provides a system that enables visually impaired people to determine the location of specific products in a supermarket or other store. In this system, a user takes a photo of a specific area in the store using a device such as a smartphone. The captured image is automatically uploaded by the system, and a URL for the image is generated. The generated URL is linked to a generative AI. The generative AI analyzes the location information of a specific object in the image, such as a "daikon radish," and communicates this information to the user in the form of "in front," "to the right, in front," or "not in front." This enables visually impaired people to determine the location of specific products.
[1527] "Example 2"
[1528] One embodiment of the present invention provides a system that enables visually impaired people to determine the location of specific products in a supermarket or other store. In this system, a user takes a photo of a specific area in the store using a device such as a smartphone. The captured image is automatically uploaded by the system, and a URL for the image is generated. The generated URL is linked to a generative AI. The generative AI analyzes the location information of a specific object in the image, such as a "daikon radish," and communicates this information to the user in the form of "in front," "to the right, in front," or "not in front." This enables visually impaired people to determine the location of specific products.
[1529] The processing flow of each embodiment will be described below.
[1530] "Example 1"
[1531] Step 1: A visually impaired person uses a device such as a smartphone to take a photo of a specific area inside a store such as a supermarket.
[1532] Step 2: The system will automatically upload the captured image and generate a URL for it.
[1533] Step 3: The generated URL is linked to the generative AI.
[1534] Step 4: The generative AI analyzes the location information of a specific object in the image, such as a radish.
[1535] Step 5: The generative AI communicates the location of the specific object to the user in the form of "in front of you, to the right, in front of you, not in front of you", etc.
[1536] "Example 2"
[1537] Step 1: A visually impaired person uses a device such as a smartphone to take a photo of a specific area inside a store such as a supermarket.
[1538] Step 2: The system will automatically upload the captured image and generate a URL for it.
[1539] Step 3: The generated URL is linked to the generative AI.
[1540] Step 4: The generative AI analyzes the location information of a specific object in the image, such as a radish.
[1541] Step 5: The generative AI communicates the location of the specific object to the user in the form of "in front of you, to the right, in front of you, not in front of you", etc.
[1542] Example 1
[1543] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1544] It is extremely difficult for visually impaired people to locate specific products in supermarkets and other commercial facilities. Conventional methods require the visually impaired to rely on others for assistance, making it difficult for them to shop independently. Furthermore, existing technologies lack the accuracy and real-time capabilities of image recognition, and are unable to provide sufficient support for visually impaired people.
[1545] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1546] In this invention, the server includes a means for a user to take a picture of a specific area using a mobile device, a means for uploading the taken image to the server, a means for generating a URL for the uploaded image, a means for linking the generated URL to a generative AI model, a means for the generative AI model to analyze location information of a specific object in the image, and a means for communicating the analyzed location information to the user by voice, thereby enabling visually impaired people to independently determine the location of a specific product.
[1547] "User" refers to any user, including visually impaired people, who use the system to locate a particular product.
[1548] "Mobile device" refers to a portable electronic device such as a smartphone or tablet.
[1549] "Specific area" refers to a specific location or section within a commercial establishment.
[1550] "Means for taking a photograph" refers to a method for capturing an image using the camera function built into the mobile terminal.
[1551] "Server" refers to a computer system for storing, processing, and managing data.
[1552] "Means for uploading" refers to a method for transmitting data from a mobile device to a server.
[1553] "URL" refers to an address used to identify a resource on the Internet.
[1554] "Means of generating" refers to the method of creating a URL that indicates the location where the uploaded image is saved.
[1555] A "generative AI model" refers to an artificial intelligence algorithm used for image recognition and data analysis.
[1556] "Means of collaboration" refers to the method of sending data from the server to the generative AI model.
[1557] "Means of analysis" refers to how the generative AI model identifies the location information of specific objects within an image.
[1558] "Means of communicating by voice" refers to a method of communicating the analysis results to the user by voice.
[1559] This invention is a system that enables visually impaired people to locate specific products in a commercial facility such as a supermarket. The system works by having the user take a picture of a specific area using a mobile device and uploading the image to a server.
[1560] Hardware and software used
[1561] Hardware:
[1562] Mobile devices (e.g. smartphones, tablets)
[1563] Server (e.g. cloud server, on-premise server)
[1564] software:
[1565] Mobile application for uploading images
[1566] Image management systems (e.g., cloud storage services)
[1567] API integration system (e.g. REST API)
[1568] Generative AI models (e.g., image recognition algorithms)
[1569] Voice assistants (e.g. text-to-speech software)
[1570] System Operation
[1571] The user takes a photo with their mobile device
[1572] Users can open the camera app on their mobile device and take a picture of a specific area in a shopping mall. For example, if a user is looking for daikon radishes in the vegetable section, they can take a picture of the entire vegetable section.
[1573] The device uploads the image to the server.
[1574] The captured images are automatically uploaded from the mobile device to the server, and a notification is displayed when the upload is complete.
[1575] The server generates the image URL
[1576] The server generates a URL for the uploaded image, which is then stored in the database.
[1577] The server links the URL to the generated AI model
[1578] The server connects the generated URL to the AI model, and the URL is sent via API.
[1579] A generative AI model analyzes the image
[1580] The generative AI model analyzes the location of specific objects (e.g., radishes) in the image, and the analysis results are output in text format.
[1581] Generative AI model communicates location information to user
[1582] Based on the analysis results, the generative AI model communicates the location information to the user via voice, and the voice assistant is activated to read the information aloud.
[1583] Specific examples
[1584] Let's say a user is looking for a "daikon radish" in the vegetable section of a supermarket. The user takes a photo of the entire vegetable section with their smartphone and uploads the image to the system. The system generates a URL for the image and connects it to the generative AI model. The generative AI model analyzes the image and tells the user by voice, "The daikon radish is in the front right."
[1585] Prompt Sentence Examples
[1586] "Please tell me the location of the radish in this image. Please tell the user in the form of 'in front, to the right, in front, not in front'."
[1587] This system allows visually impaired people to independently locate specific products.
[1588] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1589] Step 1:
[1590] Users take photos of the store interior using their mobile devices
[1591] Input: A user opens the camera app on their mobile device and takes a picture of a specific area.
[1592] Specific Actions: The user launches the camera app on their mobile device and presses the camera button to capture an image.
[1593] Output: The captured image data is saved on the mobile device.
[1594] Step 2:
[1595] The device uploads the image to the server.
[1596] Input: Captured image data
[1597] What it does: Your device will automatically upload the image to the server, and once the upload is complete, a notification will appear on your device.
[1598] Output: Image data is saved on the server.
[1599] Step 3:
[1600] The server generates the image URL
[1601] Input: Image data stored on the server
[1602] Specific operation: The server generates a URL indicating the location where the image is saved and saves that URL in the database.
[1603] Output: The URL of the generated image (e.g. https: / / example.com / image123.jpg)
[1604] Step 4:
[1605] The server links the URL to the generated AI model
[1606] Input: URL of the generated image
[1607] Specific operation: The server sends the image URL to the generative AI model via API, and records the completion of the integration in a log.
[1608] Output: The image URL is sent to the generative AI model.
[1609] Step 5:
[1610] A generative AI model analyzes the image
[1611] Input: The URL of the image sent to the generative AI model
[1612] How it works: The generative AI model downloads an image and analyzes the location of a specific object (e.g., a radish). The analysis results are output in text format.
[1613] Output: Analysis result (e.g. "The radish is in the front right").
[1614] Step 6:
[1615] Generative AI model communicates location information to user
[1616] Input: Text data of analysis results
[1617] How it works: The generative AI model sends the analysis results to the voice assistant, which then reads the information aloud.
[1618] Output: The location information is spoken to the user (e.g., "The radish is in front of you to the right").
[1619] (Application example 1)
[1620] Next, a description will be given of Application Example 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1621] It is very difficult for visually impaired people to find specific products in commercial facilities such as supermarkets. With conventional methods, visually impaired people need the help of others, making it difficult for them to shop independently. To solve this problem, a system that allows visually impaired people to find specific products by themselves is needed.
[1622] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and communicate location information of the specific object to the user, and a means for providing audio feedback of the location information. This enables visually impaired people to find specific products in commercial facilities using their smartphones.
[1623] An "image capture means" is a device that allows a user to capture an image of a particular area.
[1624] "Means for uploading captured images" is a function for sending captured images to a cloud or server.
[1625] "Means for linking the URL of uploaded images to generative AI" is a function for providing the URL of uploaded images to generative AI.
[1626] "Generative AI" is artificial intelligence that performs image recognition and analyzes the location information of specific objects.
[1627] "Means of communicating the location information of specific objects to users" is a function that informs users of the location information analyzed by generative AI.
[1628] "Means for providing audio feedback of location information" is a function for conveying analyzed location information to the user via audio.
[1629] A "commercial facility" is a place where consumers can purchase goods, such as a supermarket or department store.
[1630] A system for implementing this invention is for assisting visually impaired people in finding specific products in commercial facilities. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to recognize the image and communicate the location information of the specific object to the user, and a means for providing audio feedback of the location information.
[1631] Users use their smartphone camera to take pictures of specific areas within a commercial facility. The images are then uploaded to the cloud via a smartphone application. A URL is automatically generated for the uploaded image, and this URL is linked to the generative AI.
[1632] The generative AI performs image recognition based on the provided image URL and analyzes the location information of a specific object, such as a "daikon radish." As a result of the analysis, the generative AI generates location information in the form of "in front," "to the right, in front," "not in front," etc. This location information is fed back to the user via audio via a smartphone application.
[1633] Specifically, the system works as follows:
[1634] 1. Hardware: Use your smartphone camera to capture the image.
[1635] 2. Software: Use OpenCV to capture images and requests library to upload images to cloud.
[1636] 3. Data processing: Once the image is uploaded to the cloud, a URL is generated.
[1637] 4. Data calculation: Send the image URL and item name to the generative AI, which analyzes the location information of the specific product.
[1638] 5. Feedback: The acquired location information is converted into audio using gTTS (Google Text-to-Speech) and provided as feedback to the user.
[1639] For example, if the user is looking for "daikon radish," the prompt text might look like this:
[1640] Example prompt sentence:
[1641] Could you please tell me the location of the radish in the image?
[1642] By sending this prompt to the generative AI, the AI will analyze the location of the radish in the image and return location information in the form of "in front, to the right, in front, not in front," etc. This allows visually impaired people to find specific products on their own.
[1643] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1644] Step 1:
[1645] A user takes a picture of a specific area in a commercial facility using the camera on their smartphone. The input is the image captured by the user through the camera. The output is an image file stored on the smartphone.
[1646] Step 2:
[1647] The device uploads the captured image to the cloud. The input is the image file obtained in step 1. The output is the URL of the image stored on the cloud. Specifically, the device uses the requests library to send the image file to the cloud server.
[1648] Step 3:
[1649] The device connects the URL of the uploaded image to the generative AI. The input is the URL of the image generated in step 2. The output is the URL and prompt sent to the generative AI. Specifically, the device sends the generative AI a prompt saying, "Please tell me the location of a specific object in the image."
[1650] Step 4:
[1651] The server uses generative AI to perform image recognition and analyze the location information of specific objects. The input is the URL of the image sent in step 3 and the prompt text. The output is the location information of the specific object. Specifically, the generative AI analyzes the specific object in the image and generates location information such as "in front," "to the right and in front," or "not in front."
[1652] Step 5:
[1653] The device provides voice feedback of the location information obtained from the generative AI. The input is the location information obtained in step 4. The output is voice feedback provided to the user. Specifically, the device converts the location information into voice using gTTS (Google Text-to-Speech) and transmits it to the user through the smartphone speaker.
[1654] These steps allow a visually impaired person to locate a particular product within a commercial establishment.
[1655] Example 2
[1656] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1657] It is difficult for visually impaired people to locate specific products in commercial facilities such as supermarkets. With conventional methods, it is difficult for visually impaired people to find products on their own and they need to get help from others. To solve this problem, a system that allows visually impaired people to locate specific products on their own is needed.
[1658] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1659] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI model, a means for the generative AI model to perform image recognition and communicate location information of a specific object to the user, a means for transmitting the generated location information to the user's terminal, and a means for the user to check the location information on the terminal. This enables visually impaired people to independently determine the location of specific products.
[1660] An "image capturing means" is a device that allows a user to capture an image of a specific area within a commercial facility.
[1661] "Means for uploading captured images" is a function for sending image data from the user's terminal to the server.
[1662] "Means for linking the URL of uploaded images to the generative AI model" is a function that generates a URL for the image data received by the server and provides that URL to the generative AI model.
[1663] A "generative AI model" is an artificial intelligence model that uses image recognition technology to analyze the location information of specific objects within a provided image.
[1664] "Means of communicating location information of specific objects to users" is a function for communicating location information analyzed by the generative AI model to users.
[1665] "Means for transmitting the generated location information to the user's terminal" refers to a function that enables the server to transmit the location information received from the generating AI model to the user's terminal.
[1666] "Means for users to check location information on their devices" refers to a function that allows users to check received location information by voice or text using their own devices.
[1667] This invention is a system that enables visually impaired people to determine the location of specific products in a commercial facility. The system allows users to take pictures of specific areas in a store using a device such as a smartphone, and analyzes the images to provide location information for specific objects.
[1668] A user launches the camera app on their smartphone and takes a picture of a specific area in a shopping mall. For example, if the user is looking for a "daikon radish," they take a picture of the vegetable section. The captured image is automatically uploaded to the server through the smartphone's application. At this time, the device uses an Internet connection to send the image data to the server. Specifically, the image data is sent using an HTTP POST request.
[1669] The server stores the received image data and generates a URL for that image. The generated URL is linked to the generative AI model. Specifically, an API request is sent to the generative AI model, providing a prompt message including the image URL. The generative AI model analyzes the image based on the provided image URL. For example, it uses Google Cloud Vision API or Amazon Rekognition to identify the location of the "daikon radish" in the image. As a result of the analysis, location information such as "The daikon radish is in the front right" is generated.
[1670] The server sends the location information received from the generative AI model to the user's smartphone. Specifically, it sends data containing the location information as an HTTP response. The user then checks the received location information through an application on their smartphone. The application then communicates the location information to the user via voice or text. For example, it may provide a voice prompt saying, "The radish is in front of you to the right."
[1671] As a concrete example, consider the case where a user is looking for a "daikon radish" in a supermarket. The user takes a photo of the area inside the store with their smartphone and uploads the image to the system. The server generates a URL for the image and connects it to the generative AI model. The generative AI model analyzes the image and generates location information such as "The daikon radish is in the front right." The server sends this information to the user's smartphone, and the user confirms the location information by voice.
[1672] Example prompt sentence:
[1673] "Please tell me the location of the radish in this image."
[1674] This system allows visually impaired people to locate specific products.
[1675] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1676] Step 1:
[1677] The user takes a photo of a specific area inside the store using their smartphone.
[1678] Specifically, a user launches a camera app on their smartphone and takes a picture of a specific area, such as a vegetable section. The input is the image taken by the user, and the output is the image data stored on the smartphone.
[1679] Step 2:
[1680] The device uploads the captured image to the server.
[1681] Specifically, the device uses an Internet connection to send image data to a server. The input is image data stored on the smartphone, and the output is image data uploaded to the server. The image data is sent using an HTTP POST request.
[1682] Step 3:
[1683] The server generates a URL for the image and connects it to the generative AI model.
[1684] Specifically, the server saves the received image data and generates a URL for the image. The input is the image data uploaded to the server, and the output is the generated image URL. The generated URL sends an API request to the generative AI model and provides a prompt containing the image URL.
[1685] Step 4:
[1686] A generative AI model analyzes the image and generates location information for specific objects.
[1687] Specifically, the generative AI model analyzes an image based on the provided image URL. The input is the image URL and a prompt, and the output is the location information of a specific object. For example, the location of a "daikon radish" in the image can be identified using the Google Cloud Vision API or Amazon Rekognition. The analysis results in location information such as "The daikon radish is in the front right."
[1688] Step 5:
[1689] The server sends the generated location information to the user's smartphone.
[1690] Specifically, the server sends the location information received from the generative AI model to the user's smartphone. The input is the location information from the generative AI model, and the output is the location information sent to the user's smartphone. Data including the location information is sent as an HTTP response.
[1691] Step 6:
[1692] The user checks the location on their smartphone.
[1693] Specifically, the user checks the location information received through a smartphone application. The input is the location information sent from the server, and the output is the location information checked by the user. The application then communicates the location information to the user via voice or text. For example, it may provide a voice prompt saying, "The radish is in front of you to the right."
[1694] (Application example 2)
[1695] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1696] It is difficult for visually impaired people to locate specific products in stores such as supermarkets. Conventional methods have made it difficult for visually impaired people to find products on their own and have required the help of others. For this reason, there is a demand for a system that allows visually impaired people to shop independently.
[1697] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1698] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and communicate location information of a specific object to the user, a means for communicating the analysis result to the user by voice, a voice synthesis means for generating voice, and a voice playback means for playing voice, thereby enabling visually impaired people to locate specific products and shop independently.
[1699] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area inside the store.
[1700] "Means for uploading captured images" refers to devices or functions for sending captured images to a cloud or server.
[1701] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function for passing the URL of an uploaded image to the generative AI.
[1702] "Generative AI" is artificial intelligence that performs image recognition and analyzes the location information of specific objects.
[1703] "Means for communicating location information of a specific object to a user" refers to a device or function for informing a user of analyzed location information.
[1704] "Means for communicating analysis results to users via voice" refers to devices or functions that communicate the results of analysis by generative AI to users via voice.
[1705] "Speech synthesis means for generating speech" refers to a device or function for converting text information into speech.
[1706] The "audio playback means for playing back audio" refers to a device or function for allowing the user to hear the generated audio.
[1707] A system for carrying out this invention assists visually impaired people in locating specific products in a store such as a supermarket. The system includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and communicate the location information of the specific object to the user, a means for communicating the analysis results to the user by voice, a voice synthesis means for generating voice, and a voice playback means for playing back voice.
[1708] Hardware and software used
[1709] Hardware: Smartphone (camera, microphone, speaker)
[1710] software:
[1711] OpenCV (Image Capture)
[1712] requests (HTTP requests)
[1713] gTTS (Google Text-to-Speech)
[1714] mpg321 (audio playback)
[1715] System Operation
[1716] 1. Image capture:
[1717] The user takes a photo of a specific area in the store using their smartphone camera, and the image is captured using OpenCV and saved locally.
[1718] 2. Image upload:
[1719] Upload the captured image to the cloud by using the requests library to send the image as a POST request to the specified URL.
[1720] 3. Image Analysis:
[1721] The URL of the uploaded image is linked to the generative AI. The generative AI analyzes the location information of specific objects in the image. The analysis results are returned in the form of, for example, "In front, in front to the right, not in front."
[1722] 4. Providing Feedback:
[1723] The analysis results are communicated to the user via voice. gTTS is used to convert the text information into voice and play it back in mpg321.
[1724] Specific examples
[1725] If a user is looking for a "daikon radish," they can take a photo of the inside of the store with their smartphone camera. The image is uploaded to the cloud, and generative AI analyzes the location of the "daikon radish." If the analysis returns "It's in front of you on the right," the smartphone will relay that information to the user via voice.
[1726] Prompt Sentence Examples
[1727] Analyze the location of the "radish" in the image. Return the result in the form of "In front, in front right, not in front", etc.
[1728] The system allows visually impaired people to locate specific products and shop independently.
[1729] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1730] Step 1:
[1731] A user uses a smartphone camera to take a picture of a specific area in a store. The input is real-time video from the camera, and the output is a still image. Specifically, a user launches the camera app, frames a specific area, and presses the shutter button to capture the image.
[1732] Step 2:
[1733] The device saves the captured image to local storage. The input is the image data captured in step 1, and the output is an image file saved in local storage. Specifically, the image is saved in JPEG format using the OpenCV library.
[1734] Step 3:
[1735] Uploads an image saved on the device to a cloud server. The input is an image file saved in local storage, and the output is an image URL on the cloud server. Specifically, the requests library is used to send the image file to the specified URL as a POST request.
[1736] Step 4:
[1737] The server connects the URL of the uploaded image to the generative AI. The input is the image URL on the cloud server, and the output is the URL data passed to the generative AI. Specifically, the server sends the URL to the API endpoint of the generative AI.
[1738] Step 5:
[1739] The generative AI analyzes the location information of a specific object in an image. The input is the image URL, and the output is the location information of the specific object. Specifically, the generative AI uses an image recognition algorithm to analyze the location of a specific object (e.g., a "radish") in the image and returns a result in the form of "In front, in front to the right, not in front of you," etc.
[1740] Step 6:
[1741] The server generates data to communicate the analysis results to the user via voice. The input is location information from the generative AI, and the output is text data for speech synthesis. Specifically, the analysis results are converted into text format and passed to the speech synthesis API.
[1742] Step 7:
[1743] The device generates the audio and plays it back to the user. The input is text data for speech synthesis, and the output is audio that the user hears. Specifically, the gTTS library is used to convert the text to an audio file and play it back in mpg321.
[1744] The above steps enable a visually impaired person to locate specific products and shop independently.
[1745] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1746] "Example 1"
[1747] In one embodiment of the present invention, the generative AI includes an emotion engine that recognizes a user's emotions. This emotion engine recognizes emotions from the user's tone of voice. For example, if a user says, "I don't know where the radish is," with a sense of frustration in their voice, the emotion engine recognizes that frustration. The generative AI then communicates location information for a specific object based on the recognized emotion. Specifically, if the generative AI recognizes that the user is feeling frustrated, it communicates location information more politely, such as, "Don't worry, the radish is three meters in front of you to your right." This makes it possible to provide a service that responds to the user's emotions.
[1748] "Example 2"
[1749] In one embodiment of the present invention, the generative AI includes an emotion engine that recognizes a user's emotions. This emotion engine recognizes emotions from the user's tone of voice. For example, if a user says, "I don't know where the radish is," with a sense of frustration in their voice, the emotion engine recognizes that frustration. The generative AI then communicates location information for a specific object based on the recognized emotion. Specifically, if the generative AI recognizes that the user is feeling frustrated, it communicates location information more politely, such as, "Don't worry, the radish is three meters in front of you to your right." This makes it possible to provide a service that responds to the user's emotions.
[1750] The processing flow of each embodiment will be described below.
[1751] "Example 1"
[1752] Step 1: The user says with frustration in their voice, "I don't know where the radish is."
[1753] Step 2: The emotion engine recognizes frustration from the user's tone of voice. Step 3: Based on the emotion recognized by the generative AI, the location information of the radish is communicated. Specifically, the location information is communicated more politely, such as "Don't worry, the radish is three meters in front of you to your right."
[1754] "Example 2"
[1755] Step 1: The user says with frustration in their voice, "I don't know where the radish is."
[1756] Step 2: The emotion engine recognizes frustration from the user's tone of voice. Step 3: Based on the emotion recognized by the generative AI, the location information of the radish is communicated. Specifically, the location information is communicated more politely, such as "Don't worry, the radish is three meters in front of you to your right."
[1757] Example 1
[1758] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1759] It is difficult for visually impaired people to locate specific products in supermarkets and other commercial facilities. It is also difficult to provide appropriate feedback based on the user's emotions. This often causes stress for visually impaired people when shopping. Therefore, there is a need for a system that allows visually impaired people to easily locate specific products and provides feedback based on the user's emotions.
[1760] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1761] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to the generative AI, a means for the generative AI to perform image recognition and communicate location information of a specific object to the user, an emotion engine that recognizes the user's emotions, and a means for providing feedback based on the emotions recognized by the emotion engine. This enables visually impaired people to easily determine the location of specific products and provides appropriate feedback according to the user's emotions.
[1762] "Image capturing means" refers to a device or function that allows a user to capture an image of a specific area within a commercial facility.
[1763] "Means for uploading captured images" refers to a device or function that allows a user to send captured images to a server.
[1764] "Means for linking the URL of an uploaded image to the generative AI" refers to a device or function that generates a URL for the image received by the server and sends that URL to the generative AI.
[1765] "Generative AI" is an artificial intelligence system that performs image recognition and analyzes the location information of specific objects.
[1766] "Means for communicating the location information of a specific object to the user" refers to devices or functions for notifying the user of the location information analyzed by generative AI.
[1767] An "emotion engine that recognizes user emotions" is a device or function that analyzes the tone of a user's voice and other emotional expressions to recognize the user's emotions.
[1768] "Means for providing feedback based on emotions recognized by the emotion engine" refers to devices or functions that enable the generative AI to provide appropriate feedback in response to the user's emotions recognized by the emotion engine.
[1769] MODE FOR CARRYING OUT THE INVENTION
[1770] This invention is a system that allows visually impaired people to determine the location of specific products in a commercial facility. The system allows users to take pictures of specific areas in a store using a device such as a smartphone, and by analyzing the images, provides location information for specific objects. It also includes a function to recognize the user's emotions and provide feedback according to the emotions.
[1771] Hardware and software used
[1772] 1. Smartphone (device)
[1773] Camera application: Used by users to take pictures of specific areas in the store.
[1774] Upload function: Used to send captured images to the server.
[1775] 2. Server
[1776] Image saving function: Saves received images and generates their URLs.
[1777] URL generation function: Generates a URL for the saved image and connects it to the generative AI.
[1778] 3. Generative AI
[1779] Image recognition software: Uses software such as Google Cloud Vision API or Amazon Rekognition to analyze the location of specific objects in an image.
[1780] Feedback function: Used to notify the user of the analysis results.
[1781] 4. Emotion Engine
[1782] Voice analysis software: Using software such as IBM Watson Tone Analyzer, it recognizes emotions from the tone of a user's voice.
[1783] Emotion-based feedback: Generative AI provides appropriate feedback based on the emotions it recognizes.
[1784] Specific examples
[1785] Consider a scenario in which a user takes a photo of a vegetable section in a shopping mall with their smartphone and says, "I don't know where the radishes are." In this case, the system operates as follows.
[1786] 1. The user launches the camera app on their smartphone and takes a picture of the vegetable section. When the user presses the capture button, the image is saved on the device.
[1787] 2. The device automatically uploads the image to the server using the server endpoint configured in the application (e.g., https: / / example.com / upload).
[1788] 3. The server receives and saves the image, generates a URL for the saved image (e.g., https: / / example.com / images / 12345.jpg), and sends the URL to the generative AI API.
[1789] 4. The generative AI receives the image URL and uses the Google Cloud Vision API to analyze the location of the "radish" in the image. As a result, it obtains location information such as "The radish is in the front right."
[1790] 5. The generative AI notifies the user of the analysis results using the smartphone's voice synthesis function. The user's smartphone will then say aloud, "The radish is in front of you to the right."
[1791] 6. The emotion engine analyzes the user's tone of voice and recognizes frustration. The generative AI then conveys the location more politely, such as "Don't worry, the radish is three meters in front of you to your right."
[1792] Prompt Sentence Examples
[1793] "Users can upload images they take with their smartphones and then analyze and tell us the location of specific items. If users get frustrated, we'll give them a more polite way to find the location."
[1794] This system allows visually impaired people to easily locate specific products within a commercial facility and provides appropriate feedback based on the user's emotions.
[1795] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1796] Program processing flow
[1797] Step 1:
[1798] A user takes a photo of a specific area in a commercial facility using their smartphone.
[1799] Input: The user launches the camera app and presses the capture button.
[1800] Data processing: The smartphone camera captures image data and stores it on the device.
[1801] Output: Captured image data.
[1802] Specific operation: The user launches the camera app on their smartphone and takes a picture of the vegetable section. When the user presses the capture button, the image is saved on the device.
[1803] Step 2:
[1804] The device uploads the captured image to the server.
[1805] Input: Image data stored on the device.
[1806] Data processing: The device sends image data to the server using an HTTP POST request.
[1807] Output: Image data uploaded to the server.
[1808] Specific behavior: The device automatically uploads the image to the server using the server endpoint configured in the application (e.g., https: / / example.com / upload).
[1809] Step 3:
[1810] The server generates a URL for the image and connects it to the generative AI.
[1811] Input: Image data uploaded to the server.
[1812] Data processing: The server saves the image and generates a URL for it. The generated URL is sent to the generative AI API.
[1813] Output: The generated image URL.
[1814] Specific operation: The server receives and saves the image. It generates a URL for the saved image (e.g., https: / / example.com / images / 12345.jpg) and sends that URL to the generative AI API.
[1815] Step 4:
[1816] Generative AI analyzes the location of specific objects within an image.
[1817] Input: Image URL sent to the generative AI.
[1818] Data processing: Generative AI uses image recognition software (e.g., Google Cloud Vision API) to analyze the location of a specific object (e.g., a "daikon radish") in an image.
[1819] Output: Location information of a specific object.
[1820] Specific operation: The generative AI receives the image URL and uses the Google Cloud Vision API to analyze the location of the "radish" in the image. As a result of the analysis, it obtains location information such as "The radish is in the front right."
[1821] Step 5:
[1822] The generative AI communicates the analysis results to the user.
[1823] Input: Location information of a specific object.
[1824] Data processing: The generative AI notifies the user of the analysis results via voice or text message.
[1825] Output: The location information communicated to the user.
[1826] Specific operation: The generative AI notifies the user of the analysis results using the smartphone's voice synthesis function. The user's smartphone will then say aloud, "The radish is in front of you to the right."
[1827] Step 6:
[1828] The emotion engine recognizes the user's emotions, and the generative AI provides feedback based on the emotions.
[1829] Input: The user's tone of voice.
[1830] Data processing: The emotion engine uses voice analysis software (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions. The generative AI provides feedback based on the recognized emotions.
[1831] Output: Emotion-based feedback.
[1832] How it works: The emotion engine analyzes the user's tone of voice and recognizes frustration. The generative AI then conveys the location information more politely, such as, "Don't worry, the radish is three meters in front of you to your right."
[1833] (Application example 1)
[1834] Next, a description will be given of Application Example 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1835] It is extremely difficult for visually impaired people to locate specific products in commercial facilities. Conventional systems make it difficult for visually impaired people to find products on their own, often causing them great stress when shopping. Furthermore, there is no system that provides feedback based on the user's emotions, making it impossible to reduce user frustration. A new system is needed to solve these issues.
[1836] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1837] In this invention, the server includes an image capturing means, a means for uploading the captured image, a means for linking the URL of the uploaded image to a generative AI, a means for the generative AI to perform image recognition and communicate location information of a specific object to the user, an emotion recognition means for recognizing the user's emotion, and a means for adjusting feedback based on the emotion recognition result. This enables visually impaired pe...
Claims
[Claim 1] A system including a terminal and a server, an image capturing means for capturing an image of a specific area within the commercial facility using a camera function provided in the terminal; means for automatically uploading the image taken by the image taking means from the terminal to the server; The server generates a URL indicating a storage location of the uploaded image and links the URL to a generation system AI; The generative AI performs image recognition of an object specified as a recognition target based on an instruction from a user, based on the linked URL, analyzes location information of the object, and transmits the analyzed location information to the user through the terminal by voice synthesis; an emotion engine that recognizes the emotion of a user by analyzing the tone of the user's voice captured by a microphone function of the terminal; A means for the generative AI to adjust at least one of the content and tone of the analyzed location information based on the emotion recognized by the emotion engine and provide the information to the user through the terminal by voice synthesis; A system including:
Citation Information
Patent Citations
Merchandise information providing terminal and merchandise information providing system
JP2011253324A
Behavior analysis method and apparatus
JP2012000449A
Knowledge information processing server system with image recognition system
JP2013088906A
Vending machine management device, data processing method of the same and program
JP2014127168A
Image recognition system for inventory control and marketing
JP2014218313A