System
The system converts artwork data using generative AI to allow visually or hearing impaired users to experience their work through alternative sensory inputs, providing accurate feedback.
Patent Information
- Application Number
- JP2024116516
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Conventional methods are difficult for visually or hearing-impaired users to fully experience works that require visual or auditory input and receive accurate feedback.
A system that converts artwork data into formats suitable for the user's sensory organs, using generative AI to transform image files into audio data for visually impaired users and audio files into vibration patterns or visual waveforms for hearing impaired users, allowing feedback through other sensory organs.
Enables visually or hearing impaired users to experience their own artwork and receive accurate feedback through alternative sensory inputs.
Smart Images

Figure 2026015042000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional methods have been difficult for visually or hearing-impaired users to fully experience works that require visual or auditory input. These users also face difficulties in receiving feedback about the works. The present invention aims to provide a system that allows such users to experience works through other sensory organs and receive accurate feedback. [Means for solving the problem]
[0005] The present invention is a system for converting artwork into a different format depending on the state of the user's sensory organs and providing the artwork. Specifically, the system includes a means for acquiring artwork data, a means for invoking a generation AI that converts the acquired artwork data into a different format based on the state of the user's sensory organs, a means for providing the converted data to the user, a means for receiving feedback from the user, a means for analyzing the feedback, a means for re-invoking the generation AI based on the analysis results to perform additional conversion, and a means for providing the analysis results to the user. In particular, by including a means for converting artwork data into audio data if the artwork data is an image file, and a means for converting audio files into vibration patterns or visual waveforms, the system enables the artwork to be experienced through a wide range of sensory organs.
[0006] "Work Data" means visual or audio data (e.g., image files or audio files) created by a user.
[0007] "User" means an individual who uses the System to seek feedback on their work.
[0008] A "sensory organ" is a part of the user's body that senses external stimuli, such as sight, hearing, or touch.
[0009] "Generative AI" is a program that uses artificial intelligence to transform specific input data into a different format.
[0010] "Conversion" is the process of changing work data into a different format that is appropriate for the user's sensory state.
[0011] "Feedback" refers to the act of a user providing feedback on the converted data, such as their impressions and suggestions for improvement.
[0012] "Analysis" is the process of analyzing the collected feedback and extracting key points and improvement suggestions.
[0013] A "terminal" is a device used by a user (e.g., a smartphone, tablet, or PC).
[0014] The "system" is a collection of devices and programs according to the present invention that transforms the artwork according to the state of the user's sensory organs and provides feedback. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The present invention provides a system that allows users who lack a particular sensory organ to experience their own artwork through another sensory organ. Specific embodiments of this system will be described below.
[0037] 1. System Overview
[0038] This system consists of a server, a terminal, and a generation AI.
[0039] Users use their terminals to send their creation data to the server and receive feedback.
[0040] 2. Receiving a Request
[0041] A user sends a request from a terminal for feedback on their creation (e.g., a painting or a song).
[0042] 3. Acquiring work data
[0043] The server receives the work data sent from the terminal and stores, for example, image files and audio files appropriately.
[0044] 4. Check the condition of the sensory organs
[0045] The server consults a database to ascertain the user's sensory status, for example to determine whether the user is visually impaired or hearing impaired.
[0046] 5. Calling the generation AI
[0047] The server sends the acquired artwork data to the generation AI and sends a request to convert it into a different format depending on the state of the user's sensory organs.
[0048] 6. Implementing sensory transformation
[0049] The generative AI converts the received artwork data into a new format as specified, for example converting image data into audio data for visually impaired users, or audio data into vibration patterns or visual waveforms for hearing impaired users.
[0050] 7. Sending the converted data
[0051] The generation AI returns the converted data to the server, which then sends the data to the user's device.
[0052] 8. Providing conversion data
[0053] The device presents the converted data to the user, either as audio for visually impaired users or as vibrations or visual waveforms for hearing impaired users.
[0054] 9. Collecting Feedback
[0055] The user inputs feedback on the converted data through the terminal, and the terminal transmits the feedback to the server.
[0056] 10. Feedback Analysis
[0057] The server analyzes the received feedback and, if necessary, sends another request to the generating AI to perform additional transformations or improvements.
[0058] 11. Providing Final Feedback
[0059] The server generates feedback based on the analysis results, sends the final analysis results to the terminal, and provides the final feedback to the user.
[0060] As a concrete example, consider the case where a visually impaired user requests feedback on a painting (image file) they have created. The user uses their device to send the image file of the painting to a server. The server receives the image file and confirms that the user is visually impaired. The server then sends a request to the generation AI to convert the image file into audio data. The generation AI maps the image's colors and shapes to sounds and returns the result to the server as an audio file. The server then sends the audio file to the user's device, which plays it and provides it to the user. The user listens to the audio data and enters their impressions and suggestions for improvement as feedback, which the device then sends to the server. The server analyzes the feedback and, if necessary, sends another request to the generation AI to make additional conversions or improvements. Finally, the server provides the analysis results to the user in text format. This process allows visually impaired users to experience their own work as audio and receive feedback.
[0061] This allows visually or hearing impaired users to experience the work through other sensory organs and receive accurate feedback.
[0062] The processing flow will be explained below.
[0063] Step 1:
[0064] A user uses a terminal to send a request for feedback on his / her work data (e.g., an image file or an audio file).
[0065] Step 2:
[0066] The server receives the request and work data sent from the terminal, analyzes the HTTP request, and obtains the contents of the attached file or data.
[0067] Step 3:
[0068] The server retrieves the user's sensory status from a database, including whether the user is visually or hearing impaired.
[0069] Step 4:
[0070] The server sends a conversion request to the generation AI based on the acquired artwork data. This request includes the artwork data to be converted and the state of the user's sensory organs.
[0071] Step 5:
[0072] The generative AI converts the received artwork data into a format that corresponds to the state of the user's sensory organs. For example, when converting images into audio data, it maps colors and shapes to pitch and rhythm.
[0073] Step 6:
[0074] The generated AI sends the converted data (e.g., audio files, vibration data) back to the server, where it is appropriately encoded.
[0075] Step 7:
[0076] The server sends the converted data to the terminal, encoding the data as an HTTP response.
[0077] Step 8:
[0078] The terminal presents the received converted data to the user, playing audio data for visually impaired users and displaying vibrations or visual waveforms for hearing impaired users.
[0079] Step 9:
[0080] The user inputs feedback on the presented conversion data through the terminal, including their impressions and suggestions for improvement.
[0081] Step 10:
[0082] The terminal transmits the feedback from the user to the server. The feedback data is sent to the server as an HTTP request.
[0083] Step 11:
[0084] The server analyzes the received feedback using natural language processing technology to extract important points and suggestions for improvement.
[0085] Step 12:
[0086] The server may send a request to the generation AI again based on the results of analyzing the feedback, to perform additional transformations or improvements.
[0087] Step 13:
[0088] The server generates final feedback based on the analysis results and sends it to the terminal in text format.
[0089] Step 14:
[0090] The device will then present the final feedback to the user, allowing them to see how their work can be improved and the details of the feedback.
[0091] Example 1
[0092] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0093] This solves the problem of users who lack certain sensory organs having difficulty experiencing their own work through other sensory organs and receiving appropriate feedback. Conventional technologies have limited the means by which visually or hearing-impaired users can obtain feedback on their own work, and no efficient system exists for this purpose.
[0094] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0095] In this invention, the server includes means for receiving a request for feedback on work data from a user, means for acquiring work data, means for saving the acquired work data, means for checking the state of the user's sensory organs, means for calling a generation AI that converts the acquired work data into another format based on the state of the user's sensory organs, means for sending a conversion request to the generation AI, means for acquiring the converted data, means for providing the converted data to the user, means for receiving feedback from the user, means for analyzing the feedback, means for again calling the generation AI based on the analysis result to perform additional conversion, and means for providing the analysis result to the user. This enables users with visual or hearing impairments to experience the work through their other sensory organs and receive accurate feedback based on the results.
[0096] "User" refers to a person who uses this system to submit work data and receive feedback.
[0097] "Work data" refers to digital content such as images and audio created by users.
[0098] "Feedback" refers to evaluations, opinions, and comments for improvement provided regarding work data.
[0099] "Status of sensory organs" refers to the functional status of the sensory organs held by the user, for example, whether or not there is a visual or hearing impairment.
[0100] "Generative AI" refers to artificial intelligence used to transform work data into another format.
[0101] "Conversion request" refers to a command that instructs the generation AI to convert work data into a different format.
[0102] "Converted data" refers to data in a new format that has been converted from the original work data by the generating AI.
[0103] "Server" refers to a computer system that receives requests from users, stores work data, and converts and provides the data in cooperation with the generative AI.
[0104] A "request" refers to a request made by a user to a server, for example, a request for feedback.
[0105] "Analysis" refers to the process by which the server analyzes the feedback obtained from the user.
[0106] This invention provides a system that allows users who lack a particular sensory organ to experience their own artwork through another sensory organ. Specific embodiments of this system are described below.
[0107] Hardware and software used
[0108] This system consists of a server, a terminal, and a generative AI. Specific hardware includes high-performance computers (e.g., AWS EC2 instances), and terminals include users' smartphones and tablets. Generative AI uses GPT-4, Stable Diffusion, DALL-E, and other algorithms.
[0109] System Operation Overview
[0110] The following is an overview of how this system works.
[0111] User operation: A user sends a request for feedback on their work through a dedicated application. The request includes the work data (image data and audio data) and the user ID.
[0112] Server processing: The server receives the request sent by the user and saves the work data. The server also references the database based on the user ID to check the state of the user's sensory organs. For example, if the user is visually impaired, the server uses this information to send a request to the generation AI to convert the work data into audio data.
[0113] Processing by the generation AI: The generation AI analyzes the artwork data received from the server and converts it into the specified format. For visually impaired users, it converts image data into audio data, mapping color and shape information as sound. This converted data is sent back to the server as an audio file.
[0114] Server reprocessing: The server sends the converted data to the user's device. The user then experiences the converted data through the device and inputs their feedback. The feedback sent from the device is received and analyzed by the server. If necessary, a request is sent again to the generation AI for additional conversion or improvement.
[0115] Specific examples
[0116] If a visually impaired user wants feedback on a drawing they have made (an image file), the specific steps are as follows:
[0117] 1. The user uses the terminal to send an image file of the painting to the server.
[0118] 2. The server receives the image file and verifies that the user is visually impaired.
[0119] 3. The server sends a request to the AI to convert the image file into audio data. The prompt used is as follows:
[0120] > "Convert the following image data into audio data. The audio data should include information about the colors and shapes contained in the image."
[0121] 4. The generative AI maps the image's color and shape to sound and returns the result to the server as an audio file.
[0122] 5. The server sends the audio file to the user's terminal, which plays the audio file and provides it to the user.
[0123] 6. The user listens to the audio data and enters their impressions and areas for improvement as feedback, which the device then sends to the server.
[0124] 7. The server analyzes the feedback and, if necessary, sends another request to the generating AI for additional transformations or refinements, and provides the final result to the user in text format.
[0125] This allows users with visual or hearing impairments to experience their creations through other senses and receive detailed feedback.
[0126] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0127] Step 1: Receiving the request
[0128] Input: A user sends a request for feedback on a work through a dedicated application. The request includes the work data (image files and audio files) and the user ID.
[0129] How it works: The device accepts the user's request and sends an HTTP request to the server.
[0130] Output: The request data is sent to the server.
[0131] Step 2: Obtaining the work data
[0132] Input: The server receives the HTTP request sent from the device.
[0133] How it works: The server extracts the artwork data and user ID from the request data and saves the image and audio files to cloud storage.
[0134] Output: The work data is saved to cloud storage, and the user ID is saved to a database.
[0135] Step 3: Check the condition of your sensory organs
[0136] Input: The server retrieves the stored user ID.
[0137] How it works: The server checks the database to determine the user's sensory status (visual or hearing impairment), then executes an SQL query to retrieve the user information.
[0138] Output: The state data of the user's sensory organs is obtained.
[0139] Step 4: Calling the generation AI
[0140] Input: Artwork data and data on the state of the user's sensory organs are collected.
[0141] How it works: The server sends a request to the generation AI, including a prompt to convert the artwork data into a new format. An example prompt is "Convert the following image data into audio data. The audio data should include information about the colors and shapes contained in the image."
[0142] Output: A request containing the prompt is sent to the generation AI.
[0143] Step 5: Performing sensory transformations
[0144] Input: The generation AI receives the work data and prompt text sent from the server.
[0145] How it works: Generative AI uses a specified algorithm to convert, for example, image data into audio data, mapping color and shape information to sound.
[0146] Output: The converted data (e.g. audio data) is generated and returned to the server.
[0147] Step 6: Submit the conversion data
[0148] Input: The server receives the transformation data from the generation AI.
[0149] Operation: The server sends the converted data to the user's device using an HTTP response.
[0150] Output: The converted data is sent to the user's terminal.
[0151] Step 7: Provide transformation data
[0152] Input: The device receives the conversion data sent from the server.
[0153] What it does: The device plays or displays the converted data, presenting it as sound for visually impaired users and as vibrations or visual waveforms for hearing impaired users.
[0154] Output: The user experiences the transformed data.
[0155] Step 8: Gather feedback
[0156] Input: The user enters feedback on the transformation data.
[0157] How it works: The device accepts user feedback and sends it to the server using an HTTP request.
[0158] Output: The feedback data is sent to the server.
[0159] Step 9: Analyze feedback
[0160] Input: The server receives the feedback data sent from the device.
[0161] How it works: The server analyzes the feedback data and uses text analysis algorithms to extract user opinions and improvement requests.
[0162] Output: Analysis results are generated.
[0163] Step 10: Provide final feedback
[0164] Input: Feedback analysis results are complete.
[0165] How it works: The server generates feedback based on the analysis results and sends the final analysis results to the user's device, usually in text format.
[0166] Output: Final feedback is displayed on the user's device.
[0167] This allows users with visual or hearing impairments to experience their creations through other sensory organs and receive feedback.
[0168] (Application example 1)
[0169] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0170] It is difficult for users who lack certain sensory organs to experience their own artwork through other sensory organs. For example, when a visually impaired person evaluates their own painting, they cannot experience the artwork visually and must obtain feedback through other means. Similarly, when a hearing impaired person evaluates a musical piece, they need a way to experience it through vision or vibration patterns. However, these methods are not currently provided efficiently or effectively. Furthermore, there is no system that can properly analyze user feedback and convert and optimize artwork data using regenerative AI.
[0171] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0172] In this invention, the server includes means for acquiring work data, means for calling a generation AI that converts the work data into another format based on the state of the user's sensory organs, means for setting the state of the sensory organs, means for saving the sensory organ state information, means for using a generation AI that converts the work data into a corresponding format using a prompt sentence, means for providing the converted data to the user, means for receiving feedback from the user, means for analyzing the feedback, means for again calling the generation AI based on the analysis result to perform additional conversion, and means for providing the analysis result to the user. This enables users who are deficient in certain sensory organs to experience their own work in a different sensory format and receive accurate feedback.
[0173] "Work data" refers to digital files that represent content created by a user, and includes image files, audio files, video files, and the like.
[0174] "Means of acquisition" refers to the technical means by which a user uses a device to send and store work data on a server, and the server receives that data.
[0175] "Generative AI" is an artificial intelligence system that converts a user's work data into a different sensory format, for example, one that has the ability to convert image data into audio data.
[0176] The "calling means" refers to the communication protocol means by which the server sends a conversion request to the generating AI, and the generating AI performs data conversion based on that request.
[0177] "Sensory organ status" refers to information on whether the user's sensory organs, such as vision and hearing, are normal or missing, and serves as the condition for the generation AI to convert data based on this information.
[0178] A "prompt" is an instruction statement that lists specific requests to the generating AI, and is generated based on the work data and the state of the sensory organs.
[0179] "Means for converting" refers to the algorithms and processing system that the generative AI uses to convert the work data into another sensory format based on the prompt text.
[0180] The "means for providing" refers to a communication means and a user interface for sending the converted data to the user's terminal and allowing the user to experience it.
[0181] "Feedback" refers to the opinions and evaluations provided by users after experiencing the converted data, and is important data for the system to use to make improvements.
[0182] The "means of analysis" refers to an analysis system that has the function of collecting feedback from users, analyzing it, and evaluating the accuracy and quality of the conversion by the generation AI.
[0183] "Means for performing additional conversion" refers to technical means for sending a request to the generation AI again based on the analysis results to further optimize and improve the format of the work data.
[0184] This invention is a system that allows users who lack certain sensory organs to experience their own artwork through other sensory organs. This system is primarily composed of a server, a terminal, and a generative AI model. A detailed explanation of how to implement this system is provided below.
[0185] System configuration and hardware / software
[0186] This system uses smartphones and head-mounted displays (HoloLens 2, Oculus Quest 2, etc.) as terminals. AWS EC2 and Heroku are used as servers, and AWS RDS and Firebase are used as databases. OpenAI's GPT-4 is used as the generative AI model.
[0187] Obtaining work data
[0188] The user uses a device to send their creation data to the server. The creation data includes image files, audio files, video files, etc. The user's creation is then stored in digital format on the server.
[0189] Setting and saving the state of the sensory organs
[0190] The user sets the state of their sensory organs via the terminal. This information is sent to the server and stored in the user profile. For example, if the user is visually impaired, that information is stored and used for later processing.
[0191] Conversion process by generative AI
[0192] The server creates a request to send the submitted artwork data and sensory organ status information to the generation AI. This request is generated as a prompt, and the generation AI converts the data based on it. Examples of prompts include the following:
[0193] Please convert the following work (image data) into audio data so that it can be understood by visually impaired people: [Work data]
[0194] Data conversion and provisioning
[0195] Based on the prompt, the AI converts the artwork data into a format that matches the user's sensory state. For example, it converts image data into audio data and returns the result as an audio file to the server. The server then sends the audio data to the device, allowing the user to experience it.
[0196] User feedback and analysis
[0197] Users experience the converted data and enter their evaluations and opinions (feedback) through their devices. The feedback is sent to the server, which analyzes it. The analysis includes whether the feedback is positive or negative, and points for improvement.
[0198] Additional transformations and final feedback
[0199] The server analyzes the feedback and sends another request to the generation AI based on the results. The generation AI performs additional transformations and optimizations and returns the results to the server. This process is repeated until the final feedback is provided to the user.
[0200] Specific examples
[0201] A specific example of this system is shown below.
[0202] 1. Data upload: The user takes a photo of the painting with their smartphone and uploads it to the app.
[0203] 2. Sensory settings: The user sets "visual impairment" and the information is saved on the server.
[0204] 3. Prompt generation: The server generates a prompt for the AI to convert the image data into audio data.
[0205] 4. Data conversion: The generative AI converts image data into audio data.
[0206] 5. Feedback and refinement: The user listens to the audio data and provides feedback. The server analyzes the feedback and, if necessary, converts and refines the data again.
[0207] In this way, users who lack certain sensory organs can experience their creations in an alternative sensory format and receive accurate feedback.
[0208] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0209] Step 1:
[0210] The user uploads the work data using the device.
[0211] Input: Work data (image files, audio files, etc.) stored on the user's device.
[0212] Output: The work data is uploaded to the server and stored in the server database.
[0213] Specific operation: The user uses a smartphone or head-mounted display to select the artwork data through the app and press the upload button. The uploaded data is sent to the server via an HTTP request, and the server stores it in a database.
[0214] Step 2:
[0215] The user sets the state of the sensory organs using the terminal.
[0216] Input: Sensory status (e.g., visual impairment, hearing impairment) set by the user in the device application.
[0217] Output: The sensory organ status information is sent to the server and stored in a database.
[0218] What happens: The user selects the state of their sensory organs within the application and presses the settings button. This information is sent to the server and saved in the user profile.
[0219] Step 3:
[0220] The server generates a prompt for the generated AI.
[0221] Input: Uploaded artwork data and sensory organ status information.
[0222] Output: The prompt to send to the generation AI.
[0223] Specific operation: The server obtains the artwork data and the state of the sensory organs, and generates a prompt based on this. For example, a prompt that converts image data into audio data is created: "Please convert the following artwork (image data) into audio data so that it can be understood by visually impaired people: [artwork data]."
[0224] Step 4:
[0225] The server sends a prompt to the generation AI, requesting data conversion.
[0226] Input: The generated prompt statement.
[0227] Output: Data converted by the generating AI (e.g., audio data).
[0228] Specific operation: The server sends the generated prompt to the generation AI via the OpenAI API. The generation AI performs data conversion based on the prompt.
[0229] Step 5:
[0230] The generation AI returns the converted data to the server.
[0231] Input: Data converted by the generation AI based on the prompt.
[0232] Output: The transformed data is returned to the server.
[0233] Specific operation: The generation AI generates the conversion result and returns it to the server as an HTTP response. The server receives this data and stores it in a database as needed.
[0234] Step 6:
[0235] The server transmits the converted data to the user's terminal.
[0236] Input: Transformation data returned from the generation AI.
[0237] Output: The converted data is displayed and played on the user's device.
[0238] Specific operation: The server obtains the converted data and sends it to the user's device as an HTTP response. The user's device receives the data and displays or plays it in the appropriate format.
[0239] Step 7:
[0240] The user inputs feedback on the converted data.
[0241] Input: The transformation data experienced by the user.
[0242] Output: The feedback data is sent to the server.
[0243] Specific operation: The user uses the device to input feedback about the conversion data they have experienced. Once the input is complete, it is sent to the server and saved.
[0244] Step 8:
[0245] The server analyzes the feedback and requests additional transformations from the generation AI if necessary.
[0246] Input: User feedback data.
[0247] Output: A new prompt statement requesting additional transformations.
[0248] Specific operation: The server analyzes the feedback data and determines whether improvements are necessary. If so, it generates a new prompt for additional conversion and sends it to the generation AI.
[0249] Step 9:
[0250] The generation AI performs additional transformations and returns the results to the server.
[0251] Input: New prompt statement.
[0252] Output: Additional conversion result data.
[0253] Specific operation: The generation AI performs further optimized data conversion based on the new prompt sentence, and sends the conversion result back to the server.
[0254] Step 10:
[0255] The server sends the final feedback data to the user's terminal.
[0256] Input: Analysis results and additional transformation data.
[0257] Output: The final feedback data is provided to the user's terminal.
[0258] What it does: The server compiles the final analysis results and additional conversion data and sends them to the device in the most useful format for the user, allowing the user to re-experience the data and check the final feedback.
[0259] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0260] This invention combines a system that allows users who lack a particular sensory organ to experience their own work through another sensory organ with an emotion engine that recognizes the user's emotions. Specific embodiments of this system will be described below.
[0261] 1. System Overview
[0262] This system consists of a server, a terminal, a generative AI, and an emotion engine.
[0263] Users use their devices to send their creation data to the server and receive feedback. The emotion engine also recognizes the user's emotions and reflects them in the system's processing.
[0264] 2. Receiving a Request
[0265] A user sends a request from a terminal for feedback on their creation (e.g., a painting or a song).
[0266] 3. Acquiring work data
[0267] The server receives the work data sent from the terminal and stores, for example, image files and audio files appropriately.
[0268] 4. Check the condition of the sensory organs
[0269] The server consults a database to ascertain the user's sensory status, for example to determine whether the user is visually impaired or hearing impaired.
[0270] 5. Calling the generation AI
[0271] The server sends a conversion request to the generation AI based on the acquired artwork data. This request includes the artwork data to be converted and the state of the user's sensory organs.
[0272] 6. Emotional Engine Activation
[0273] The server activates an emotion engine to recognize the user's emotions in real time. The emotion engine acquires emotion data from the user's facial expressions and voice.
[0274] 7. Implementing sensory transformation
[0275] The generative AI converts the received artwork data into a format that corresponds to the state of the user's sensory organs. For example, when converting images into audio data, it maps colors and shapes to pitch and rhythm.
[0276] 8. Emotion-Based Regulation
[0277] The server adjusts the generated conversion data based on the emotion data obtained from the emotion engine. For example, if the user is surprised or sad, the server adjusts the conversion data to take those emotions into consideration.
[0278] 9. Sending the converted data
[0279] The generated AI sends the converted data (e.g., audio files, vibration data) back to the server, where it is appropriately encoded.
[0280] 10. Provision of conversion data
[0281] The server sends the converted data to the terminal, encoding the data as an HTTP response.
[0282] The device presents the received converted data to the user, playing audio data for visually impaired users and displaying vibrations or visual waveforms for hearing impaired users.
[0283] 11. Collecting Feedback
[0284] The user inputs feedback on the presented conversion data through the terminal, including their impressions and suggestions for improvement.
[0285] The emotion engine also collects emotional data while users are entering feedback, making it clear what emotions are behind the feedback provided.
[0286] 12. Feedback Analysis
[0287] The server analyzes the received feedback using natural language processing technology to extract important points and suggestions for improvement.
[0288] Emotion data collected from the emotion engine is also used in the analysis to evaluate how the user's emotional state influenced the feedback.
[0289] 13. Recall of generated AI
[0290] Based on the feedback and the results of the analysis of the emotional data, the server sends another request to the generative AI to perform additional transformations and improvements.
[0291] 14. Providing Final Feedback
[0292] The server generates final feedback based on the analysis results and sends it to the terminal in text format.
[0293] The device will then present the final feedback to the user, allowing them to see how their work can be improved and the details of the feedback.
[0294] As a concrete example, consider the case where a visually impaired user requests feedback on a painting (image file) they have created. The user uses their device to send the image file of the painting to a server. The server receives the image file and confirms that the user is visually impaired. The server then sends a request to the generation AI to convert the image file into audio data. Meanwhile, it activates an emotion engine and monitors the user's emotions in real time. The generation AI maps the image's color and shape to sound and returns the result to the server as an audio file. The server adjusts the audio data based on the emotional data obtained from the emotion engine and sends the audio file to the user's device. The device plays the audio file and provides it to the user. The user listens to the audio data and enters their impressions and suggestions for improvement as feedback, and the emotion engine also collects their emotions at this time. The server analyzes the feedback and emotional data and sends another request to the generation AI to make additional conversions or improvements. Finally, the server provides the analysis results in text format to the user, allowing them to check the improvements to the work and details of the feedback.
[0295] This allows visually or hearing impaired users to experience the work through other sensory organs, and by taking into account their emotions in the process, they can receive richer and more accurate feedback.
[0296] The processing flow will be explained below.
[0297] Step 1:
[0298] A user uses a terminal to send a request for feedback on his / her work data (e.g., an image file or an audio file), and the terminal sends the work data along with the request to the server.
[0299] Step 2:
[0300] The server receives the request and work data sent from the device. The server analyzes the HTTP request, obtains the attached file data, and saves it.
[0301] Step 3:
[0302] The server accesses a database to check the user's sensory status, specifically to determine whether the user has a visual or hearing impairment.
[0303] Step 4:
[0304] The server creates a request to send the acquired artwork data to the generation AI. This request includes the artwork data to be converted and the state of the user's sensory organs.
[0305] Step 5:
[0306] The server sends a request to the generation AI, asking it to convert the artwork data into a format that corresponds to the state of the user's sensory organs.
[0307] Step 6:
[0308] The generative AI then converts the received artwork data. For example, for visually impaired users, it converts image data into audio data. In this process, it maps colors and shapes to pitch and rhythm.
[0309] Step 7:
[0310] The generation AI sends the converted data back to the server, where it becomes an audio file, vibration pattern, etc.
[0311] Step 8:
[0312] The server runs an emotion engine to monitor the user's emotions in real time, analyzing facial expressions and voices while the user is experiencing the device.
[0313] Step 9:
[0314] The server receives the converted data returned by the generation AI and makes adjustments based on the emotional data from the emotion engine, adjusting the speed of the voice data or changing the vibration pattern depending on the user's emotional state.
[0315] Step 10:
[0316] The server sends the adjusted converted data to the terminal, which receives the data as an HTTP response and presents it to the user.
[0317] Step 11:
[0318] The terminal provides the converted data to the user, playing it as audio for visually impaired users, or displaying vibrations or visual waveforms for hearing impaired users.
[0319] Step 12:
[0320] The user inputs feedback on the presented data through the terminal, including their impressions of the presented data and suggestions for improvement.
[0321] Step 13:
[0322] The emotion engine collects the emotions of users who enter feedback in real time by analyzing their facial expressions and tone of voice to obtain emotional data.
[0323] Step 14:
[0324] The device sends the user's feedback and emotional data to the server. The feedback is sent in text format, and the emotional data is sent in numerical and categorical format.
[0325] Step 15:
[0326] The server analyzes the received feedback and sentiment data, and uses natural language processing technology to analyze the feedback content and extract key points and improvement suggestions.
[0327] Step 16:
[0328] The server then sends a request to the generative AI again based on the feedback and emotion data to perform additional transformations and refinements.
[0329] Step 17:
[0330] The generation AI performs additional transformations and refinements and sends the results back to the server, which receives the results.
[0331] Step 18:
[0332] The server generates final feedback based on the analysis results and sends it to the terminal in text format.
[0333] Step 19:
[0334] The device will then present the final feedback to the user, allowing them to see how their work can be improved and the details of the feedback.
[0335] Example 2
[0336] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0337] The challenge is to provide a system that allows users who lack certain sensory organs to experience their own work through other sensory organs and receive feedback that takes into account their emotions during the process.
[0338] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0339] In this invention, the server includes means for acquiring work data, means for calling a generation AI that converts the acquired work data into another format based on the state of the user's sensory organs, means for recognizing the user's emotions in real time and acquiring emotional data, means for adjusting the converted data based on the acquired emotional data, means for providing the adjusted converted data to the user, means for receiving feedback from the user and collecting the user's emotional data during that time, means for analyzing the feedback and emotional data, means for calling the generation AI again based on the analysis results and emotional data to perform additional conversion, and means for providing the analysis results to the user. This enables users with visual or hearing impairments to experience the work through their other sensory organs and receive rich and accurate feedback that takes into account the emotions experienced during the process.
[0340] "Work data" refers to content created by users, and refers to digital data such as image files and audio files.
[0341] "Sensory organ status" refers to information indicating the health or absence of impairment of the user's vision, hearing, or other senses.
[0342] "Generative AI" refers to artificial intelligence techniques that transform input data into a different format.
[0343] An "emotion engine" refers to a system that recognizes a user's emotions in real time and collects emotional data from facial expressions and voice.
[0344] "Emotion data" refers to data that indicates the user's emotional state obtained by the emotion engine.
[0345] "Converted data" refers to the new data format converted from the original work data by the generating AI.
[0346] "Feedback" refers to the impressions and evaluations that users provide about the sensory-converted work.
[0347] "Analysis results" refers to the results of the analysis performed by the server based on the feedback and emotional data received from the user.
[0348] "Re-invoking" refers to invoking the generation AI again after the initial invocation to perform additional transformations or improvements.
[0349] "Another format" refers to a data format that can be experienced through a different sensory organ than the original work data (e.g., converting an image into sound).
[0350] This invention combines a system that allows users who lack a particular sensory organ to experience their own work through another sensory organ with an emotion engine that recognizes the user's emotions. Specific embodiments of this system will be described below.
[0351] 1. System Overview
[0352] This system consists of a server, a terminal, a generation AI, and an emotion engine. Users use their terminal to send their creation data to the server and receive feedback. The emotion engine also recognizes the user's emotions and reflects them in each process of the system.
[0353] 2. Hardware and Software Configuration
[0354] Server: A computer system with a powerful processor and large memory capacity that runs the emotion engine and generative AI.
[0355] Device: A device used by a user, such as a smartphone, tablet, or computer, that is connected to the internet.
[0356] Generative AI: Artificial intelligence that uses deep learning models to perform specific transformations, such as converting visual data into audio data.
[0357] Emotion engine: Software that analyzes the user's facial expressions and voice to obtain emotional data in real time.
[0358] 3. Specific Examples
[0359] As a concrete example, consider the case where a visually impaired user requests feedback on a painting (image file) they have created. The user uses their device to send the image file of the painting to a server. The server receives the image file and confirms that the user is visually impaired. The server then sends a request to the generation AI to convert the image file into audio data. Meanwhile, it activates an emotion engine and monitors the user's emotions in real time. The generation AI maps the image's color and shape to sound and returns the result to the server as an audio file. The server adjusts the audio data based on the emotional data obtained from the emotion engine and sends the audio file to the user's device. The device plays the audio file and provides it to the user. The user listens to the audio data and enters their impressions and suggestions for improvement as feedback, and the emotion engine also collects their emotions at this time. The server analyzes the feedback and emotional data and sends another request to the generation AI to make additional conversions or improvements. Finally, the server provides the analysis results in text format to the user, allowing them to check the improvements to the work and details of the feedback.
[0360] 4. Examples of prompts
[0361] "We have an image file of a painting drawn by a user. Based on this image file, please convert it into audio data by mapping the colors and shapes to pitch and rhythm. The user is visually impaired."
[0362] This allows users who lack certain sensory organs to experience their own work through other sensory organs and receive rich feedback that takes into account the emotions based on that experience.
[0363] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0364] Step 1:
[0365] A user uses a terminal to send his / her own artwork data (e.g., an image file of a painting) to the server. The artwork data provided by the user is the input, and the artwork data received by the server is the output.
[0366] Step 2:
[0367] The server receives the artwork data sent from the device, converts it into the appropriate format, and saves it. As part of the data processing, it checks the format of the sent data and saves it in the appropriate storage. The saved artwork data is obtained as output.
[0368] Step 3:
[0369] The server references the database to check the status of the user's sensory organs. The input is identification information such as the user's ID, and the server uses this data to calculate whether or not the sensory organs are impaired. The output is information about the status of the sensory organs.
[0370] Step 4:
[0371] The server sends a conversion request to the generation AI based on the acquired artwork data and information about the user's sensory organs. The inputs include artwork data and information about the state of the sensory organs, and a prompt is generated and sent to the generation AI. For example, a prompt such as "We have an image file of a painting drawn by the user. Based on this image file, please map the color and shape to pitch and rhythm and convert it into audio data" is sent. The output is the data resulting from the conversion by the generation AI (for example, an audio file).
[0372] Step 5:
[0373] The server runs an emotion engine and monitors the user's emotions in real time. The input is the user's facial expressions and voice collected from sensors such as cameras and microphones, and the output is the user's emotional data acquired in real time.
[0374] Step 6:
[0375] The generation AI converts the artwork data it receives into a format that corresponds to the state of the user's sensory organs. For example, data processing involves mapping the color of an image to the pitch of a sound, and the shape to the rhythm of a sound. The converted data (for example, an audio file) is obtained as output.
[0376] Step 7:
[0377] The server adjusts the generated converted data based on the emotional data obtained from the emotion engine. The converted data and emotional data are input, and adjustments are made to take the emotion into consideration through data calculations. For example, if the user is surprised or sad, the tone of the sound may be changed. The adjusted converted data is obtained as output.
[0378] Step 8:
[0379] The server sends the adjusted transformed data to the terminal. The input is the adjusted transformed data, and the output is the data encoded as an HTTP response.
[0380] Step 9:
[0381] The device presents the received converted data to the user. The input is the converted data, and the specific action is to play audio data or transmit a vibration pattern. The output allows the user to experience the work with different sensory organs.
[0382] Step 10:
[0383] The user inputs feedback on the presented converted data through the terminal. At this time, the emotion engine also collects the user's emotions. The input is the user's feedback and the emotional data during that time, and the output is the collected feedback and emotional data.
[0384] Step 11:
[0385] The server analyzes the received feedback and emotional data. The collected feedback and emotional data are input, and natural language processing technology is used to analyze the feedback and extract important points and improvement suggestions. The analysis results are obtained as output.
[0386] Step 12:
[0387] The server then sends a request to the AI generator based on the analysis results to perform additional transformations and improvements. The input is the analysis results, and the output is the regenerated transformation data.
[0388] Step 13:
[0389] The server generates the final feedback and sends it to the terminal in text format. The input is the final analysis result, and the output is the textual feedback provided to the user.
[0390] (Application example 2)
[0391] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0392] In previous systems, users with sensory impairments lacked the means to experience their artwork through other senses, and the feedback they received did not take into account the user's emotions. This resulted in low user satisfaction and the quality of their experience with the artwork. Furthermore, the lack of sufficient feedback collection and analysis made it difficult to further improve the generative AI.
[0393] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to upload artwork data to a terminal at a physical store, means for calling a generation AI that converts the acquired artwork data into another format based on the state of the user's sensory organs, means for providing the converted data to the user, means for activating an emotion engine that acquires emotional data from the user's facial expressions and voice, means for adjusting the converted data according to the emotional data, means for receiving feedback from the user, means for analyzing the feedback and emotional data, means for re-invoking the generation AI based on the analysis results to perform additional conversion, and means for providing the analysis results to the user. This enables users with sensory organ deficiencies to experience their own artwork through their other sensory organs and receive emotional feedback.
[0394] "Means for users to upload artwork data to a device at a physical store" refers to a means for users to photograph or record artworks exhibited at a store they actually visit using a device such as a smartphone or tablet, and then send that data to an online server.
[0395] "Means for invoking a generation AI that converts acquired artwork data into a different format based on the state of the user's sensory organs" refers to a means by which the server launches a generation AI that automatically converts images into an appropriate format, such as converting them into audio, based on information about the state of the user's sensory organs, such as visual or hearing impairments.
[0396] The "means for providing the converted data to the user" refers to a means for transmitting the converted voice data, vibration pattern, or other data to the user's terminal and reading it out loud or notifying it by vibration.
[0397] The "means for activating an emotion engine that acquires emotion data from the user's facial expressions and voice" refers to a means by which the server activates an emotion recognition system that analyzes the user's real-time facial expressions and voice tone to grasp their emotional state.
[0398] The "means for adjusting the converted data in accordance with the emotional data" refers to a means for changing the tone of the voice of the converted data or adjusting the strength of the vibration pattern based on the acquired emotional data.
[0399] The "means for receiving feedback from the user" refers to a means by which the server provides an interface for the user to input their impressions and suggestions for improvement after experiencing the converted data.
[0400] The "means for analyzing feedback and emotional data" refers to a means for analyzing the emotional data acquired simultaneously with the received feedback content, and for deeply understanding and analyzing the user's experiences and opinions.
[0401] "Means of calling the generation AI again based on the analysis results to perform additional conversion" refers to means of starting the generation AI again based on information obtained from the analysis results to generate more accurate conversion data.
[0402] "Means for providing analysis results to users" refers to means for providing the results of the analyzed feedback to users in an easy-to-understand format and suggesting improvements to the work or new experiences.
[0403] This invention combines a system that allows users who lack certain sensory organs to experience their own artwork through other sensory organs with an emotion engine that recognizes the user's emotions. The system consists of a server, a terminal, a generative AI, and an emotion engine.
[0404] First, the user takes a photo or records a piece of artwork on display in a physical store using a device such as a smartphone or tablet. Next, the user uses the device to upload the artwork data to a server. The server stores the received artwork data and retrieves information about the user's sensory organ status (e.g., visual or hearing impairment) from a database.
[0405] The acquired artwork data is converted by the generative AI into an appropriate format based on the state of the user's sensory organs. For example, image files are converted into audio data, and audio files are converted into vibration patterns or visual waveforms. At this time, the server activates an emotion engine to obtain emotional data in real time from the user's facial expressions and voice. The emotional data is reflected in the conversion process by the generative AI, and the converted data is adjusted based on the user's emotional state.
[0406] The converted data is sent from the server to the device and provided to the user. The user experiences the audio and vibration data and sends feedback about the experience to the server via the device. This feedback includes their impressions and areas for improvement, and also includes emotional data collected by the emotion engine.
[0407] The server analyzes the received feedback and emotional data, and sends a re-request to the generation AI to further transform and improve the work data. Finally, the improved data generated based on the analysis results is provided to the user, allowing the user to check the improvements to the work and the details of the feedback.
[0408] As a concrete example, let's say a visually impaired user visits an art museum and takes a photo of a painting. In this case, the app converts the painting into audio and adjusts the audio data while checking the user's emotions using an emotion engine. The user can then provide feedback on the audio they heard, which is then further analyzed by the AI generator and provided to the user as optimized audio data.
[0409] An example of a prompt for a generative AI model is:
[0410] "Text-to-speech prompts for the visually impaired:
[0411] 1. Convert the drawn image into audio data.
[0412] 2. Map colors and shapes to pitch and rhythm.
[0413] 3. Adjust the voice depending on the user's emotions.
[0414] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0415] Step 1:
[0416] A user photographs or records artwork data (e.g., paintings or sounds) in a physical store using a device and uploads it to the server. The input here is the photographed or recorded artwork data file. The device sends this data to the server as an HTTP request. The output is the artwork data file, which is saved on the server.
[0417] Step 2:
[0418] The server stores the received artwork data and retrieves information about the user's sensory organ status (visual impairment, hearing impairment, etc.) from a database. The uploaded artwork data and user ID are used as input. Based on this, the server references the user's profile and retrieves the status of the sensory organs. The output is the status of the user's sensory organs.
[0419] Step 3:
[0420] The server calls the generative AI to convert the acquired artwork data into an appropriate format (sound, vibration pattern, visual waveform, etc.). The inputs are the artwork data and the state of the user's sensory organs. The generative AI model converts the data based on this information. The converted data (sound file, vibration pattern, etc.) is generated as output.
[0421] Step 4:
[0422] The server activates the emotion engine and analyzes the user's facial expressions and voice while experiencing the converted data to obtain emotion data. The input is the user's real-time facial video and voice data. The emotion engine analyzes this data to determine the user's emotional state. The output is emotion data.
[0423] Step 5:
[0424] The server adjusts the conversion data based on the acquired emotional data. For example, the tone of the voice or the intensity of the vibrations is changed according to the user's emotion. The conversion data and emotional data are used as input. The server adjusts the data appropriately based on this. The output is conversion data adjusted according to the emotion.
[0425] Step 6:
[0426] The server sends the adjusted transformed data to the user's device. The adjusted transformed data is used as input. The server sends this as an HTTP response to the device. The transformed data is received as output by the user's device. The user experiences the data through the device.
[0427] Step 7:
[0428] The user experiences the provided converted data and sends feedback about the experience from the device to the server. The user's feedback content is used as input. The device sends this as an HTTP request to the server. The feedback data is saved on the server as output.
[0429] Step 8:
[0430] The server analyzes the received feedback and emotion data and sends a re-request to the generation AI to perform additional transformations and improvements to the work data. The feedback data and emotion data are used as input. The server analyzes this data and sends new instructions to the generation AI. The improved transformation data is generated as output.
[0431] Step 9:
[0432] The server generates the final feedback content based on the analysis results and sends it to the terminal. The parsed feedback data is used as input. The server converts it into text format and sends it to the terminal as an HTTP response. The final feedback is displayed on the user's terminal as output.
[0433] This allows users to experience the artwork and receive emotional feedback regardless of sensory deficiencies.
[0434] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0435] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0436] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0437] [Second embodiment]
[0438] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0439] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0440] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0441] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0442] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0443] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0444] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0445] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0446] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0447] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0448] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0449] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0450] The present invention provides a system that allows users who lack a particular sensory organ to experience their own artwork through another sensory organ. Specific embodiments of this system will be described below.
[0451] 1. System Overview
[0452] This system consists of a server, a terminal, and a generation AI.
[0453] Users use their terminals to send their creation data to the server and receive feedback.
[0454] 2. Receiving a Request
[0455] A user sends a request from a terminal for feedback on their creation (e.g., a painting or a song).
[0456] 3. Acquiring work data
[0457] The server receives the work data sent from the terminal and stores, for example, image files and audio files appropriately.
[0458] 4. Check the condition of the sensory organs
[0459] The server consults a database to ascertain the user's sensory status, for example to determine whether the user is visually impaired or hearing impaired.
[0460] 5. Calling the generation AI
[0461] The server sends the acquired artwork data to the generation AI and sends a request to convert it into a different format depending on the state of the user's sensory organs.
[0462] 6. Implementing sensory transformation
[0463] The generative AI converts the received artwork data into a new format as specified, for example converting image data into audio data for visually impaired users, or audio data into vibration patterns or visual waveforms for hearing impaired users.
[0464] 7. Sending the converted data
[0465] The generation AI returns the converted data to the server, which then sends the data to the user's device.
[0466] 8. Providing conversion data
[0467] The device presents the converted data to the user, either as audio for visually impaired users or as vibrations or visual waveforms for hearing impaired users.
[0468] 9. Collecting Feedback
[0469] The user inputs feedback on the converted data through the terminal, and the terminal transmits the feedback to the server.
[0470] 10. Feedback Analysis
[0471] The server analyzes the received feedback and, if necessary, sends another request to the generating AI to perform additional transformations or improvements.
[0472] 11. Providing Final Feedback
[0473] The server generates feedback based on the analysis results, sends the final analysis results to the terminal, and provides the final feedback to the user.
[0474] As a concrete example, consider the case where a visually impaired user requests feedback on a painting (image file) they have created. The user uses their device to send the image file of the painting to a server. The server receives the image file and confirms that the user is visually impaired. The server then sends a request to the generation AI to convert the image file into audio data. The generation AI maps the image's colors and shapes to sounds and returns the result to the server as an audio file. The server then sends the audio file to the user's device, which plays it and provides it to the user. The user listens to the audio data and enters their impressions and suggestions for improvement as feedback, which the device then sends to the server. The server analyzes the feedback and, if necessary, sends another request to the generation AI to make additional conversions or improvements. Finally, the server provides the analysis results to the user in text format. This process allows visually impaired users to experience their own work as audio and receive feedback.
[0475] This allows visually or hearing impaired users to experience the work through other sensory organs and receive accurate feedback.
[0476] The processing flow will be explained below.
[0477] Step 1:
[0478] A user uses a terminal to send a request for feedback on his / her work data (e.g., an image file or an audio file).
[0479] Step 2:
[0480] The server receives the request and work data sent from the terminal, analyzes the HTTP request, and obtains the contents of the attached file or data.
[0481] Step 3:
[0482] The server retrieves the user's sensory status from a database, including whether the user is visually or hearing impaired.
[0483] Step 4:
[0484] The server sends a conversion request to the generation AI based on the acquired artwork data. This request includes the artwork data to be converted and the state of the user's sensory organs.
[0485] Step 5:
[0486] The generative AI converts the received artwork data into a format that corresponds to the state of the user's sensory organs. For example, when converting images into audio data, it maps colors and shapes to pitch and rhythm.
[0487] Step 6:
[0488] The generated AI sends the converted data (e.g., audio files, vibration data) back to the server, where it is appropriately encoded.
[0489] Step 7:
[0490] The server sends the converted data to the terminal, encoding the data as an HTTP response.
[0491] Step 8:
[0492] The terminal presents the received converted data to the user, playing audio data for visually impaired users and displaying vibrations or visual waveforms for hearing impaired users.
[0493] Step 9:
[0494] The user inputs feedback on the presented conversion data through the terminal, including their impressions and suggestions for improvement.
[0495] Step 10:
[0496] The terminal transmits the feedback from the user to the server. The feedback data is sent to the server as an HTTP request.
[0497] Step 11:
[0498] The server analyzes the received feedback using natural language processing technology to extract important points and suggestions for improvement.
[0499] Step 12:
[0500] The server may send a request to the generation AI again based on the results of analyzing the feedback, to perform additional transformations or improvements.
[0501] Step 13:
[0502] The server generates final feedback based on the analysis results and sends it to the terminal in text format.
[0503] Step 14:
[0504] The device will then present the final feedback to the user, allowing them to see how their work can be improved and the details of the feedback.
[0505] Example 1
[0506] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0507] This solves the problem of users who lack certain sensory organs having difficulty experiencing their own work through other sensory organs and receiving appropriate feedback. Conventional technologies have limited the means by which visually or hearing-impaired users can obtain feedback on their own work, and no efficient system exists for this purpose.
[0508] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0509] In this invention, the server includes means for receiving a request for feedback on work data from a user, means for acquiring work data, means for saving the acquired work data, means for checking the state of the user's sensory organs, means for calling a generation AI that converts the acquired work data into another format based on the state of the user's sensory organs, means for sending a conversion request to the generation AI, means for acquiring the converted data, means for providing the converted data to the user, means for receiving feedback from the user, means for analyzing the feedback, means for again calling the generation AI based on the analysis result to perform additional conversion, and means for providing the analysis result to the user. This enables users with visual or hearing impairments to experience the work through their other sensory organs and receive accurate feedback based on the results.
[0510] "User" refers to a person who uses this system to submit work data and receive feedback.
[0511] "Work data" refers to digital content such as images and audio created by users.
[0512] "Feedback" refers to evaluations, opinions, and comments for improvement provided regarding work data.
[0513] "Status of sensory organs" refers to the functional status of the sensory organs held by the user, for example, whether or not there is a visual or hearing impairment.
[0514] "Generative AI" refers to artificial intelligence used to transform work data into another format.
[0515] "Conversion request" refers to a command that instructs the generation AI to convert work data into a different format.
[0516] "Converted data" refers to data in a new format that has been converted from the original work data by the generating AI.
[0517] "Server" refers to a computer system that receives requests from users, stores work data, and converts and provides the data in cooperation with the generative AI.
[0518] A "request" refers to a request made by a user to a server, for example, a request for feedback.
[0519] "Analysis" refers to the process by which the server analyzes the feedback obtained from the user.
[0520] This invention provides a system that allows users who lack a particular sensory organ to experience their own artwork through another sensory organ. Specific embodiments of this system are described below.
[0521] Hardware and software used
[0522] This system consists of a server, a terminal, and a generative AI. Specific hardware includes high-performance computers (e.g., AWS EC2 instances), and terminals include users' smartphones and tablets. Generative AI uses GPT-4, Stable Diffusion, DALL-E, and other algorithms.
[0523] System Operation Overview
[0524] The following is an overview of how this system works.
[0525] User operation: A user sends a request for feedback on their work through a dedicated application. The request includes the work data (image data and audio data) and the user ID.
[0526] Server processing: The server receives the request sent by the user and saves the work data. The server also references the database based on the user ID to check the state of the user's sensory organs. For example, if the user is visually impaired, the server uses this information to send a request to the generation AI to convert the work data into audio data.
[0527] Processing by the generation AI: The generation AI analyzes the artwork data received from the server and converts it into the specified format. For visually impaired users, it converts image data into audio data, mapping color and shape information as sound. This converted data is sent back to the server as an audio file.
[0528] Server reprocessing: The server sends the converted data to the user's device. The user then experiences the converted data through the device and inputs their feedback. The feedback sent from the device is received and analyzed by the server. If necessary, a request is sent again to the generation AI for additional conversion or improvement.
[0529] Specific examples
[0530] If a visually impaired user wants feedback on a drawing they have made (an image file), the specific steps are as follows:
[0531] 1. The user uses the terminal to send an image file of the painting to the server.
[0532] 2. The server receives the image file and verifies that the user is visually impaired.
[0533] 3. The server sends a request to the AI to convert the image file into audio data. The prompt used is as follows:
[0534] > "Convert the following image data into audio data. The audio data should include information about the colors and shapes contained in the image."
[0535] 4. The generative AI maps the image's color and shape to sound and returns the result to the server as an audio file.
[0536] 5. The server sends the audio file to the user's terminal, which plays the audio file and provides it to the user.
[0537] 6. The user listens to the audio data and enters their impressions and areas for improvement as feedback, which the device then sends to the server.
[0538] 7. The server analyzes the feedback and, if necessary, sends another request to the generating AI for additional transformations or refinements, and provides the final result to the user in text format.
[0539] This allows users with visual or hearing impairments to experience their creations through other senses and receive detailed feedback.
[0540] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0541] Step 1: Receiving the request
[0542] Input: A user sends a request for feedback on a work through a dedicated application. The request includes the work data (image files and audio files) and the user ID.
[0543] How it works: The device accepts the user's request and sends an HTTP request to the server.
[0544] Output: The request data is sent to the server.
[0545] Step 2: Obtaining the work data
[0546] Input: The server receives the HTTP request sent from the device.
[0547] How it works: The server extracts the artwork data and user ID from the request data and saves the image and audio files to cloud storage.
[0548] Output: The work data is saved to cloud storage, and the user ID is saved to a database.
[0549] Step 3: Check the condition of your sensory organs
[0550] Input: The server retrieves the stored user ID.
[0551] How it works: The server checks the database to determine the user's sensory status (visual or hearing impairment), then executes an SQL query to retrieve the user information.
[0552] Output: The state data of the user's sensory organs is obtained.
[0553] Step 4: Calling the generation AI
[0554] Input: Artwork data and data on the state of the user's sensory organs are collected.
[0555] How it works: The server sends a request to the generation AI, including a prompt to convert the artwork data into a new format. An example prompt is "Convert the following image data into audio data. The audio data should include information about the colors and shapes contained in the image."
[0556] Output: A request containing the prompt is sent to the generation AI.
[0557] Step 5: Performing sensory transformations
[0558] Input: The generation AI receives the work data and prompt text sent from the server.
[0559] How it works: Generative AI uses a specified algorithm to convert, for example, image data into audio data, mapping color and shape information to sound.
[0560] Output: The converted data (e.g. audio data) is generated and returned to the server.
[0561] Step 6: Submit the conversion data
[0562] Input: The server receives the transformation data from the generation AI.
[0563] Operation: The server sends the converted data to the user's device using an HTTP response.
[0564] Output: The converted data is sent to the user's terminal.
[0565] Step 7: Provide transformation data
[0566] Input: The device receives the conversion data sent from the server.
[0567] What it does: The device plays or displays the converted data, presenting it as sound for visually impaired users and as vibrations or visual waveforms for hearing impaired users.
[0568] Output: The user experiences the transformed data.
[0569] Step 8: Gather feedback
[0570] Input: The user enters feedback on the transformation data.
[0571] How it works: The device accepts user feedback and sends it to the server using an HTTP request.
[0572] Output: The feedback data is sent to the server.
[0573] Step 9: Analyze feedback
[0574] Input: The server receives the feedback data sent from the device.
[0575] How it works: The server analyzes the feedback data and uses text analysis algorithms to extract user opinions and improvement requests.
[0576] Output: Analysis results are generated.
[0577] Step 10: Provide final feedback
[0578] Input: Feedback analysis results are complete.
[0579] How it works: The server generates feedback based on the analysis results and sends the final analysis results to the user's device, usually in text format.
[0580] Output: Final feedback is displayed on the user's device.
[0581] This allows users with visual or hearing impairments to experience their creations through other sensory organs and receive feedback.
[0582] (Application example 1)
[0583] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0584] It is difficult for users who lack certain sensory organs to experience their own artwork through other sensory organs. For example, when a visually impaired person evaluates their own painting, they cannot experience the artwork visually and must obtain feedback through other means. Similarly, when a hearing impaired person evaluates a musical piece, they need a way to experience it through vision or vibration patterns. However, these methods are not currently provided efficiently or effectively. Furthermore, there is no system that can properly analyze user feedback and convert and optimize artwork data using regenerative AI.
[0585] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0586] In this invention, the server includes means for acquiring work data, means for calling a generation AI that converts the work data into another format based on the state of the user's sensory organs, means for setting the state of the sensory organs, means for saving the sensory organ state information, means for using a generation AI that converts the work data into a corresponding format using a prompt sentence, means for providing the converted data to the user, means for receiving feedback from the user, means for analyzing the feedback, means for again calling the generation AI based on the analysis result to perform additional conversion, and means for providing the analysis result to the user. This enables users who are deficient in certain sensory organs to experience their own work in a different sensory format and receive accurate feedback.
[0587] "Work data" refers to digital files that represent content created by a user, and includes image files, audio files, video files, and the like.
[0588] "Means of acquisition" refers to the technical means by which a user uses a device to send and store work data on a server, and the server receives that data.
[0589] "Generative AI" is an artificial intelligence system that converts a user's work data into a different sensory format, for example, one that has the ability to convert image data into audio data.
[0590] The "calling means" refers to the communication protocol means by which the server sends a conversion request to the generating AI, and the generating AI performs data conversion based on that request.
[0591] "Sensory organ status" refers to information on whether the user's sensory organs, such as vision and hearing, are normal or missing, and serves as the condition for the generation AI to convert data based on this information.
[0592] A "prompt" is an instruction statement that lists specific requests to the generating AI, and is generated based on the work data and the state of the sensory organs.
[0593] "Means for converting" refers to the algorithms and processing system that the generative AI uses to convert the work data into another sensory format based on the prompt text.
[0594] The "means for providing" refers to a communication means and a user interface for sending the converted data to the user's terminal and allowing the user to experience it.
[0595] "Feedback" refers to the opinions and evaluations provided by users after experiencing the converted data, and is important data for the system to use to make improvements.
[0596] The "means of analysis" refers to an analysis system that has the function of collecting feedback from users, analyzing it, and evaluating the accuracy and quality of the conversion by the generation AI.
[0597] "Means for performing additional conversion" refers to technical means for sending a request to the generation AI again based on the analysis results to further optimize and improve the format of the work data.
[0598] This invention is a system that allows users who lack certain sensory organs to experience their own artwork through other sensory organs. This system is primarily composed of a server, a terminal, and a generative AI model. A detailed explanation of how to implement this system is provided below.
[0599] System configuration and hardware / software
[0600] This system uses smartphones and head-mounted displays (HoloLens 2, Oculus Quest 2, etc.) as terminals. AWS EC2 and Heroku are used as servers, and AWS RDS and Firebase are used as databases. OpenAI's GPT-4 is used as the generative AI model.
[0601] Obtaining work data
[0602] The user uses a device to send their creation data to the server. The creation data includes image files, audio files, video files, etc. The user's creation is then stored in digital format on the server.
[0603] Setting and saving the state of the sensory organs
[0604] The user sets the state of their sensory organs via the terminal. This information is sent to the server and stored in the user profile. For example, if the user is visually impaired, that information is stored and used for later processing.
[0605] Conversion process by generative AI
[0606] The server creates a request to send the submitted artwork data and sensory organ status information to the generation AI. This request is generated as a prompt, and the generation AI converts the data based on it. Examples of prompts include the following:
[0607] Please convert the following work (image data) into audio data so that it can be understood by visually impaired people: [Work data]
[0608] Data conversion and provisioning
[0609] Based on the prompt, the AI converts the artwork data into a format that matches the user's sensory state. For example, it converts image data into audio data and returns the result as an audio file to the server. The server then sends the audio data to the device, allowing the user to experience it.
[0610] User feedback and analysis
[0611] Users experience the converted data and enter their evaluations and opinions (feedback) through their devices. The feedback is sent to the server, which analyzes it. The analysis includes whether the feedback is positive or negative, and points for improvement.
[0612] Additional transformations and final feedback
[0613] The server analyzes the feedback and sends another request to the generation AI based on the results. The generation AI performs additional transformations and optimizations and returns the results to the server. This process is repeated until the final feedback is provided to the user.
[0614] Specific examples
[0615] A specific example of this system is shown below.
[0616] 1. Data upload: The user takes a photo of the painting with their smartphone and uploads it to the app.
[0617] 2. Sensory settings: The user sets "visual impairment" and the information is saved on the server.
[0618] 3. Prompt generation: The server generates a prompt for the AI to convert the image data into audio data.
[0619] 4. Data conversion: The generative AI converts image data into audio data.
[0620] 5. Feedback and refinement: The user listens to the audio data and provides feedback. The server analyzes the feedback and, if necessary, converts and refines the data again.
[0621] In this way, users who lack certain sensory organs can experience their creations in an alternative sensory format and receive accurate feedback.
[0622] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0623] Step 1:
[0624] The user uploads the work data using the device.
[0625] Input: Work data (image files, audio files, etc.) stored on the user's device.
[0626] Output: The work data is uploaded to the server and stored in the server database.
[0627] Specific operation: The user uses a smartphone or head-mounted display to select the artwork data through the app and press the upload button. The uploaded data is sent to the server via an HTTP request, and the server stores it in a database.
[0628] Step 2:
[0629] The user sets the state of the sensory organs using the terminal.
[0630] Input: Sensory status (e.g., visual impairment, hearing impairment) set by the user in the device application.
[0631] Output: The sensory organ status information is sent to the server and stored in a database.
[0632] What happens: The user selects the state of their sensory organs within the application and presses the settings button. This information is sent to the server and saved in the user profile.
[0633] Step 3:
[0634] The server generates a prompt for the generated AI.
[0635] Input: Uploaded artwork data and sensory organ status information.
[0636] Output: The prompt to send to the generation AI.
[0637] Specific operation: The server obtains the artwork data and the state of the sensory organs, and generates a prompt based on this. For example, a prompt that converts image data into audio data is created: "Please convert the following artwork (image data) into audio data so that it can be understood by visually impaired people: [artwork data]."
[0638] Step 4:
[0639] The server sends a prompt to the generation AI, requesting data conversion.
[0640] Input: The generated prompt statement.
[0641] Output: Data converted by the generating AI (e.g., audio data).
[0642] Specific operation: The server sends the generated prompt to the generation AI via the OpenAI API. The generation AI performs data conversion based on the prompt.
[0643] Step 5:
[0644] The generation AI returns the converted data to the server.
[0645] Input: Data converted by the generation AI based on the prompt.
[0646] Output: The transformed data is returned to the server.
[0647] Specific operation: The generation AI generates the conversion result and returns it to the server as an HTTP response. The server receives this data and stores it in a database as needed.
[0648] Step 6:
[0649] The server transmits the converted data to the user's terminal.
[0650] Input: Transformation data returned from the generation AI.
[0651] Output: The converted data is displayed and played on the user's device.
[0652] Specific operation: The server obtains the converted data and sends it to the user's device as an HTTP response. The user's device receives the data and displays or plays it in the appropriate format.
[0653] Step 7:
[0654] The user inputs feedback on the converted data.
[0655] Input: The transformation data experienced by the user.
[0656] Output: The feedback data is sent to the server.
[0657] Specific operation: The user uses the device to input feedback about the conversion data they have experienced. Once the input is complete, it is sent to the server and saved.
[0658] Step 8:
[0659] The server analyzes the feedback and requests additional transformations from the generation AI if necessary.
[0660] Input: User feedback data.
[0661] Output: A new prompt statement requesting additional transformations.
[0662] Specific operation: The server analyzes the feedback data and determines whether improvements are necessary. If so, it generates a new prompt for additional conversion and sends it to the generation AI.
[0663] Step 9:
[0664] The generation AI performs additional transformations and returns the results to the server.
[0665] Input: New prompt statement.
[0666] Output: Additional conversion result data.
[0667] Specific operation: The generation AI performs further optimized data conversion based on the new prompt sentence, and sends the conversion result back to the server.
[0668] Step 10:
[0669] The server sends the final feedback data to the user's terminal.
[0670] Input: Analysis results and additional transformation data.
[0671] Output: The final feedback data is provided to the user's terminal.
[0672] What it does: The server compiles the final analysis results and additional conversion data and sends them to the device in the most useful format for the user, allowing the user to re-experience the data and check the final feedback.
[0673] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0674] This invention combines a system that allows users who lack a particular sensory organ to experience their own work through another sensory organ with an emotion engine that recognizes the user's emotions. Specific embodiments of this system will be described below.
[0675] 1. System Overview
[0676] This system consists of a server, a terminal, a generative AI, and an emotion engine.
[0677] Users use their devices to send their creation data to the server and receive feedback. The emotion engine also recognizes the user's emotions and reflects them in the system's processing.
[0678] 2. Receiving a Request
[0679] A user sends a request from a terminal for feedback on their creation (e.g., a painting or a song).
[0680] 3. Acquiring work data
[0681] The server receives the work data sent from the terminal and stores, for example, image files and audio files appropriately.
[0682] 4. Check the condition of the sensory organs
[0683] The server consults a database to ascertain the user's sensory status, for example to determine whether the user is visually impaired or hearing impaired.
[0684] 5. Calling the generation AI
[0685] The server sends a conversion request to the generation AI based on the acquired artwork data. This request includes the artwork data to be converted and the state of the user's sensory organs.
[0686] 6. Emotional Engine Activation
[0687] The server activates an emotion engine to recognize the user's emotions in real time. The emotion engine acquires emotion data from the user's facial expressions and voice.
[0688] 7. Implementing sensory transformation
[0689] The generative AI converts the received artwork data into a format that corresponds to the state of the user's sensory organs. For example, when converting images into audio data, it maps colors and shapes to pitch and rhythm.
[0690] 8. Emotion-Based Regulation
[0691] The server adjusts the generated conversion data based on the emotion data obtained from the emotion engine. For example, if the user is surprised or sad, the server adjusts the conversion data to take those emotions into consideration.
[0692] 9. Sending the converted data
[0693] The generated AI sends the converted data (e.g., audio files, vibration data) back to the server, where it is appropriately encoded.
[0694] 10. Provision of conversion data
[0695] The server sends the converted data to the terminal, encoding the data as an HTTP response.
[0696] The device presents the received converted data to the user, playing audio data for visually impaired users and displaying vibrations or visual waveforms for hearing impaired users.
[0697] 11. Collecting Feedback
[0698] The user inputs feedback on the presented conversion data through the terminal, including their impressions and suggestions for improvement.
[0699] The emotion engine also collects emotional data while users are entering feedback, making it clear what emotions are behind the feedback provided.
[0700] 12. Feedback Analysis
[0701] The server analyzes the received feedback using natural language processing technology to extract important points and suggestions for improvement.
[0702] Emotion data collected from the emotion engine is also used in the analysis to evaluate how the user's emotional state influenced the feedback.
[0703] 13. Recall of generated AI
[0704] Based on the feedback and the results of the analysis of the emotional data, the server sends another request to the generative AI to perform additional transformations and improvements.
[0705] 14. Providing Final Feedback
[0706] The server generates final feedback based on the analysis results and sends it to the terminal in text format.
[0707] The device will then present the final feedback to the user, allowing them to see how their work can be improved and the details of the feedback.
[0708] As a concrete example, consider the case where a visually impaired user requests feedback on a painting (image file) they have created. The user uses their device to send the image file of the painting to a server. The server receives the image file and confirms that the user is visually impaired. The server then sends a request to the generation AI to convert the image file into audio data. Meanwhile, it activates an emotion engine and monitors the user's emotions in real time. The generation AI maps the image's color and shape to sound and returns the result to the server as an audio file. The server adjusts the audio data based on the emotional data obtained from the emotion engine and sends the audio file to the user's device. The device plays the audio file and provides it to the user. The user listens to the audio data and enters their impressions and suggestions for improvement as feedback, and the emotion engine also collects their emotions at this time. The server analyzes the feedback and emotional data and sends another request to the generation AI to make additional conversions or improvements. Finally, the server provides the analysis results in text format to the user, allowing them to check the improvements to the work and details of the feedback.
[0709] This allows visually or hearing impaired users to experience the work through other sensory organs, and by taking into account their emotions in the process, they can receive richer and more accurate feedback.
[0710] The processing flow will be explained below.
[0711] Step 1:
[0712] A user uses a terminal to send a request for feedback on his / her work data (e.g., an image file or an audio file), and the terminal sends the work data along with the request to the server.
[0713] Step 2:
[0714] The server receives the request and work data sent from the device. The server analyzes the HTTP request, obtains the attached file data, and saves it.
[0715] Step 3:
[0716] The server accesses a database to check the user's sensory status, specifically to determine whether the user has a visual or hearing impairment.
[0717] Step 4:
[0718] The server creates a request to send the acquired artwork data to the generation AI. This request includes the artwork data to be converted and the state of the user's sensory organs.
[0719] Step 5:
[0720] The server sends a request to the generation AI, asking it to convert the artwork data into a format that corresponds to the state of the user's sensory organs.
[0721] Step 6:
[0722] The generative AI then converts the received artwork data. For example, for visually impaired users, it converts image data into audio data. In this process, it maps colors and shapes to pitch and rhythm.
[0723] Step 7:
[0724] The generation AI sends the converted data back to the server, where it becomes an audio file, vibration pattern, etc.
[0725] Step 8:
[0726] The server runs an emotion engine to monitor the user's emotions in real time, analyzing facial expressions and voices while the user is experiencing the device.
[0727] Step 9:
[0728] The server receives the converted data returned by the generation AI and makes adjustments based on the emotional data from the emotion engine, adjusting the speed of the voice data or changing the vibration pattern depending on the user's emotional state.
[0729] Step 10:
[0730] The server sends the adjusted converted data to the terminal, which receives the data as an HTTP response and presents it to the user.
[0731] Step 11:
[0732] The terminal provides the converted data to the user, playing it as audio for visually impaired users, or displaying vibrations or visual waveforms for hearing impaired users.
[0733] Step 12:
[0734] The user inputs feedback on the presented data through the terminal, including their impressions of the presented data and suggestions for improvement.
[0735] Step 13:
[0736] The emotion engine collects the emotions of users who enter feedback in real time by analyzing their facial expressions and tone of voice to obtain emotional data.
[0737] Step 14:
[0738] The device sends the user's feedback and emotional data to the server. The feedback is sent in text format, and the emotional data is sent in numerical and categorical format.
[0739] Step 15:
[0740] The server analyzes the received feedback and sentiment data, and uses natural language processing technology to analyze the feedback content and extract key points and improvement suggestions.
[0741] Step 16:
[0742] The server then sends a request to the generative AI again based on the feedback and emotion data to perform additional transformations and refinements.
[0743] Step 17:
[0744] The generation AI performs additional transformations and refinements and sends the results back to the server, which receives the results.
[0745] Step 18:
[0746] The server generates final feedback based on the analysis results and sends it to the terminal in text format.
[0747] Step 19:
[0748] The device will then present the final feedback to the user, allowing them to see how their work can be improved and the details of the feedback.
[0749] Example 2
[0750] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0751] The challenge is to provide a system that allows users who lack certain sensory organs to experience their own work through other sensory organs and receive feedback that takes into account their emotions during the process.
[0752] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0753] In this invention, the server includes means for acquiring work data, means for calling a generation AI that converts the acquired work data into another format based on the state of the user's sensory organs, means for recognizing the user's emotions in real time and acquiring emotional data, means for adjusting the converted data based on the acquired emotional data, means for providing the adjusted converted data to the user, means for receiving feedback from the user and collecting the user's emotional data during that time, means for analyzing the feedback and emotional data, means for calling the generation AI again based on the analysis results and emotional data to perform additional conversion, and means for providing the analysis results to the user. This enables users with visual or hearing impairments to experience the work through their other sensory organs and receive rich and accurate feedback that takes into account the emotions experienced during the process.
[0754] "Work data" refers to content created by users, and refers to digital data such as image files and audio files.
[0755] "Sensory organ status" refers to information indicating the health or absence of impairment of the user's vision, hearing, or other senses.
[0756] "Generative AI" refers to artificial intelligence techniques that transform input data into a different format.
[0757] An "emotion engine" refers to a system that recognizes a user's emotions in real time and collects emotional data from facial expressions and voice.
[0758] "Emotion data" refers to data that indicates the user's emotional state obtained by the emotion engine.
[0759] "Converted data" refers to the new data format converted from the original work data by the generating AI.
[0760] "Feedback" refers to the impressions and evaluations that users provide about the sensory-converted work.
[0761] "Analysis results" refers to the results of the analysis performed by the server based on the feedback and emotional data received from the user.
[0762] "Re-invoking" refers to invoking the generation AI again after the initial invocation to perform additional transformations or improvements.
[0763] "Another format" refers to a data format that can be experienced through a different sensory organ than the original work data (e.g., converting an image into sound).
[0764] This invention combines a system that allows users who lack a particular sensory organ to experience their own work through another sensory organ with an emotion engine that recognizes the user's emotions. Specific embodiments of this system will be described below.
[0765] 1. System Overview
[0766] This system consists of a server, a terminal, a generation AI, and an emotion engine. Users use their terminal to send their creation data to the server and receive feedback. The emotion engine also recognizes the user's emotions and reflects them in each process of the system.
[0767] 2. Hardware and Software Configuration
[0768] Server: A computer system with a powerful processor and large memory capacity that runs the emotion engine and generative AI.
[0769] Device: A device used by a user, such as a smartphone, tablet, or computer, that is connected to the internet.
[0770] Generative AI: Artificial intelligence that uses deep learning models to perform specific transformations, such as converting visual data into audio data.
[0771] Emotion engine: Software that analyzes the user's facial expressions and voice to obtain emotional data in real time.
[0772] 3. Specific Examples
[0773] As a concrete example, consider the case where a visually impaired user requests feedback on a painting (image file) they have created. The user uses their device to send the image file of the painting to a server. The server receives the image file and confirms that the user is visually impaired. The server then sends a request to the generation AI to convert the image file into audio data. Meanwhile, it activates an emotion engine and monitors the user's emotions in real time. The generation AI maps the image's color and shape to sound and returns the result to the server as an audio file. The server adjusts the audio data based on the emotional data obtained from the emotion engine and sends the audio file to the user's device. The device plays the audio file and provides it to the user. The user listens to the audio data and enters their impressions and suggestions for improvement as feedback, and the emotion engine also collects their emotions at this time. The server analyzes the feedback and emotional data and sends another request to the generation AI to make additional conversions or improvements. Finally, the server provides the analysis results in text format to the user, allowing them to check the improvements to the work and details of the feedback.
[0774] 4. Examples of prompts
[0775] "We have an image file of a painting drawn by a user. Based on this image file, please convert it into audio data by mapping the colors and shapes to pitch and rhythm. The user is visually impaired."
[0776] This allows users who lack certain sensory organs to experience their own work through other sensory organs and receive rich feedback that takes into account the emotions based on that experience.
[0777] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0778] Step 1:
[0779] A user uses a terminal to send his / her own artwork data (e.g., an image file of a painting) to the server. The artwork data provided by the user is the input, and the artwork data received by the server is the output.
[0780] Step 2:
[0781] The server receives the artwork data sent from the device, converts it into the appropriate format, and saves it. As part of the data processing, it checks the format of the sent data and saves it in the appropriate storage. The saved artwork data is obtained as output.
[0782] Step 3:
[0783] The server references the database to check the status of the user's sensory organs. The input is identification information such as the user's ID, and the server uses this data to calculate whether or not the sensory organs are impaired. The output is information about the status of the sensory organs.
[0784] Step 4:
[0785] The server sends a conversion request to the generation AI based on the acquired artwork data and information about the user's sensory organs. The inputs include artwork data and information about the state of the sensory organs, and a prompt is generated and sent to the generation AI. For example, a prompt such as "We have an image file of a painting drawn by the user. Based on this image file, please map the color and shape to pitch and rhythm and convert it into audio data" is sent. The output is the data resulting from the conversion by the generation AI (for example, an audio file).
[0786] Step 5:
[0787] The server runs an emotion engine and monitors the user's emotions in real time. The input is the user's facial expressions and voice collected from sensors such as cameras and microphones, and the output is the user's emotional data acquired in real time.
[0788] Step 6:
[0789] The generation AI converts the artwork data it receives into a format that corresponds to the state of the user's sensory organs. For example, data processing involves mapping the color of an image to the pitch of a sound, and the shape to the rhythm of a sound. The converted data (for example, an audio file) is obtained as output.
[0790] Step 7:
[0791] The server adjusts the generated converted data based on the emotional data obtained from the emotion engine. The converted data and emotional data are input, and adjustments are made to take the emotion into consideration through data calculations. For example, if the user is surprised or sad, the tone of the sound may be changed. The adjusted converted data is obtained as output.
[0792] Step 8:
[0793] The server sends the adjusted transformed data to the terminal. The input is the adjusted transformed data, and the output is the data encoded as an HTTP response.
[0794] Step 9:
[0795] The device presents the received converted data to the user. The input is the converted data, and the specific action is to play audio data or transmit a vibration pattern. The output allows the user to experience the work with different sensory organs.
[0796] Step 10:
[0797] The user inputs feedback on the presented converted data through the terminal. At this time, the emotion engine also collects the user's emotions. The input is the user's feedback and the emotional data during that time, and the output is the collected feedback and emotional data.
[0798] Step 11:
[0799] The server analyzes the received feedback and emotional data. The collected feedback and emotional data are input, and natural language processing technology is used to analyze the feedback and extract important points and improvement suggestions. The analysis results are obtained as output.
[0800] Step 12:
[0801] The server then sends a request to the AI generator based on the analysis results to perform additional transformations and improvements. The input is the analysis results, and the output is the regenerated transformation data.
[0802] Step 13:
[0803] The server generates the final feedback and sends it to the terminal in text format. The input is the final analysis result, and the output is the textual feedback provided to the user.
[0804] (Application example 2)
[0805] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0806] In previous systems, users with sensory impairments lacked the means to experience their artwork through other senses, and the feedback they received did not take into account the user's emotions. This resulted in low user satisfaction and the quality of their experience with the artwork. Furthermore, the lack of sufficient feedback collection and analysis made it difficult to further improve the generative AI.
[0807] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to upload artwork data to a terminal at a physical store, means for calling a generation AI that converts the acquired artwork data into another format based on the state of the user's sensory organs, means for providing the converted data to the user, means for activating an emotion engine that acquires emotional data from the user's facial expressions and voice, means for adjusting the converted data according to the emotional data, means for receiving feedback from the user, means for analyzing the feedback and emotional data, means for re-invoking the generation AI based on the analysis results to perform additional conversion, and means for providing the analysis results to the user. This enables users with sensory organ deficiencies to experience their own artwork through their other sensory organs and receive emotional feedback.
[0808] "Means for users to upload artwork data to a device at a physical store" refers to a means for users to photograph or record artworks exhibited at a store they actually visit using a device such as a smartphone or tablet, and then send that data to an online server.
[0809] "Means for invoking a generation AI that converts acquired artwork data into a different format based on the state of the user's sensory organs" refers to a means by which the server launches a generation AI that automatically converts images into an appropriate format, such as converting them into audio, based on information about the state of the user's sensory organs, such as visual or hearing impairments.
[0810] The "means for providing the converted data to the user" refers to a means for transmitting the converted voice data, vibration pattern, or other data to the user's terminal and reading it out loud or notifying it by vibration.
[0811] The "means for activating an emotion engine that acquires emotion data from the user's facial expressions and voice" refers to a means by which the server activates an emotion recognition system that analyzes the user's real-time facial expressions and voice tone to grasp their emotional state.
[0812] The "means for adjusting the converted data in accordance with the emotional data" refers to a means for changing the tone of the voice of the converted data or adjusting the strength of the vibration pattern based on the acquired emotional data.
[0813] The "means for receiving feedback from the user" refers to a means by which the server provides an interface for the user to input their impressions and suggestions for improvement after experiencing the converted data.
[0814] The "means for analyzing feedback and emotional data" refers to a means for analyzing the emotional data acquired simultaneously with the received feedback content, and for deeply understanding and analyzing the user's experiences and opinions.
[0815] "Means of calling the generation AI again based on the analysis results to perform additional conversion" refers to means of starting the generation AI again based on information obtained from the analysis results to generate more accurate conversion data.
[0816] "Means for providing analysis results to users" refers to means for providing the results of the analyzed feedback to users in an easy-to-understand format and suggesting improvements to the work or new experiences.
[0817] This invention combines a system that allows users who lack certain sensory organs to experience their own artwork through other sensory organs with an emotion engine that recognizes the user's emotions. The system consists of a server, a terminal, a generative AI, and an emotion engine.
[0818] First, the user takes a photo or records a piece of artwork on display in a physical store using a device such as a smartphone or tablet. Next, the user uses the device to upload the artwork data to a server. The server stores the received artwork data and retrieves information about the user's sensory organ status (e.g., visual or hearing impairment) from a database.
[0819] The acquired artwork data is converted by the generative AI into an appropriate format based on the state of the user's sensory organs. For example, image files are converted into audio data, and audio files are converted into vibration patterns or visual waveforms. At this time, the server activates an emotion engine to obtain emotional data in real time from the user's facial expressions and voice. The emotional data is reflected in the conversion process by the generative AI, and the converted data is adjusted based on the user's emotional state.
[0820] The converted data is sent from the server to the device and provided to the user. The user experiences the audio and vibration data and sends feedback about the experience to the server via the device. This feedback includes their impressions and areas for improvement, and also includes emotional data collected by the emotion engine.
[0821] The server analyzes the received feedback and emotional data, and sends a re-request to the generation AI to further transform and improve the work data. Finally, the improved data generated based on the analysis results is provided to the user, allowing the user to check the improvements to the work and the details of the feedback.
[0822] As a concrete example, let's say a visually impaired user visits an art museum and takes a photo of a painting. In this case, the app converts the painting into audio and adjusts the audio data while checking the user's emotions using an emotion engine. The user can then provide feedback on the audio they heard, which is then further analyzed by the AI generator and provided to the user as optimized audio data.
[0823] An example of a prompt for a generative AI model is:
[0824] "Text-to-speech prompts for the visually impaired:
[0825] 1. Convert the drawn image into audio data.
[0826] 2. Map colors and shapes to pitch and rhythm.
[0827] 3. Adjust the voice depending on the user's emotions.
[0828] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0829] Step 1:
[0830] A user photographs or records artwork data (e.g., paintings or sounds) in a physical store using a device and uploads it to the server. The input here is the photographed or recorded artwork data file. The device sends this data to the server as an HTTP request. The output is the artwork data file, which is saved on the server.
[0831] Step 2:
[0832] The server stores the received artwork data and retrieves information about the user's sensory organ status (visual impairment, hearing impairment, etc.) from a database. The uploaded artwork data and user ID are used as input. Based on this, the server references the user's profile and retrieves the status of the sensory organs. The output is the status of the user's sensory organs.
[0833] Step 3:
[0834] The server calls the generative AI to convert the acquired artwork data into an appropriate format (sound, vibration pattern, visual waveform, etc.). The inputs are the artwork data and the state of the user's sensory organs. The generative AI model converts the data based on this information. The converted data (sound file, vibration pattern, etc.) is generated as output.
[0835] Step 4:
[0836] The server activates the emotion engine and analyzes the user's facial expressions and voice while experiencing the converted data to obtain emotion data. The input is the user's real-time facial video and voice data. The emotion engine analyzes this data to determine the user's emotional state. The output is emotion data.
[0837] Step 5:
[0838] The server adjusts the conversion data based on the acquired emotional data. For example, the tone of the voice or the intensity of the vibrations is changed according to the user's emotion. The conversion data and emotional data are used as input. The server adjusts the data appropriately based on this. The output is conversion data adjusted according to the emotion.
[0839] Step 6:
[0840] The server sends the adjusted transformed data to the user's device. The adjusted transformed data is used as input. The server sends this as an HTTP response to the device. The transformed data is received as output by the user's device. The user experiences the data through the device.
[0841] Step 7:
[0842] The user experiences the provided converted data and sends feedback about the experience from the device to the server. The user's feedback content is used as input. The device sends this as an HTTP request to the server. The feedback data is saved on the server as output.
[0843] Step 8:
[0844] The server analyzes the received feedback and emotion data and sends a re-request to the generation AI to perform additional transformations and improvements to the work data. The feedback data and emotion data are used as input. The server analyzes this data and sends new instructions to the generation AI. The improved transformation data is generated as output.
[0845] Step 9:
[0846] The server generates the final feedback content based on the analysis results and sends it to the terminal. The parsed feedback data is used as input. The server converts it into text format and sends it to the terminal as an HTTP response. The final feedback is displayed on the user's terminal as output.
[0847] This allows users to experience the artwork and receive emotional feedback regardless of sensory deficiencies.
[0848] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0849] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0850] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0851] [Third embodiment]
[0852] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0853] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0854] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0855] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0856] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0857] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0858] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0859] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0860] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0861] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0862] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0863] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0864] The present invention provides a system that allows users who lack a particular sensory organ to experience their own artwork through another sensory organ. Specific embodiments of this system will be described below.
[0865] 1. System Overview
[0866] This system consists of a server, a terminal, and a generation AI.
[0867] Users use their terminals to send their creation data to the server and receive feedback.
[0868] 2. Receiving a Request
[0869] A user sends a request from a terminal for feedback on their creation (e.g., a painting or a song).
[0870] 3. Acquiring work data
[0871] The server receives the work data sent from the terminal and stores, for example, image files and audio files appropriately.
[0872] 4. Check the condition of the sensory organs
[0873] The server consults a database to ascertain the user's sensory status, for example to determine whether the user is visually impaired or hearing impaired.
[0874] 5. Calling the generation AI
[0875] The server sends the acquired artwork data to the generation AI and sends a request to convert it into a different format depending on the state of the user's sensory organs.
[0876] 6. Implementing sensory transformation
[0877] The generative AI converts the received artwork data into a new format as specified, for example converting image data into audio data for visually impaired users, or audio data into vibration patterns or visual waveforms for hearing impaired users.
[0878] 7. Sending the converted data
[0879] The generation AI returns the converted data to the server, which then sends the data to the user's device.
[0880] 8. Providing conversion data
[0881] The device presents the converted data to the user, either as audio for visually impaired users or as vibrations or visual waveforms for hearing impaired users.
[0882] 9. Collecting Feedback
[0883] The user inputs feedback on the converted data through the terminal, and the terminal transmits the feedback to the server.
[0884] 10. Feedback Analysis
[0885] The server analyzes the received feedback and, if necessary, sends another request to the generating AI to perform additional transformations or improvements.
[0886] 11. Providing Final Feedback
[0887] The server generates feedback based on the analysis results, sends the final analysis results to the terminal, and provides the final feedback to the user.
[0888] As a concrete example, consider the case where a visually impaired user requests feedback on a painting (image file) they have created. The user uses their device to send the image file of the painting to a server. The server receives the image file and confirms that the user is visually impaired. The server then sends a request to the generation AI to convert the image file into audio data. The generation AI maps the image's colors and shapes to sounds and returns the result to the server as an audio file. The server then sends the audio file to the user's device, which plays it and provides it to the user. The user listens to the audio data and enters their impressions and suggestions for improvement as feedback, which the device then sends to the server. The server analyzes the feedback and, if necessary, sends another request to the generation AI to make additional conversions or improvements. Finally, the server provides the analysis results to the user in text format. This process allows visually impaired users to experience their own work as audio and receive feedback.
[0889] This allows visually or hearing impaired users to experience the work through other sensory organs and receive accurate feedback.
[0890] The processing flow will be explained below.
[0891] Step 1:
[0892] A user uses a terminal to send a request for feedback on his / her work data (e.g., an image file or an audio file).
[0893] Step 2:
[0894] The server receives the request and work data sent from the terminal, analyzes the HTTP request, and obtains the contents of the attached file or data.
[0895] Step 3:
[0896] The server retrieves the user's sensory status from a database, including whether the user is visually or hearing impaired.
[0897] Step 4:
[0898] The server sends a conversion request to the generation AI based on the acquired artwork data. This request includes the artwork data to be converted and the state of the user's sensory organs.
[0899] Step 5:
[0900] The generative AI converts the received artwork data into a format that corresponds to the state of the user's sensory organs. For example, when converting images into audio data, it maps colors and shapes to pitch and rhythm.
[0901] Step 6:
[0902] The generated AI sends the converted data (e.g., audio files, vibration data) back to the server, where it is appropriately encoded.
[0903] Step 7:
[0904] The server sends the converted data to the terminal, encoding the data as an HTTP response.
[0905] Step 8:
[0906] The terminal presents the received converted data to the user, playing audio data for visually impaired users and displaying vibrations or visual waveforms for hearing impaired users.
[0907] Step 9:
[0908] The user inputs feedback on the presented conversion data through the terminal, including their impressions and suggestions for improvement.
[0909] Step 10:
[0910] The terminal transmits the feedback from the user to the server. The feedback data is sent to the server as an HTTP request.
[0911] Step 11:
[0912] The server analyzes the received feedback using natural language processing technology to extract important points and suggestions for improvement.
[0913] Step 12:
[0914] The server may send a request to the generation AI again based on the results of analyzing the feedback, to perform additional transformations or improvements.
[0915] Step 13:
[0916] The server generates final feedback based on the analysis results and sends it to the terminal in text format.
[0917] Step 14:
[0918] The device will then present the final feedback to the user, allowing them to see how their work can be improved and the details of the feedback.
[0919] Example 1
[0920] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0921] This solves the problem of users who lack certain sensory organs having difficulty experiencing their own work through other sensory organs and receiving appropriate feedback. Conventional technologies have limited the means by which visually or hearing-impaired users can obtain feedback on their own work, and no efficient system exists for this purpose.
[0922] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0923] In this invention, the server includes means for receiving a request for feedback on work data from a user, means for acquiring work data, means for saving the acquired work data, means for checking the state of the user's sensory organs, means for calling a generation AI that converts the acquired work data into another format based on the state of the user's sensory organs, means for sending a conversion request to the generation AI, means for acquiring the converted data, means for providing the converted data to the user, means for receiving feedback from the user, means for analyzing the feedback, means for again calling the generation AI based on the analysis result to perform additional conversion, and means for providing the analysis result to the user. This enables users with visual or hearing impairments to experience the work through their other sensory organs and receive accurate feedback based on the results.
[0924] "User" refers to a person who uses this system to submit work data and receive feedback.
[0925] "Work data" refers to digital content such as images and audio created by users.
[0926] "Feedback" refers to evaluations, opinions, and comments for improvement provided regarding work data.
[0927] "Status of sensory organs" refers to the functional status of the sensory organs held by the user, for example, whether or not there is a visual or hearing impairment.
[0928] "Generative AI" refers to artificial intelligence used to transform work data into another format.
[0929] "Conversion request" refers to a command that instructs the generation AI to convert work data into a different format.
[0930] "Converted data" refers to data in a new format that has been converted from the original work data by the generating AI.
[0931] "Server" refers to a computer system that receives requests from users, stores work data, and converts and provides the data in cooperation with the generative AI.
[0932] A "request" refers to a request made by a user to a server, for example, a request for feedback.
[0933] "Analysis" refers to the process by which the server analyzes the feedback obtained from the user.
[0934] This invention provides a system that allows users who lack a particular sensory organ to experience their own artwork through another sensory organ. Specific embodiments of this system are described below.
[0935] Hardware and software used
[0936] This system consists of a server, a terminal, and a generative AI. Specific hardware includes high-performance computers (e.g., AWS EC2 instances), and terminals include users' smartphones and tablets. Generative AI uses GPT-4, Stable Diffusion, DALL-E, and other algorithms.
[0937] System Operation Overview
[0938] The following is an overview of how this system works.
[0939] User operation: A user sends a request for feedback on their work through a dedicated application. The request includes the work data (image data and audio data) and the user ID.
[0940] Server processing: The server receives the request sent by the user and saves the work data. The server also references the database based on the user ID to check the state of the user's sensory organs. For example, if the user is visually impaired, the server uses this information to send a request to the generation AI to convert the work data into audio data.
[0941] Processing by the generation AI: The generation AI analyzes the artwork data received from the server and converts it into the specified format. For visually impaired users, it converts image data into audio data, mapping color and shape information as sound. This converted data is sent back to the server as an audio file.
[0942] Server reprocessing: The server sends the converted data to the user's device. The user then experiences the converted data through the device and inputs their feedback. The feedback sent from the device is received and analyzed by the server. If necessary, a request is sent again to the generation AI for additional conversion or improvement.
[0943] Specific examples
[0944] If a visually impaired user wants feedback on a drawing they have made (an image file), the specific steps are as follows:
[0945] 1. The user uses the terminal to send an image file of the painting to the server.
[0946] 2. The server receives the image file and verifies that the user is visually impaired.
[0947] 3. The server sends a request to the AI to convert the image file into audio data. The prompt used is as follows:
[0948] > "Convert the following image data into audio data. The audio data should include information about the colors and shapes contained in the image."
[0949] 4. The generative AI maps the image's color and shape to sound and returns the result to the server as an audio file.
[0950] 5. The server sends the audio file to the user's terminal, which plays the audio file and provides it to the user.
[0951] 6. The user listens to the audio data and enters their impressions and areas for improvement as feedback, which the device then sends to the server.
[0952] 7. The server analyzes the feedback and, if necessary, sends another request to the generating AI for additional transformations or refinements, and provides the final result to the user in text format.
[0953] This allows users with visual or hearing impairments to experience their creations through other senses and receive detailed feedback.
[0954] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0955] Step 1: Receiving the request
[0956] Input: A user sends a request for feedback on a work through a dedicated application. The request includes the work data (image files and audio files) and the user ID.
[0957] How it works: The device accepts the user's request and sends an HTTP request to the server.
[0958] Output: The request data is sent to the server.
[0959] Step 2: Obtaining the work data
[0960] Input: The server receives the HTTP request sent from the device.
[0961] How it works: The server extracts the artwork data and user ID from the request data and saves the image and audio files to cloud storage.
[0962] Output: The work data is saved to cloud storage, and the user ID is saved to a database.
[0963] Step 3: Check the condition of your sensory organs
[0964] Input: The server retrieves the stored user ID.
[0965] How it works: The server checks the database to determine the user's sensory status (visual or hearing impairment), then executes an SQL query to retrieve the user information.
[0966] Output: The state data of the user's sensory organs is obtained.
[0967] Step 4: Calling the generation AI
[0968] Input: Artwork data and data on the state of the user's sensory organs are collected.
[0969] How it works: The server sends a request to the generation AI, including a prompt to convert the artwork data into a new format. An example prompt is "Convert the following image data into audio data. The audio data should include information about the colors and shapes contained in the image."
[0970] Output: A request containing the prompt is sent to the generation AI.
[0971] Step 5: Performing sensory transformations
[0972] Input: The generation AI receives the work data and prompt text sent from the server.
[0973] How it works: Generative AI uses a specified algorithm to convert, for example, image data into audio data, mapping color and shape information to sound.
[0974] Output: The converted data (e.g. audio data) is generated and returned to the server.
[0975] Step 6: Submit the conversion data
[0976] Input: The server receives the transformation data from the generation AI.
[0977] Operation: The server sends the converted data to the user's device using an HTTP response.
[0978] Output: The converted data is sent to the user's terminal.
[0979] Step 7: Provide transformation data
[0980] Input: The device receives the conversion data sent from the server.
[0981] What it does: The device plays or displays the converted data, presenting it as sound for visually impaired users and as vibrations or visual waveforms for hearing impaired users.
[0982] Output: The user experiences the transformed data.
[0983] Step 8: Gather feedback
[0984] Input: The user enters feedback on the transformation data.
[0985] How it works: The device accepts user feedback and sends it to the server using an HTTP request.
[0986] Output: The feedback data is sent to the server.
[0987] Step 9: Analyze feedback
[0988] Input: The server receives the feedback data sent from the device.
[0989] How it works: The server analyzes the feedback data and uses text analysis algorithms to extract user opinions and improvement requests.
[0990] Output: Analysis results are generated.
[0991] Step 10: Provide final feedback
[0992] Input: Feedback analysis results are complete.
[0993] How it works: The server generates feedback based on the analysis results and sends the final analysis results to the user's device, usually in text format.
[0994] Output: Final feedback is displayed on the user's device.
[0995] This allows users with visual or hearing impairments to experience their creations through other sensory organs and receive feedback.
[0996] (Application example 1)
[0997] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0998] It is difficult for users who lack certain sensory organs to experience their own artwork through other sensory organs. For example, when a visually impaired person evaluates their own painting, they cannot experience the artwork visually and must obtain feedback through other means. Similarly, when a hearing impaired person evaluates a musical piece, they need a way to experience it through vision or vibration patterns. However, these methods are not currently provided efficiently or effectively. Furthermore, there is no system that can properly analyze user feedback and convert and optimize artwork data using regenerative AI.
[0999] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1000] In this invention, the server includes means for acquiring work data, means for calling a generation AI that converts the work data into another format based on the state of the user's sensory organs, means for setting the state of the sensory organs, means for saving the sensory organ state information, means for using a generation AI that converts the work data into a corresponding format using a prompt sentence, means for providing the converted data to the user, means for receiving feedback from the user, means for analyzing the feedback, means for again calling the generation AI based on the analysis result to perform additional conversion, and means for providing the analysis result to the user. This enables users who are deficient in certain sensory organs to experience their own work in a different sensory format and receive accurate feedback.
[1001] "Work data" refers to digital files that represent content created by a user, and includes image files, audio files, video files, and the like.
[1002] "Means of acquisition" refers to the technical means by which a user uses a device to send and store work data on a server, and the server receives that data.
[1003] "Generative AI" is an artificial intelligence system that converts a user's work data into a different sensory format, for example, one that has the ability to convert image data into audio data.
[1004] The "calling means" refers to the communication protocol means by which the server sends a conversion request to the generating AI, and the generating AI performs data conversion based on that request.
[1005] "Sensory organ status" refers to information on whether the user's sensory organs, such as vision and hearing, are normal or missing, and serves as the condition for the generation AI to convert data based on this information.
[1006] A "prompt" is an instruction statement that lists specific requests to the generating AI, and is generated based on the work data and the state of the sensory organs.
[1007] "Means for converting" refers to the algorithms and processing system that the generative AI uses to convert the work data into another sensory format based on the prompt text.
[1008] The "means for providing" refers to a communication means and a user interface for sending the converted data to the user's terminal and allowing the user to experience it.
[1009] "Feedback" refers to the opinions and evaluations provided by users after experiencing the converted data, and is important data for the system to use to make improvements.
[1010] The "means of analysis" refers to an analysis system that has the function of collecting feedback from users, analyzing it, and evaluating the accuracy and quality of the conversion by the generation AI.
[1011] "Means for performing additional conversion" refers to technical means for sending a request to the generation AI again based on the analysis results to further optimize and improve the format of the work data.
[1012] This invention is a system that allows users who lack certain sensory organs to experience their own artwork through other sensory organs. This system is primarily composed of a server, a terminal, and a generative AI model. A detailed explanation of how to implement this system is provided below.
[1013] System configuration and hardware / software
[1014] This system uses smartphones and head-mounted displays (HoloLens 2, Oculus Quest 2, etc.) as terminals. AWS EC2 and Heroku are used as servers, and AWS RDS and Firebase are used as databases. OpenAI's GPT-4 is used as the generative AI model.
[1015] Obtaining work data
[1016] The user uses a device to send their creation data to the server. The creation data includes image files, audio files, video files, etc. The user's creation is then stored in digital format on the server.
[1017] Setting and saving the state of the sensory organs
[1018] The user sets the state of their sensory organs via the terminal. This information is sent to the server and stored in the user profile. For example, if the user is visually impaired, that information is stored and used for later processing.
[1019] Conversion process by generative AI
[1020] The server creates a request to send the submitted artwork data and sensory organ status information to the generation AI. This request is generated as a prompt, and the generation AI converts the data based on it. Examples of prompts include the following:
[1021] Please convert the following work (image data) into audio data so that it can be understood by visually impaired people: [Work data]
[1022] Data conversion and provisioning
[1023] Based on the prompt, the AI converts the artwork data into a format that matches the user's sensory state. For example, it converts image data into audio data and returns the result as an audio file to the server. The server then sends the audio data to the device, allowing the user to experience it.
[1024] User feedback and analysis
[1025] Users experience the converted data and enter their evaluations and opinions (feedback) through their devices. The feedback is sent to the server, which analyzes it. The analysis includes whether the feedback is positive or negative, and points for improvement.
[1026] Additional transformations and final feedback
[1027] The server analyzes the feedback and sends another request to the generation AI based on the results. The generation AI performs additional transformations and optimizations and returns the results to the server. This process is repeated until the final feedback is provided to the user.
[1028] Specific examples
[1029] A specific example of this system is shown below.
[1030] 1. Data upload: The user takes a photo of the painting with their smartphone and uploads it to the app.
[1031] 2. Sensory settings: The user sets "visual impairment" and the information is saved on the server.
[1032] 3. Prompt generation: The server generates a prompt for the AI to convert the image data into audio data.
[1033] 4. Data conversion: The generative AI converts image data into audio data.
[1034] 5. Feedback and refinement: The user listens to the audio data and provides feedback. The server analyzes the feedback and, if necessary, converts and refines the data again.
[1035] In this way, users who lack certain sensory organs can experience their creations in an alternative sensory format and receive accurate feedback.
[1036] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1037] Step 1:
[1038] The user uploads the work data using the device.
[1039] Input: Work data (image files, audio files, etc.) stored on the user's device.
[1040] Output: The work data is uploaded to the server and stored in the server database.
[1041] Specific operation: The user uses a smartphone or head-mounted display to select the artwork data through the app and press the upload button. The uploaded data is sent to the server via an HTTP request, and the server stores it in a database.
[1042] Step 2:
[1043] The user sets the state of the sensory organs using the terminal.
[1044] Input: Sensory status (e.g., visual impairment, hearing impairment) set by the user in the device application.
[1045] Output: The sensory organ status information is sent to the server and stored in a database.
[1046] What happens: The user selects the state of their sensory organs within the application and presses the settings button. This information is sent to the server and saved in the user profile.
[1047] Step 3:
[1048] The server generates a prompt for the generated AI.
[1049] Input: Uploaded artwork data and sensory organ status information.
[1050] Output: The prompt to send to the generation AI.
[1051] Specific operation: The server obtains the artwork data and the state of the sensory organs, and generates a prompt based on this. For example, a prompt that converts image data into audio data is created: "Please convert the following artwork (image data) into audio data so that it can be understood by visually impaired people: [artwork data]."
[1052] Step 4:
[1053] The server sends a prompt to the generation AI, requesting data conversion.
[1054] Input: The generated prompt statement.
[1055] Output: Data converted by the generating AI (e.g., audio data).
[1056] Specific operation: The server sends the generated prompt to the generation AI via the OpenAI API. The generation AI performs data conversion based on the prompt.
[1057] Step 5:
[1058] The generation AI returns the converted data to the server.
[1059] Input: Data converted by the generation AI based on the prompt.
[1060] Output: The transformed data is returned to the server.
[1061] Specific operation: The generation AI generates the conversion result and returns it to the server as an HTTP response. The server receives this data and stores it in a database as needed.
[1062] Step 6:
[1063] The server transmits the converted data to the user's terminal.
[1064] Input: Transformation data returned from the generation AI.
[1065] Output: The converted data is displayed and played on the user's device.
[1066] Specific operation: The server obtains the converted data and sends it to the user's device as an HTTP response. The user's device receives the data and displays or plays it in the appropriate format.
[1067] Step 7:
[1068] The user inputs feedback on the converted data.
[1069] Input: The transformation data experienced by the user.
[1070] Output: The feedback data is sent to the server.
[1071] Specific operation: The user uses the device to input feedback about the conversion data they have experienced. Once the input is complete, it is sent to the server and saved.
[1072] Step 8:
[1073] The server analyzes the feedback and requests additional transformations from the generation AI if necessary.
[1074] Input: User feedback data.
[1075] Output: A new prompt statement requesting additional transformations.
[1076] Specific operation: The server analyzes the feedback data and determines whether improvements are necessary. If so, it generates a new prompt for additional conversion and sends it to the generation AI.
[1077] Step 9:
[1078] The generation AI performs additional transformations and returns the results to the server.
[1079] Input: New prompt statement.
[1080] Output: Additional conversion result data.
[1081] Specific operation: The generation AI performs further optimized data conversion based on the new prompt sentence, and sends the conversion result back to the server.
[1082] Step 10:
[1083] The server sends the final feedback data to the user's terminal.
[1084] Input: Analysis results and additional transformation data.
[1085] Output: The final feedback data is provided to the user's terminal.
[1086] What it does: The server compiles the final analysis results and additional conversion data and sends them to the device in the most useful format for the user, allowing the user to re-experience the data and check the final feedback.
[1087] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1088] This invention combines a system that allows users who lack a particular sensory organ to experience their own work through another sensory organ with an emotion engine that recognizes the user's emotions. Specific embodiments of this system will be described below.
[1089] 1. System Overview
[1090] This system consists of a server, a terminal, a generative AI, and an emotion engine.
[1091] Users use their devices to send their creation data to the server and receive feedback. The emotion engine also recognizes the user's emotions and reflects them in the system's processing.
[1092] 2. Receiving a Request
[1093] A user sends a request from a terminal for feedback on their creation (e.g., a painting or a song).
[1094] 3. Acquiring work data
[1095] The server receives the work data sent from the terminal and stores, for example, image files and audio files appropriately.
[1096] 4. Check the condition of the sensory organs
[1097] The server consults a database to ascertain the user's sensory status, for example to determine whether the user is visually impaired or hearing impaired.
[1098] 5. Calling the generation AI
[1099] The server sends a conversion request to the generation AI based on the acquired artwork data. This request includes the artwork data to be converted and the state of the user's sensory organs.
[1100] 6. Emotional Engine Activation
[1101] The server activates an emotion engine to recognize the user's emotions in real time. The emotion engine acquires emotion data from the user's facial expressions and voice.
[1102] 7. Implementing sensory transformation
[1103] The generative AI converts the received artwork data into a format that corresponds to the state of the user's sensory organs. For example, when converting images into audio data, it maps colors and shapes to pitch and rhythm.
[1104] 8. Emotion-Based Regulation
[1105] The server adjusts the generated conversion data based on the emotion data obtained from the emotion engine. For example, if the user is surprised or sad, the server adjusts the conversion data to take those emotions into consideration.
[1106] 9. Sending the converted data
[1107] The generated AI sends the converted data (e.g., audio files, vibration data) back to the server, where it is appropriately encoded.
[1108] 10. Provision of conversion data
[1109] The server sends the converted data to the terminal, encoding the data as an HTTP response.
[1110] The device presents the received converted data to the user, playing audio data for visually impaired users and displaying vibrations or visual waveforms for hearing impaired users.
[1111] 11. Collecting Feedback
[1112] The user inputs feedback on the presented conversion data through the terminal, including their impressions and suggestions for improvement.
[1113] The emotion engine also collects emotional data while users are entering feedback, making it clear what emotions are behind the feedback provided.
[1114] 12. Feedback Analysis
[1115] The server analyzes the received feedback using natural language processing technology to extract important points and suggestions for improvement.
[1116] Emotion data collected from the emotion engine is also used in the analysis to evaluate how the user's emotional state influenced the feedback.
[1117] 13. Recall of generated AI
[1118] Based on the feedback and the results of the analysis of the emotional data, the server sends another request to the generative AI to perform additional transformations and improvements.
[1119] 14. Providing Final Feedback
[1120] The server generates final feedback based on the analysis results and sends it to the terminal in text format.
[1121] The device will then present the final feedback to the user, allowing them to see how their work can be improved and the details of the feedback.
[1122] As a concrete example, consider the case where a visually impaired user requests feedback on a painting (image file) they have created. The user uses their device to send the image file of the painting to a server. The server receives the image file and confirms that the user is visually impaired. The server then sends a request to the generation AI to convert the image file into audio data. Meanwhile, it activates an emotion engine and monitors the user's emotions in real time. The generation AI maps the image's color and shape to sound and returns the result to the server as an audio file. The server adjusts the audio data based on the emotional data obtained from the emotion engine and sends the audio file to the user's device. The device plays the audio file and provides it to the user. The user listens to the audio data and enters their impressions and suggestions for improvement as feedback, and the emotion engine also collects their emotions at this time. The server analyzes the feedback and emotional data and sends another request to the generation AI to make additional conversions or improvements. Finally, the server provides the analysis results in text format to the user, allowing them to check the improvements to the work and details of the feedback.
[1123] This allows visually or hearing impaired users to experience the work through other sensory organs, and by taking into account their emotions in the process, they can receive richer and more accurate feedback.
[1124] The processing flow will be explained below.
[1125] Step 1:
[1126] A user uses a terminal to send a request for feedback on his / her work data (e.g., an image file or an audio file), and the terminal sends the work data along with the request to the server.
[1127] Step 2:
[1128] The server receives the request and work data sent from the device. The server analyzes the HTTP request, obtains the attached file data, and saves it.
[1129] Step 3:
[1130] The server accesses a database to check the user's sensory status, specifically to determine whether the user has a visual or hearing impairment.
[1131] Step 4:
[1132] The server creates a request to send the acquired artwork data to the generation AI. This request includes the artwork data to be converted and the state of the user's sensory organs.
[1133] Step 5:
[1134] The server sends a request to the generation AI, asking it to convert the artwork data into a format that corresponds to the state of the user's sensory organs.
[1135] Step 6:
[1136] The generative AI then converts the received artwork data. For example, for visually impaired users, it converts image data into audio data. In this process, it maps colors and shapes to pitch and rhythm.
[1137] Step 7:
[1138] The generation AI sends the converted data back to the server, where it becomes an audio file, vibration pattern, etc.
[1139] Step 8:
[1140] The server runs an emotion engine to monitor the user's emotions in real time, analyzing facial expressions and voices while the user is experiencing the device.
[1141] Step 9:
[1142] The server receives the converted data returned by the generation AI and makes adjustments based on the emotional data from the emotion engine, adjusting the speed of the voice data or changing the vibration pattern depending on the user's emotional state.
[1143] Step 10:
[1144] The server sends the adjusted converted data to the terminal, which receives the data as an HTTP response and presents it to the user.
[1145] Step 11:
[1146] The terminal provides the converted data to the user, playing it as audio for visually impaired users, or displaying vibrations or visual waveforms for hearing impaired users.
[1147] Step 12:
[1148] The user inputs feedback on the presented data through the terminal, including their impressions of the presented data and suggestions for improvement.
[1149] Step 13:
[1150] The emotion engine collects the emotions of users who enter feedback in real time by analyzing their facial expressions and tone of voice to obtain emotional data.
[1151] Step 14:
[1152] The device sends the user's feedback and emotional data to the server. The feedback is sent in text format, and the emotional data is sent in numerical and categorical format.
[1153] Step 15:
[1154] The server analyzes the received feedback and sentiment data, and uses natural language processing technology to analyze the feedback content and extract key points and improvement suggestions.
[1155] Step 16:
[1156] The server then sends a request to the generative AI again based on the feedback and emotion data to perform additional transformations and refinements.
[1157] Step 17:
[1158] The generation AI performs additional transformations and refinements and sends the results back to the server, which receives the results.
[1159] Step 18:
[1160] The server generates final feedback based on the analysis results and sends it to the terminal in text format.
[1161] Step 19:
[1162] The device will then present the final feedback to the user, allowing them to see how their work can be improved and the details of the feedback.
[1163] Example 2
[1164] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1165] The challenge is to provide a system that allows users who lack certain sensory organs to experience their own work through other sensory organs and receive feedback that takes into account their emotions during the process.
[1166] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1167] In this invention, the server includes means for acquiring work data, means for calling a generation AI that converts the acquired work data into another format based on the state of the user's sensory organs, means for recognizing the user's emotions in real time and acquiring emotional data, means for adjusting the converted data based on the acquired emotional data, means for providing the adjusted converted data to the user, means for receiving feedback from the user and collecting the user's emotional data during that time, means for analyzing the feedback and emotional data, means for calling the generation AI again based on the analysis results and emotional data to perform additional conversion, and means for providing the analysis results to the user. This enables users with visual or hearing impairments to experience the work through their other sensory organs and receive rich and accurate feedback that takes into account the emotions experienced during the process.
[1168] "Work data" refers to content created by users, and refers to digital data such as image files and audio files.
[1169] "Sensory organ status" refers to information indicating the health or absence of impairment of the user's vision, hearing, or other senses.
[1170] "Generative AI" refers to artificial intelligence techniques that transform input data into a different format.
[1171] An "emotion engine" refers to a system that recognizes a user's emotions in real time and collects emotional data from facial expressions and voice.
[1172] "Emotion data" refers to data that indicates the user's emotional state obtained by the emotion engine.
[1173] "Converted data" refers to the new data format converted from the original work data by the generating AI.
[1174] "Feedback" refers to the impressions and evaluations that users provide about the sensory-converted work.
[1175] "Analysis results" refers to the results of the analysis performed by the server based on the feedback and emotional data received from the user.
[1176] "Re-invoking" refers to invoking the generation AI again after the initial invocation to perform additional transformations or improvements.
[1177] "Another format" refers to a data format that can be experienced through a different sensory organ than the original work data (e.g., converting an image into sound).
[1178] This invention combines a system that allows users who lack a particular sensory organ to experience their own work through another sensory organ with an emotion engine that recognizes the user's emotions. Specific embodiments of this system will be described below.
[1179] 1. System Overview
[1180] This system consists of a server, a terminal, a generation AI, and an emotion engine. Users use their terminal to send their creation data to the server and receive feedback. The emotion engine also recognizes the user's emotions and reflects them in each process of the system.
[1181] 2. Hardware and Software Configuration
[1182] Server: A computer system with a powerful processor and large memory capacity that runs the emotion engine and generative AI.
[1183] Device: A device used by a user, such as a smartphone, tablet, or computer, that is connected to the internet.
[1184] Generative AI: Artificial intelligence that uses deep learning models to perform specific transformations, such as converting visual data into audio data.
[1185] Emotion engine: Software that analyzes the user's facial expressions and voice to obtain emotional data in real time.
[1186] 3. Specific Examples
[1187] As a concrete example, consider the case where a visually impaired user requests feedback on a painting (image file) they have created. The user uses their device to send the image file of the painting to a server. The server receives the image file and confirms that the user is visually impaired. The server then sends a request to the generation AI to convert the image file into audio data. Meanwhile, it activates an emotion engine and monitors the user's emotions in real time. The generation AI maps the image's color and shape to sound and returns the result to the server as an audio file. The server adjusts the audio data based on the emotional data obtained from the emotion engine and sends the audio file to the user's device. The device plays the audio file and provides it to the user. The user listens to the audio data and enters their impressions and suggestions for improvement as feedback, and the emotion engine also collects their emotions at this time. The server analyzes the feedback and emotional data and sends another request to the generation AI to make additional conversions or improvements. Finally, the server provides the analysis results in text format to the user, allowing them to check the improvements to the work and details of the feedback.
[1188] 4. Examples of prompts
[1189] "We have an image file of a painting drawn by a user. Based on this image file, please convert it into audio data by mapping the colors and shapes to pitch and rhythm. The user is visually impaired."
[1190] This allows users who lack certain sensory organs to experience their own work through other sensory organs and receive rich feedback that takes into account the emotions based on that experience.
[1191] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1192] Step 1:
[1193] A user uses a terminal to send his / her own artwork data (e.g., an image file of a painting) to the server. The artwork data provided by the user is the input, and the artwork data received by the server is the output.
[1194] Step 2:
[1195] The server receives the artwork data sent from the device, converts it into the appropriate format, and saves it. As part of the data processing, it checks the format of the sent data and saves it in the appropriate storage. The saved artwork data is obtained as output.
[1196] Step 3:
[1197] The server references the database to check the status of the user's sensory organs. The input is identification information such as the user's ID, and the server uses this data to calculate whether or not the sensory organs are impaired. The output is information about the status of the sensory organs.
[1198] Step 4:
[1199] The server sends a conversion request to the generation AI based on the acquired artwork data and information about the user's sensory organs. The inputs include artwork data and information about the state of the sensory organs, and a prompt is generated and sent to the generation AI. For example, a prompt such as "We have an image file of a painting drawn by the user. Based on this image file, please map the color and shape to pitch and rhythm and convert it into audio data" is sent. The output is the data resulting from the conversion by the generation AI (for example, an audio file).
[1200] Step 5:
[1201] The server runs an emotion engine and monitors the user's emotions in real time. The input is the user's facial expressions and voice collected from sensors such as cameras and microphones, and the output is the user's emotional data acquired in real time.
[1202] Step 6:
[1203] The generation AI converts the artwork data it receives into a format that corresponds to the state of the user's sensory organs. For example, data processing involves mapping the color of an image to the pitch of a sound, and the shape to the rhythm of a sound. The converted data (for example, an audio file) is obtained as output.
[1204] Step 7:
[1205] The server adjusts the generated converted data based on the emotional data obtained from the emotion engine. The converted data and emotional data are input, and adjustments are made to take the emotion into consideration through data calculations. For example, if the user is surprised or sad, the tone of the sound may be changed. The adjusted converted data is obtained as output.
[1206] Step 8:
[1207] The server sends the adjusted transformed data to the terminal. The input is the adjusted transformed data, and the output is the data encoded as an HTTP response.
[1208] Step 9:
[1209] The device presents the received converted data to the user. The input is the converted data, and the specific action is to play audio data or transmit a vibration pattern. The output allows the user to experience the work with different sensory organs.
[1210] Step 10:
[1211] The user inputs feedback on the presented converted data through the terminal. At this time, the emotion engine also collects the user's emotions. The input is the user's feedback and the emotional data during that time, and the output is the collected feedback and emotional data.
[1212] Step 11:
[1213] The server analyzes the received feedback and emotional data. The collected feedback and emotional data are input, and natural language processing technology is used to analyze the feedback and extract important points and improvement suggestions. The analysis results are obtained as output.
[1214] Step 12:
[1215] The server then sends a request to the AI generator based on the analysis results to perform additional transformations and improvements. The input is the analysis results, and the output is the regenerated transformation data.
[1216] Step 13:
[1217] The server generates the final feedback and sends it to the terminal in text format. The input is the final analysis result, and the output is the textual feedback provided to the user.
[1218] (Application example 2)
[1219] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1220] In previous systems, users with sensory impairments lacked the means to experience their artwork through other senses, and the feedback they received did not take into account the user's emotions. This resulted in low user satisfaction and the quality of their experience with the artwork. Furthermore, the lack of sufficient feedback collection and analysis made it difficult to further improve the generative AI.
[1221] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to upload artwork data to a terminal at a physical store, means for calling a generation AI that converts the acquired artwork data into another format based on the state of the user's sensory organs, means for providing the converted data to the user, means for activating an emotion engine that acquires emotional data from the user's facial expressions and voice, means for adjusting the converted data according to the emotional data, means for receiving feedback from the user, means for analyzing the feedback and emotional data, means for re-invoking the generation AI based on the analysis results to perform additional conversion, and means for providing the analysis results to the user. This enables users with sensory organ deficiencies to experience their own artwork through their other sensory organs and receive emotional feedback.
[1222] "Means for users to upload artwork data to a device at a physical store" refers to a means for users to photograph or record artworks exhibited at a store they actually visit using a device such as a smartphone or tablet, and then send that data to an online server.
[1223] "Means for invoking a generation AI that converts acquired artwork data into a different format based on the state of the user's sensory organs" refers to a means by which the server launches a generation AI that automatically converts images into an appropriate format, such as converting them into audio, based on information about the state of the user's sensory organs, such as visual or hearing impairments.
[1224] The "means for providing the converted data to the user" refers to a means for transmitting the converted voice data, vibration pattern, or other data to the user's terminal and reading it out loud or notifying it by vibration.
[1225] The "means for activating an emotion engine that acquires emotion data from the user's facial expressions and voice" refers to a means by which the server activates an emotion recognition system that analyzes the user's real-time facial expressions and voice tone to grasp their emotional state.
[1226] The "means for adjusting the converted data in accordance with the emotional data" refers to a means for changing the tone of the voice of the converted data or adjusting the strength of the vibration pattern based on the acquired emotional data.
[1227] The "means for receiving feedback from the user" refers to a means by which the server provides an interface for the user to input their impressions and suggestions for improvement after experiencing the converted data.
[1228] The "means for analyzing feedback and emotional data" refers to a means for analyzing the emotional data acquired simultaneously with the received feedback content, and for deeply understanding and analyzing the user's experiences and opinions.
[1229] "Means of calling the generation AI again based on the analysis results to perform additional conversion" refers to means of starting the generation AI again based on information obtained from the analysis results to generate more accurate conversion data.
[1230] "Means for providing analysis results to users" refers to means for providing the results of the analyzed feedback to users in an easy-to-understand format and suggesting improvements to the work or new experiences.
[1231] This invention combines a system that allows users who lack certain sensory organs to experience their own artwork through other sensory organs with an emotion engine that recognizes the user's emotions. The system consists of a server, a terminal, a generative AI, and an emotion engine.
[1232] First, the user takes a photo or records a piece of artwork on display in a physical store using a device such as a smartphone or tablet. Next, the user uses the device to upload the artwork data to a server. The server stores the received artwork data and retrieves information about the user's sensory organ status (e.g., visual or hearing impairment) from a database.
[1233] The acquired artwork data is converted by the generative AI into an appropriate format based on the state of the user's sensory organs. For example, image files are converted into audio data, and audio files are converted into vibration patterns or visual waveforms. At this time, the server activates an emotion engine to obtain emotional data in real time from the user's facial expressions and voice. The emotional data is reflected in the conversion process by the generative AI, and the converted data is adjusted based on the user's emotional state.
[1234] The converted data is sent from the server to the device and provided to the user. The user experiences the audio and vibration data and sends feedback about the experience to the server via the device. This feedback includes their impressions and areas for improvement, and also includes emotional data collected by the emotion engine.
[1235] The server analyzes the received feedback and emotional data, and sends a re-request to the generation AI to further transform and improve the work data. Finally, the improved data generated based on the analysis results is provided to the user, allowing the user to check the improvements to the work and the details of the feedback.
[1236] As a concrete example, let's say a visually impaired user visits an art museum and takes a photo of a painting. In this case, the app converts the painting into audio and adjusts the audio data while checking the user's emotions using an emotion engine. The user can then provide feedback on the audio they heard, which is then further analyzed by the AI generator and provided to the user as optimized audio data.
[1237] An example of a prompt for a generative AI model is:
[1238] "Text-to-speech prompts for the visually impaired:
[1239] 1. Convert the drawn image into audio data.
[1240] 2. Map colors and shapes to pitch and rhythm.
[1241] 3. Adjust the voice depending on the user's emotions.
[1242] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1243] Step 1:
[1244] A user photographs or records artwork data (e.g., paintings or sounds) in a physical store using a device and uploads it to the server. The input here is the photographed or recorded artwork data file. The device sends this data to the server as an HTTP request. The output is the artwork data file, which is saved on the server.
[1245] Step 2:
[1246] The server stores the received artwork data and retrieves information about the user's sensory organ status (visual impairment, hearing impairment, etc.) from a database. The uploaded artwork data and user ID are used as input. Based on this, the server references the user's profile and retrieves the status of the sensory organs. The output is the status of the user's sensory organs.
[1247] Step 3:
[1248] The server calls the generative AI to convert the acquired artwork data into an appropriate format (sound, vibration pattern, visual waveform, etc.). The inputs are the artwork data and the state of the user's sensory organs. The generative AI model converts the data based on this information. The converted data (sound file, vibration pattern, etc.) is generated as output.
[1249] Step 4:
[1250] The server activates the emotion engine and analyzes the user's facial expressions and voice while experiencing the converted data to obtain emotion data. The input is the user's real-time facial video and voice data. The emotion engine analyzes this data to determine the user's emotional state. The output is emotion data.
[1251] Step 5:
[1252] The server adjusts the conversion data based on the acquired emotional data. For example, the tone of the voice or the intensity of the vibrations is changed according to the user's emotion. The conversion data and emotional data are used as input. The server adjusts the data appropriately based on this. The output is conversion data adjusted according to the emotion.
[1253] Step 6:
[1254] The server sends the adjusted transformed data to the user's device. The adjusted transformed data is used as input. The server sends this as an HTTP response to the device. The transformed data is received as output by the user's device. The user experiences the data through the device.
[1255] Step 7:
[1256] The user experiences the provided converted data and sends feedback about the experience from the device to the server. The user's feedback content is used as input. The device sends this as an HTTP request to the server. The feedback data is saved on the server as output.
[1257] Step 8:
[1258] The server analyzes the received feedback and emotion data and sends a re-request to the generation AI to perform additional transformations and improvements to the work data. The feedback data and emotion data are used as input. The server analyzes this data and sends new instructions to the generation AI. The improved transformation data is generated as output.
[1259] Step 9:
[1260] The server generates the final feedback content based on the analysis results and sends it to the terminal. The parsed feedback data is used as input. The server converts it into text format and sends it to the terminal as an HTTP response. The final feedback is displayed on the user's terminal as output.
[1261] This allows users to experience the artwork and receive emotional feedback regardless of sensory deficiencies.
[1262] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1263] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1264] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1265] [Fourth embodiment]
[1266] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1267] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1268] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1269] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1270] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1271] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1272] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1273] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1274] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1275] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1276] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1277] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1278] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1279] The present invention provides a system that allows users who lack a particular sensory organ to experience their own artwork through another sensory organ. Specific embodiments of this system will be described below.
[1280] 1. System Overview
[1281] This system consists of a server, a terminal, and a generation AI.
[1282] Users use their terminals to send their creation data to the server and receive feedback.
[1283] 2. Receiving a Request
[1284] A user sends a request from a terminal for feedback on their creation (e.g., a painting or a song).
[1285] 3. Acquiring work data
[1286] The server receives the work data sent from the terminal and stores, for example, image files and audio files appropriately.
[1287] 4. Check the condition of the sensory organs
[1288] The server consults a database to ascertain the user's sensory status, for example to determine whether the user is visually impaired or hearing impaired.
[1289] 5. Calling the generation AI
[1290] The server sends the acquired artwork data to the generation AI and sends a request to convert it into a different format depending on the state of the user's sensory organs.
[1291] 6. Implementing sensory transformation
[1292] The generative AI converts the received artwork data into a new format as specified, for example converting image data into audio data for visually impaired users, or audio data into vibration patterns or visual waveforms for hearing impaired users.
[1293] 7. Sending the converted data
[1294] The generation AI returns the converted data to the server, which then sends the data to the user's device.
[1295] 8. Providing conversion data
[1296] The device presents the converted data to the user, either as audio for visually impaired users or as vibrations or visual waveforms for hearing impaired users.
[1297] 9. Collecting Feedback
[1298] The user inputs feedback on the converted data through the terminal, and the terminal transmits the feedback to the server.
[1299] 10. Feedback Analysis
[1300] The server analyzes the received feedback and, if necessary, sends another request to the generating AI to perform additional transformations or improvements.
[1301] 11. Providing Final Feedback
[1302] The server generates feedback based on the analysis results, sends the final analysis results to the terminal, and provides the final feedback to the user.
[1303] As a concrete example, consider the case where a visually impaired user requests feedback on a painting (image file) they have created. The user uses their device to send the image file of the painting to a server. The server receives the image file and confirms that the user is visually impaired. The server then sends a request to the generation AI to convert the image file into audio data. The generation AI maps the image's colors and shapes to sounds and returns the result to the server as an audio file. The server then sends the audio file to the user's device, which plays it and provides it to the user. The user listens to the audio data and enters their impressions and suggestions for improvement as feedback, which the device then sends to the server. The server analyzes the feedback and, if necessary, sends another request to the generation AI to make additional conversions or improvements. Finally, the server provides the analysis results to the user in text format. This process allows visually impaired users to experience their own work as audio and receive feedback.
[1304] This allows visually or hearing impaired users to experience the work through other sensory organs and receive accurate feedback.
[1305] The processing flow will be explained below.
[1306] Step 1:
[1307] A user uses a terminal to send a request for feedback on his / her work data (e.g., an image file or an audio file).
[1308] Step 2:
[1309] The server receives the request and work data sent from the terminal, analyzes the HTTP request, and obtains the contents of the attached file or data.
[1310] Step 3:
[1311] The server retrieves the user's sensory status from a database, including whether the user is visually or hearing impaired.
[1312] Step 4:
[1313] The server sends a conversion request to the generation AI based on the acquired artwork data. This request includes the artwork data to be converted and the state of the user's sensory organs.
[1314] Step 5:
[1315] The generative AI converts the received artwork data into a format that corresponds to the state of the user's sensory organs. For example, when converting images into audio data, it maps colors and shapes to pitch and rhythm.
[1316] Step 6:
[1317] The generated AI sends the converted data (e.g., audio files, vibration data) back to the server, where it is appropriately encoded.
[1318] Step 7:
[1319] The server sends the converted data to the terminal, encoding the data as an HTTP response.
[1320] Step 8:
[1321] The terminal presents the received converted data to the user, playing audio data for visually impaired users and displaying vibrations or visual waveforms for hearing impaired users.
[1322] Step 9:
[1323] The user inputs feedback on the presented conversion data through the terminal, including their impressions and suggestions for improvement.
[1324] Step 10:
[1325] The terminal transmits the feedback from the user to the server. The feedback data is sent to the server as an HTTP request.
[1326] Step 11:
[1327] The server analyzes the received feedback using natural language processing technology to extract important points and suggestions for improvement.
[1328] Step 12:
[1329] The server may send a request to the generation AI again based on the results of analyzing the feedback, to perform additional transformations or improvements.
[1330] Step 13:
[1331] The server generates final feedback based on the analysis results and sends it to the terminal in text format.
[1332] Step 14:
[1333] The device will then present the final feedback to the user, allowing them to see how their work can be improved and the details of the feedback.
[1334] Example 1
[1335] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1336] This solves the problem of users who lack certain sensory organs having difficulty experiencing their own work through other sensory organs and receiving appropriate feedback. Conventional technologies have limited the means by which visually or hearing-impaired users can obtain feedback on their own work, and no efficient system exists for this purpose.
[1337] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1338] In this invention, the server includes means for receiving a request for feedback on work data from a user, means for acquiring work data, means for saving the acquired work data, means for checking the state of the user's sensory organs, means for calling a generation AI that converts the acquired work data into another format based on the state of the user's sensory organs, means for sending a conversion request to the generation AI, means for acquiring the converted data, means for providing the converted data to the user, means for receiving feedback from the user, means for analyzing the feedback, means for again calling the generation AI based on the analysis result to perform additional conversion, and means for providing the analysis result to the user. This enables users with visual or hearing impairments to experience the work through their other sensory organs and receive accurate feedback based on the results.
[1339] "User" refers to a person who uses this system to submit work data and receive feedback.
[1340] "Work data" refers to digital content such as images and audio created by users.
[1341] "Feedback" refers to evaluations, opinions, and comments for improvement provided regarding work data.
[1342] "Status of sensory organs" refers to the functional status of the sensory organs held by the user, for example, whether or not there is a visual or hearing impairment.
[1343] "Generative AI" refers to artificial intelligence used to transform work data into another format.
[1344] "Conversion request" refers to a command that instructs the generation AI to convert work data into a different format.
[1345] "Converted data" refers to data in a new format that has been converted from the original work data by the generating AI.
[1346] "Server" refers to a computer system that receives requests from users, stores work data, and converts and provides the data in cooperation with the generative AI.
[1347] A "request" refers to a request made by a user to a server, for example, a request for feedback.
[1348] "Analysis" refers to the process by which the server analyzes the feedback obtained from the user.
[1349] This invention provides a system that allows users who lack a particular sensory organ to experience their own artwork through another sensory organ. Specific embodiments of this system are described below.
[1350] Hardware and software used
[1351] This system consists of a server, a terminal, and a generative AI. Specific hardware includes high-performance computers (e.g., AWS EC2 instances), and terminals include users' smartphones and tablets. Generative AI uses GPT-4, Stable Diffusion, DALL-E, and other algorithms.
[1352] System Operation Overview
[1353] The following is an overview of how this system works.
[1354] User operation: A user sends a request for feedback on their work through a dedicated application. The request includes the work data (image data and audio data) and the user ID.
[1355] Server processing: The server receives the request sent by the user and saves the work data. The server also references the database based on the user ID to check the state of the user's sensory organs. For example, if the user is visually impaired, the server uses this information to send a request to the generation AI to convert the work data into audio data.
[1356] Processing by the generation AI: The generation AI analyzes the artwork data received from the server and converts it into the specified format. For visually impaired users, it converts image data into audio data, mapping color and shape information as sound. This converted data is sent back to the server as an audio file.
[1357] Server reprocessing: The server sends the converted data to the user's device. The user then experiences the converted data through the device and inputs their feedback. The feedback sent from the device is received and analyzed by the server. If necessary, a request is sent again to the generation AI for additional conversion or improvement.
[1358] Specific examples
[1359] If a visually impaired user wants feedback on a drawing they have made (an image file), the specific steps are as follows:
[1360] 1. The user uses the terminal to send an image file of the painting to the server.
[1361] 2. The server receives the image file and verifies that the user is visually impaired.
[1362] 3. The server sends a request to the AI to convert the image file into audio data. The prompt used is as follows:
[1363] > "Convert the following image data into audio data. The audio data should include information about the colors and shapes contained in the image."
[1364] 4. The generative AI maps the image's color and shape to sound and returns the result to the server as an audio file.
[1365] 5. The server sends the audio file to the user's terminal, which plays the audio file and provides it to the user.
[1366] 6. The user listens to the audio data and enters their impressions and areas for improvement as feedback, which the device then sends to the server.
[1367] 7. The server analyzes the feedback and, if necessary, sends another request to the generating AI for additional transformations or refinements, and provides the final result to the user in text format.
[1368] This allows users with visual or hearing impairments to experience their creations through other senses and receive detailed feedback.
[1369] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1370] Step 1: Receiving the request
[1371] Input: A user sends a request for feedback on a work through a dedicated application. The request includes the work data (image files and audio files) and the user ID.
[1372] How it works: The device accepts the user's request and sends an HTTP request to the server.
[1373] Output: The request data is sent to the server.
[1374] Step 2: Obtaining the work data
[1375] Input: The server receives the HTTP request sent from the device.
[1376] How it works: The server extracts the artwork data and user ID from the request data and saves the image and audio files to cloud storage.
[1377] Output: The work data is saved to cloud storage, and the user ID is saved to a database.
[1378] Step 3: Check the condition of your sensory organs
[1379] Input: The server retrieves the stored user ID.
[1380] How it works: The server checks the database to determine the user's sensory status (visual or hearing impairment), then executes an SQL query to retrieve the user information.
[1381] Output: The state data of the user's sensory organs is obtained.
[1382] Step 4: Calling the generation AI
[1383] Input: Artwork data and data on the state of the user's sensory organs are collected.
[1384] How it works: The server sends a request to the generation AI, including a prompt to convert the artwork data into a new format. An example prompt is "Convert the following image data into audio data. The audio data should include information about the colors and shapes contained in the image."
[1385] Output: A request containing the prompt is sent to the generation AI.
[1386] Step 5: Performing sensory transformations
[1387] Input: The generation AI receives the work data and prompt text sent from the server.
[1388] How it works: Generative AI uses a specified algorithm to convert, for example, image data into audio data, mapping color and shape information to sound.
[1389] Output: The converted data (e.g. audio data) is generated and returned to the server.
[1390] Step 6: Submit the conversion data
[1391] Input: The server receives the transformation data from the generation AI.
[1392] Operation: The server sends the converted data to the user's device using an HTTP response.
[1393] Output: The converted data is sent to the user's terminal.
[1394] Step 7: Provide transformation data
[1395] Input: The device receives the conversion data sent from the server.
[1396] What it does: The device plays or displays the converted data, presenting it as sound for visually impaired users and as vibrations or visual waveforms for hearing impaired users.
[1397] Output: The user experiences the transformed data.
[1398] Step 8: Gather feedback
[1399] Input: The user enters feedback on the transformation data.
[1400] How it works: The device accepts user feedback and sends it to the server using an HTTP request.
[1401] Output: The feedback data is sent to the server.
[1402] Step 9: Analyze feedback
[1403] Input: The server receives the feedback data sent from the device.
[1404] How it works: The server analyzes the feedback data and uses text analysis algorithms to extract user opinions and improvement requests.
[1405] Output: Analysis results are generated.
[1406] Step 10: Provide final feedback
[1407] Input: Feedback analysis results are complete.
[1408] How it works: The server generates feedback based on the analysis results and sends the final analysis results to the user's device, usually in text format.
[1409] Output: Final feedback is displayed on the user's device.
[1410] This allows users with visual or hearing impairments to experience their creations through other sensory organs and receive feedback.
[1411] (Application example 1)
[1412] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1413] It is difficult for users who lack certain sensory organs to experience their own artwork through other sensory organs. For example, when a visually impaired person evaluates their own painting, they cannot experience the artwork visually and must obtain feedback through other means. Similarly, when a hearing impaired person evaluates a musical piece, they need a way to experience it through vision or vibration patterns. However, these methods are not currently provided efficiently or effectively. Furthermore, there is no system that can properly analyze user feedback and convert and optimize artwork data using regenerative AI.
[1414] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1415] In this invention, the server includes means for acquiring work data, means for calling a generation AI that converts the work data into another format based on the state of the user's sensory organs, means for setting the state of the sensory organs, means for saving the sensory organ state information, means for using a generation AI that converts the work data into a corresponding format using a prompt sentence, means for providing the converted data to the user, means for receiving feedback from the user, means for analyzing the feedback, means for again calling the generation AI based on the analysis result to perform additional conversion, and means for providing the analysis result to the user. This enables users who are deficient in certain sensory organs to experience their own work in a different sensory format and receive accurate feedback.
[1416] "Work data" refers to digital files that represent content created by a user, and includes image files, audio files, video files, and the like.
[1417] "Means of acquisition" refers to the technical means by which a user uses a device to send and store work data on a server, and the server receives that data.
[1418] "Generative AI" is an artificial intelligence system that converts a user's work data into a different sensory format, for example, one that has the ability to convert image data into audio data.
[1419] The "calling means" refers to the communication protocol means by which the server sends a conversion request to the generating AI, and the generating AI performs data conversion based on that request.
[1420] "Sensory organ status" refers to information on whether the user's sensory organs, such as vision and hearing, are normal or missing, and serves as the condition for the generation AI to convert data based on this information.
[1421] A "prompt" is an instruction statement that lists specific requests to the generating AI, and is generated based on the work data and the state of the sensory organs.
[1422] "Means for converting" refers to the algorithms and processing system that the generative AI uses to convert the work data into another sensory format based on the prompt text.
[1423] The "means for providing" refers to a communication means and a user interface for sending the converted data to the user's terminal and allowing the user to experience it.
[1424] "Feedback" refers to the opinions and evaluations provided by users after experiencing the converted data, and is important data for the system to use to make improvements.
[1425] The "means of analysis" refers to an analysis system that has the function of collecting feedback from users, analyzing it, and evaluating the accuracy and quality of the conversion by the generation AI.
[1426] "Means for performing additional conversion" refers to technical means for sending a request to the generation AI again based on the analysis results to further optimize and improve the format of the work data.
[1427] This invention is a system that allows users who lack certain sensory organs to experience their own artwork through other sensory organs. This system is primarily composed of a server, a terminal, and a generative AI model. A detailed explanation of how to implement this system is provided below.
[1428] System configuration and hardware / software
[1429] This system uses smartphones and head-mounted displays (HoloLens 2, Oculus Quest 2, etc.) as terminals. AWS EC2 and Heroku are used as servers, and AWS RDS and Firebase are used as databases. OpenAI's GPT-4 is used as the generative AI model.
[1430] Obtaining work data
[1431] The user uses a device to send their creation data to the server. The creation data includes image files, audio files, video files, etc. The user's creation is then stored in digital format on the server.
[1432] Setting and saving the state of the sensory organs
[1433] The user sets the state of their sensory organs via the terminal. This information is sent to the server and stored in the user profile. For example, if the user is visually impaired, that information is stored and used for later processing.
[1434] Conversion process by generative AI
[1435] The server creates a request to send the submitted artwork data and sensory organ status information to the generation AI. This request is generated as a prompt, and the generation AI converts the data based on it. Examples of prompts include the following:
[1436] Please convert the following work (image data) into audio data so that it can be understood by visually impaired people: [Work data]
[1437] Data conversion and provisioning
[1438] Based on the prompt, the AI converts the artwork data into a format that matches the user's sensory state. For example, it converts image data into audio data and returns the result as an audio file to the server. The server then sends the audio data to the device, allowing the user to experience it.
[1439] User feedback and analysis
[1440] Users experience the converted data and enter their evaluations and opinions (feedback) through their devices. The feedback is sent to the server, which analyzes it. The analysis includes whether the feedback is positive or negative, and points for improvement.
[1441] Additional transformations and final feedback
[1442] The server analyzes the feedback and sends another request to the generation AI based on the results. The generation AI performs additional transformations and optimizations and returns the results to the server. This process is repeated until the final feedback is provided to the user.
[1443] Specific examples
[1444] A specific example of this system is shown below.
[1445] 1. Data upload: The user takes a photo of the painting with their smartphone and uploads it to the app.
[1446] 2. Sensory settings: The user sets "visual impairment" and the information is saved on the server.
[1447] 3. Prompt generation: The server generates a prompt for the AI to convert the image data into audio data.
[1448] 4. Data conversion: The generative AI converts image data into audio data.
[1449] 5. Feedback and refinement: The user listens to the audio data and provides feedback. The server analyzes the feedback and, if necessary, converts and refines the data again.
[1450] In this way, users who lack certain sensory organs can experience their creations in an alternative sensory format and receive accurate feedback.
[1451] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1452] Step 1:
[1453] The user uploads the work data using the device.
[1454] Input: Work data (image files, audio files, etc.) stored on the user's device.
[1455] Output: The work data is uploaded to the server and stored in the server database.
[1456] Specific operation: The user uses a smartphone or head-mounted display to select the artwork data through the app and press the upload button. The uploaded data is sent to the server via an HTTP request, and the server stores it in a database.
[1457] Step 2:
[1458] The user sets the state of the sensory organs using the terminal.
[1459] Input: Sensory status (e.g., visual impairment, hearing impairment) set by the user in the device application.
[1460] Output: The sensory organ status information is sent to the server and stored in a database.
[1461] What happens: The user selects the state of their sensory organs within the application and presses the settings button. This information is sent to the server and saved in the user profile.
[1462] Step 3:
[1463] The server generates a prompt for the generated AI.
[1464] Input: Uploaded artwork data and sensory organ status information.
[1465] Output: The prompt to send to the generation AI.
[1466] Specific operation: The server obtains the artwork data and the state of the sensory organs, and generates a prompt based on this. For example, a prompt that converts image data into audio data is created: "Please convert the following artwork (image data) into audio data so that it can be understood by visually impaired people: [artwork data]."
[1467] Step 4:
[1468] The server sends a prompt to the generation AI, requesting data conversion.
[1469] Input: The generated prompt statement.
[1470] Output: Data converted by the generating AI (e.g., audio data).
[1471] Specific operation: The server sends the generated prompt to the generation AI via the OpenAI API. The generation AI performs data conversion based on the prompt.
[1472] Step 5:
[1473] The generation AI returns the converted data to the server.
[1474] Input: Data converted by the generation AI based on the prompt.
[1475] Output: The transformed data is returned to the server.
[1476] Specific operation: The generation AI generates the conversion result and returns it to the server as an HTTP response. The server receives this data and stores it in a database as needed.
[1477] Step 6:
[1478] The server transmits the converted data to the user's terminal.
[1479] Input: Transformation data returned from the generation AI.
[1480] Output: The converted data is displayed and played on the user's device.
[1481] Specific operation: The server obtains the converted data and sends it to the user's device as an HTTP response. The user's device receives the data and displays or plays it in the appropriate format.
[1482] Step 7:
[1483] The user inputs feedback on the converted data.
[1484] Input: The transformation data experienced by the user.
[1485] Output: The feedback data is sent to the server.
[1486] Specific operation: The user uses the device to input feedback about the conversion data they have experienced. Once the input is complete, it is sent to the server and saved.
[1487] Step 8:
[1488] The server analyzes the feedback and requests additional transformations from the generation AI if necessary.
[1489] Input: User feedback data.
[1490] Output: A new prompt statement requesting additional transformations.
[1491] Specific operation: The server analyzes the feedback data and determines whether improvements are necessary. If so, it generates a new prompt for additional conversion and sends it to the generation AI.
[1492] Step 9:
[1493] The generation AI performs additional transformations and returns the results to the server.
[1494] Input: New prompt statement.
[1495] Output: Additional conversion result data.
[1496] Specific operation: The generation AI performs further optimized data conversion based on the new prompt sentence, and sends the conversion result back to the server.
[1497] Step 10:
[1498] The server sends the final feedback data to the user's terminal.
[1499] Input: Analysis results and additional transformation data.
[1500] Output: The final feedback data is provided to the user's terminal.
[1501] What it does: The server compiles the final analysis results and additional conversion data and sends them to the device in the most useful format for the user, allowing the user to re-experience the data and check the final feedback.
[1502] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1503] This invention combines a system that allows users who lack a particular sensory organ to experience their own work through another sensory organ with an emotion engine that recognizes the user's emotions. Specific embodiments of this system will be described below.
[1504] 1. System Overview
[1505] This system consists of a server, a terminal, a generative AI, and an emotion engine.
[1506] Users use their devices to send their creation data to the server and receive feedback. The emotion engine also recognizes the user's emotions and reflects them in the system's processing.
[1507] 2. Receiving a Request
[1508] A user sends a request from a terminal for feedback on their creation (e.g., a painting or a song).
[1509] 3. Acquiring work data
[1510] The server receives the work data sent from the terminal and stores, for example, image files and audio files appropriately.
[1511] 4. Check the condition of the sensory organs
[1512] The server consults a database to ascertain the user's sensory status, for example to determine whether the user is visually impaired or hearing impaired.
[1513] 5. Calling the generation AI
[1514] The server sends a conversion request to the generation AI based on the acquired artwork data. This request includes the artwork data to be converted and the state of the user's sensory organs.
[1515] 6. Emotional Engine Activation
[1516] The server activates an emotion engine to recognize the user's emotions in real time. The emotion engine acquires emotion data from the user's facial expressions and voice.
[1517] 7. Implementing sensory transformation
[1518] The generative AI converts the received artwork data into a format that corresponds to the state of the user's sensory organs. For example, when converting images into audio data, it maps colors and shapes to pitch and rhythm.
[1519] 8. Emotion-Based Regulation
[1520] The server adjusts the generated conversion data based on the emotion data obtained from the emotion engine. For example, if the user is surprised or sad, the server adjusts the conversion data to take those emotions into consideration.
[1521] 9. Sending the converted data
[1522] The generated AI sends the converted data (e.g., audio files, vibration data) back to the server, where it is appropriately encoded.
[1523] 10. Provision of conversion data
[1524] The server sends the converted data to the terminal, encoding the data as an HTTP response.
[1525] The device presents the received converted data to the user, playing audio data for visually impaired users and displaying vibrations or visual waveforms for hearing impaired users.
[1526] 11. Collecting Feedback
[1527] The user inputs feedback on the presented conversion data through the terminal, including their impressions and suggestions for improvement.
[1528] The emotion engine also collects emotional data while users are entering feedback, making it clear what emotions are behind the feedback provided.
[1529] 12. Feedback Analysis
[1530] The server analyzes the received feedback using natural language processing technology to extract important points and suggestions for improvement.
[1531] Emotion data collected from the emotion engine is also used in the analysis to evaluate how the user's emotional state influenced the feedback.
[1532] 13. Recall of generated AI
[1533] Based on the feedback and the results of the analysis of the emotional data, the server sends another request to the generative AI to perform additional transformations and improvements.
[1534] 14. Providing Final Feedback
[1535] The server generates final feedback based on the analysis results and sends it to the terminal in text format.
[1536] The device will then present the final feedback to the user, allowing them to see how their work can be improved and the details of the feedback.
[1537] As a concrete example, consider the case where a visually impaired user requests feedback on a painting (image file) they have created. The user uses their device to send the image file of the painting to a server. The server receives the image file and confirms that the user is visually impaired. The server then sends a request to the generation AI to convert the image file into audio data. Meanwhile, it activates an emotion engine and monitors the user's emotions in real time. The generation AI maps the image's color and shape to sound and returns the result to the server as an audio file. The server adjusts the audio data based on the emotional data obtained from the emotion engine and sends the audio file to the user's device. The device plays the audio file and provides it to the user. The user listens to the audio data and enters their impressions and suggestions for improvement as feedback, and the emotion engine also collects their emotions at this time. The server analyzes the feedback and emotional data and sends another request to the generation AI to make additional conversions or improvements. Finally, the server provides the analysis results in text format to the user, allowing them to check the improvements to the work and details of the feedback.
[1538] This allows visually or hearing impaired users to experience the work through other sensory organs, and by taking into account their emotions in the process, they can receive richer and more accurate feedback.
[1539] The processing flow will be explained below.
[1540] Step 1:
[1541] A user uses a terminal to send a request for feedback on his / her work data (e.g., an image file or an audio file), and the terminal sends the work data along with the request to the server.
[1542] Step 2:
[1543] The server receives the request and work data sent from the device. The server analyzes the HTTP request, obtains the attached file data, and saves it.
[1544] Step 3:
[1545] The server accesses a database to check the user's sensory status, specifically to determine whether the user has a visual or hearing impairment.
[1546] Step 4:
[1547] The server creates a request to send the acquired artwork data to the generation AI. This request includes the artwork data to be converted and the state of the user's sensory organs.
[1548] Step 5:
[1549] The server sends a request to the generation AI, asking it to convert the artwork data into a format that corresponds to the state of the user's sensory organs.
[1550] Step 6:
[1551] The generative AI then converts the received artwork data. For example, for visually impaired users, it converts image data into audio data. In this process, it maps colors and shapes to pitch and rhythm.
[1552] Step 7:
[1553] The generation AI sends the converted data back to the server, where it becomes an audio file, vibration pattern, etc.
[1554] Step 8:
[1555] The server runs an emotion engine to monitor the user's emotions in real time, analyzing facial expressions and voices while the user is experiencing the device.
[1556] Step 9:
[1557] The server receives the converted data returned by the generation AI and makes adjustments based on the emotional data from the emotion engine, adjusting the speed of the voice data or changing the vibration pattern depending on the user's emotional state.
[1558] Step 10:
[1559] The server sends the adjusted converted data to the terminal, which receives the data as an HTTP response and presents it to the user.
[1560] Step 11:
[1561] The terminal provides the converted data to the user, playing it as audio for visually impaired users, or displaying vibrations or visual waveforms for hearing impaired users.
[1562] Step 12:
[1563] The user inputs feedback on the presented data through the terminal, including their impressions of the presented data and suggestions for improvement.
[1564] Step 13:
[1565] The emotion engine collects the emotions of users who enter feedback in real time by analyzing their facial expressions and tone of voice to obtain emotional data.
[1566] Step 14:
[1567] The device sends the user's feedback and emotional data to the server. The feedback is sent in text format, and the emotional data is sent in numerical and categorical format.
[1568] Step 15:
[1569] The server analyzes the received feedback and sentiment data, and uses natural language processing technology to analyze the feedback content and extract key points and improvement suggestions.
[1570] Step 16:
[1571] The server then sends a request to the generative AI again based on the feedback and emotion data to perform additional transformations and refinements.
[1572] Step 17:
[1573] The generation AI performs additional transformations and refinements and sends the results back to the server, which receives the results.
[1574] Step 18:
[1575] The server generates final feedback based on the analysis results and sends it to the terminal in text format.
[1576] Step 19:
[1577] The device will then present the final feedback to the user, allowing them to see how their work can be improved and the details of the feedback.
[1578] Example 2
[1579] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1580] The challenge is to provide a system that allows users who lack certain sensory organs to experience their own work through other sensory organs and receive feedback that takes into account their emotions during the process.
[1581] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1582] In this invention, the server includes means for acquiring work data, means for calling a generation AI that converts the acquired work data into another format based on the state of the user's sensory organs, means for recognizing the user's emotions in real time and acquiring emotional data, means for adjusting the converted data based on the acquired emotional data, means for providing the adjusted converted data to the user, means for receiving feedback from the user and collecting the user's emotional data during that time, means for analyzing the feedback and emotional data, means for calling the generation AI again based on the analysis results and emotional data to perform additional conversion, and means for providing the analysis results to the user. This enables users with visual or hearing impairments to experience the work through their other sensory organs and receive rich and accurate feedback that takes into account the emotions experienced during the process.
[1583] "Work data" refers to content created by users, and refers to digital data such as image files and audio files.
[1584] "Sensory organ status" refers to information indicating the health or absence of impairment of the user's vision, hearing, or other senses.
[1585] "Generative AI" refers to artificial intelligence techniques that transform input data into a different format.
[1586] An "emotion engine" refers to a system that recognizes a user's emotions in real time and collects emotional data from facial expressions and voice.
[1587] "Emotion data" refers to data that indicates the user's emotional state obtained by the emotion engine.
[1588] "Converted data" refers to the new data format converted from the original work data by the generating AI.
[1589] "Feedback" refers to the impressions and evaluations that users provide about the sensory-converted work.
[1590] "Analysis results" refers to the results of the analysis performed by the server based on the feedback and emotional data received from the user.
[1591] "Re-invoking" refers to invoking the generation AI again after the initial invocation to perform additional transformations or improvements.
[1592] "Another format" refers to a data format that can be experienced through a different sensory organ than the original work data (e.g., converting an image into sound).
[1593] This invention combines a system that allows users who lack a particular sensory organ to experience their own work through another sensory organ with an emotion engine that recognizes the user's emotions. Specific embodiments of this system will be described below.
[1594] 1. System Overview
[1595] This system consists of a server, a terminal, a generation AI, and an emotion engine. Users use their terminal to send their creation data to the server and receive feedback. The emotion engine also recognizes the user's emotions and reflects them in each process of the system.
[1596] 2. Hardware and Software Configuration
[1597] Server: A computer system with a powerful processor and large memory capacity that runs the emotion engine and generative AI.
[1598] Device: A device used by a user, such as a smartphone, tablet, or computer, that is connected to the internet.
[1599] Generative AI: Artificial intelligence that uses deep learning models to perform specific transformations, such as converting visual data into audio data.
[1600] Emotion engine: Software that analyzes the user's facial expressions and voice to obtain emotional data in real time.
[1601] 3. Specific Examples
[1602] As a concrete example, consider the case where a visually impaired user requests feedback on a painting (image file) they have created. The user uses their device to send the image file of the painting to a server. The server receives the image file and confirms that the user is visually impaired. The server then sends a request to the generation AI to convert the image file into audio data. Meanwhile, it activates an emotion engine and monitors the user's emotions in real time. The generation AI maps the image's color and shape to sound and returns the result to the server as an audio file. The server adjusts the audio data based on the emotional data obtained from the emotion engine and sends the audio file to the user's device. The device plays the audio file and provides it to the user. The user listens to the audio data and enters their impressions and suggestions for improvement as feedback, and the emotion engine also collects their emotions at this time. The server analyzes the feedback and emotional data and sends another request to the generation AI to make additional conversions or improvements. Finally, the server provides the analysis results in text format to the user, allowing them to check the improvements to the work and details of the feedback.
[1603] 4. Examples of prompts
[1604] "We have an image file of a painting drawn by a user. Based on this image file, please convert it into audio data by mapping the colors and shapes to pitch and rhythm. The user is visually impaired."
[1605] This allows users who lack certain sensory organs to experience their own work through other sensory organs and receive rich feedback that takes into account the emotions based on that experience.
[1606] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1607] Step 1:
[1608] A user uses a terminal to send his / her own artwork data (e.g., an image file of a painting) to the server. The artwork data provided by the user is the input, and the artwork data received by the server is the output.
[1609] Step 2:
[1610] The server receives the artwork data sent from the device, converts it into the appropriate format, and saves it. As part of the data processing, it checks the format of the sent data and saves it in the appropriate storage. The saved artwork data is obtained as output.
[1611] Step 3:
[1612] The server references the database to check the status of the user's sensory organs. The input is identification information such as the user's ID, and the server uses this data to calculate whether or not the sensory organs are impaired. The output is information about the status of the sensory organs.
[1613] Step 4:
[1614] The server sends a conversion request to the generation AI based on the acquired artwork data and information about the user's sensory organs. The inputs include artwork data and information about the state of the sensory organs, and a prompt is generated and sent to the generation AI. For example, a prompt such as "We have an image file of a painting drawn by the user. Based on this image file, please map the color and shape to pitch and rhythm and convert it into audio data" is sent. The output is the data resulting from the conversion by the generation AI (for example, an audio file).
[1615] Step 5:
[1616] The server runs an emotion engine and monitors the user's emotions in real time. The input is the user's facial expressions and voice collected from sensors such as cameras and microphones, and the output is the user's emotional data acquired in real time.
[1617] Step 6:
[1618] The generation AI converts the artwork data it receives into a format that corresponds to the state of the user's sensory organs. For example, data processing involves mapping the color of an image to the pitch of a sound, and the shape to the rhythm of a sound. The converted data (for example, an audio file) is obtained as output.
[1619] Step 7:
[1620] The server adjusts the generated converted data based on the emotional data obtained from the emotion engine. The converted data and emotional data are input, and adjustments are made to take the emotion into consideration through data calculations. For example, if the user is surprised or sad, the tone of the sound may be changed. The adjusted converted data is obtained as output.
[1621] Step 8:
[1622] The server sends the adjusted transformed data to the terminal. The input is the adjusted transformed data, and the output is the data encoded as an HTTP response.
[1623] Step 9:
[1624] The device presents the received converted data to the user. The input is the converted data, and the specific action is to play audio data or transmit a vibration pattern. The output allows the user to experience the work with different sensory organs.
[1625] Step 10:
[1626] The user inputs feedback on the presented converted data through the terminal. At this time, the emotion engine also collects the user's emotions. The input is the user's feedback and the emotional data during that time, and the output is the collected feedback and emotional data.
[1627] Step 11:
[1628] The server analyzes the received feedback and emotional data. The collected feedback and emotional data are input, and natural language processing technology is used to analyze the feedback and extract important points and improvement suggestions. The analysis results are obtained as output.
[1629] Step 12:
[1630] The server then sends a request to the AI generator based on the analysis results to perform additional transformations and improvements. The input is the analysis results, and the output is the regenerated transformation data.
[1631] Step 13:
[1632] The server generates the final feedback and sends it to the terminal in text format. The input is the final analysis result, and the output is the textual feedback provided to the user.
[1633] (Application example 2)
[1634] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1635] In previous systems, users with sensory impairments lacked the means to experience their artwork through other senses, and the feedback they received did not take into account the user's emotions. This resulted in low user satisfaction and the quality of their experience with the artwork. Furthermore, the lack of sufficient feedback collection and analysis made it difficult to further improve the generative AI.
[1636] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to upload artwork data to a terminal at a physical store, means for calling a generation AI that converts the acquired artwork data into another format based on the state of the user's sensory organs, means for providing the converted data to the user, means for activating an emotion engine that acquires emotional data from the user's facial expressions and voice, means for adjusting the converted data according to the emotional data, means for receiving feedback from the user, means for analyzing the feedback and emotional data, means for re-invoking the generation AI based on the analysis results to perform additional conversion, and means for providing the analysis results to the user. This enables users with sensory organ deficiencies to experience their own artwork through their other sensory organs and receive emotional feedback.
[1637] "Means for users to upload artwork data to a device at a physical store" refers to a means for users to photograph or record artworks exhibited at a store they actually visit using a device such as a smartphone or tablet, and then send that data to an online server.
[1638] "Means for invoking a generation AI that converts acquired artwork data into a different format based on the state of the user's sensory organs" refers to a means by which the server launches a generation AI that automatically converts images into an appropriate format, such as converting them into audio, based on information about the state of the user's sensory organs, such as visual or hearing impairments.
[1639] The "means for providing the converted data to the user" refers to a means for transmitting the converted voice data, vibration pattern, or other data to the user's terminal and reading it out loud or notifying it by vibration.
[1640] The "means for activating an emotion engine that acquires emotion data from the user's facial expressions and voice" refers to a means by which the server activates an emotion recognition system that analyzes the user's real-time facial expressions and voice tone to grasp their emotional state.
[1641] The "means for adjusting the converted data in accordance with the emotional data" refers to a means for changing the tone of the voice of the converted data or adjusting the strength of the vibration pattern based on the acquired emotional data.
[1642] The "means for receiving feedback from the user" refers to a means by which the server provides an interface for the user to input their impressions and suggestions for improvement after experiencing the converted data.
[1643] The "means for analyzing feedback and emotional data" refers to a means for analyzing the emotional data acquired simultaneously with the received feedback content, and for deeply understanding and analyzing the user's experiences and opinions.
[1644] "Means of calling the generation AI again based on the analysis results to perform additional conversion" refers to means of starting the generation AI again based on information obtained from the analysis results to generate more accurate conversion data.
[1645] "Means for providing analysis results to users" refers to means for providing the results of the analyzed feedback to users in an easy-to-understand format and suggesting improvements to the work or new experiences.
[1646] This invention combines a system that allows users who lack certain sensory organs to experience their own artwork through other sensory organs with an emotion engine that recognizes the user's emotions. The system consists of a server, a terminal, a generative AI, and an emotion engine.
[1647] First, the user takes a photo or records a piece of artwork on display in a physical store using a device such as a smartphone or tablet. Next, the user uses the device to upload the artwork data to a server. The server stores the received artwork data and retrieves information about the user's sensory organ status (e.g., visual or hearing impairment) from a database.
[1648] The acquired artwork data is converted by the generative AI into an appropriate format based on the state of the user's sensory organs. For example, image files are converted into audio data, and audio files are converted into vibration patterns or visual waveforms. At this time, the server activates an emotion engine to obtain emotional data in real time from the user's facial expressions and voice. The emotional data is reflected in the conversion process by the generative AI, and the converted data is adjusted based on the user's emotional state.
[1649] The converted data is sent from the server to the device and provided to the user. The user experiences the audio and vibration data and sends feedback about the experience to the server via the device. This feedback includes their impressions and areas for improvement, and also includes emotional data collected by the emotion engine.
[1650] The server analyzes the received feedback and emotional data, and sends a re-request to the generation AI to further transform and improve the work data. Finally, the improved data generated based on the analysis results is provided to the user, allowing the user to check the improvements to the work and the details of the feedback.
[1651] As a concrete example, let's say a visually impaired user visits an art museum and takes a photo of a painting. In this case, the app converts the painting into audio and adjusts the audio data while checking the user's emotions using an emotion engine. The user can then provide feedback on the audio they heard, which is then further analyzed by the AI generator and provided to the user as optimized audio data.
[1652] An example of a prompt for a generative AI model is:
[1653] "Text-to-speech prompts for the visually impaired:
[1654] 1. Convert the drawn image into audio data.
[1655] 2. Map colors and shapes to pitch and rhythm.
[1656] 3. Adjust the voice depending on the user's emotions.
[1657] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1658] Step 1:
[1659] A user photographs or records artwork data (e.g., paintings or sounds) in a physical store using a device and uploads it to the server. The input here is the photographed or recorded artwork data file. The device sends this data to the server as an HTTP request. The output is the artwork data file, which is saved on the server.
[1660] Step 2:
[1661] The server stores the received artwork data and retrieves information about the user's sensory organ status (visual impairment, hearing impairment, etc.) from a database. The uploaded artwork data and user ID are used as input. Based on this, the server references the user's profile and retrieves the status of the sensory organs. The output is the status of the user's sensory organs.
[1662] Step 3:
[1663] The server calls the generative AI to convert the acquired artwork data into an appropriate format (sound, vibration pattern, visual waveform, etc.). The inputs are the artwork data and the state of the user's sensory organs. The generative AI model converts the data based on this information. The converted data (sound file, vibration pattern, etc.) is generated as output.
[1664] Step 4:
[1665] The server activates the emotion engine and analyzes the user's facial expressions and voice while experiencing the converted data to obtain emotion data. The input is the user's real-time facial video and voice data. The emotion engine analyzes this data to determine the user's emotional state. The output is emotion data.
[1666] Step 5:
[1667] The server adjusts the conversion data based on the acquired emotional data. For example, the tone of the voice or the intensity of the vibrations is changed according to the user's emotion. The conversion data and emotional data are used as input. The server adjusts the data appropriately based on this. The output is conversion data adjusted according to the emotion.
[1668] Step 6:
[1669] The server sends the adjusted transformed data to the user's device. The adjusted transformed data is used as input. The server sends this as an HTTP response to the device. The transformed data is received as output by the user's device. The user experiences the data through the device.
[1670] Step 7:
[1671] The user experiences the provided converted data and sends feedback about the experience from the device to the server. The user's feedback content is used as input. The device sends this as an HTTP request to the server. The feedback data is saved on the server as output.
[1672] Step 8:
[1673] The server analyzes the received feedback and emotion data and sends a re-request to the generation AI to perform additional transformations and improvements to the work data. The feedback data and emotion data are used as input. The server analyzes this data and sends new instructions to the generation AI. The improved transformation data is generated as output.
[1674] Step 9:
[1675] The server generates the final feedback content based on the analysis results and sends it to the terminal. The parsed feedback data is used as input. The server converts it into text format and sends it to the terminal as an HTTP response. The final feedback is displayed on the user's terminal as output.
[1676] This allows users to experience the artwork and receive emotional feedback regardless of sensory deficiencies.
[1677] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1678] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1679] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1680] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1681] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1682] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1683] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1684] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1685] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1686] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1687] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1688] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1689] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1690] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1691] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1692] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1693] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1694] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1695] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1696] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1697] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1698] The following is further disclosed regarding the above embodiment.
[1699] (Claim 1)
[1700] A means for acquiring work data;
[1701] A means to call a generation AI that converts the acquired artwork data into a different format based on the state of the user's sensory organs;
[1702] means for providing the converted data to a user;
[1703] a means for receiving feedback from users;
[1704] a means for analyzing the feedback;
[1705] A method to call the generation AI again based on the analysis results and perform additional conversions,
[1706] A means of providing analysis results to users
[1707] A system including:
[1708] (Claim 2)
[1709] 2. The system according to claim 1, wherein the work data is an image file, and further comprising means for converting the image file into audio data.
[1710] (Claim 3)
[1711] 10. The system of claim 1, wherein the production data is an audio file and the system includes means for converting the audio file into a vibration pattern or a visual waveform.
[1712] "Example 1"
[1713] (Claim 1)
[1714] means for receiving a request from a user requesting feedback on the work data;
[1715] A means for acquiring work data;
[1716] A means for saving the acquired work data;
[1717] means for ascertaining the state of the user's sensory organs;
[1718] A means to call a generation AI that converts the acquired artwork data into a different format based on the state of the user's sensory organs;
[1719] A means for sending a conversion request to the generation AI;
[1720] a means for obtaining the transformed data;
[1721] means for providing the converted data to a user;
[1722] a means for receiving feedback from users;
[1723] a means for analyzing the feedback;
[1724] A method to call the generation AI again based on the analysis results and perform additional conversions,
[1725] A means of providing analysis results to users
[1726] A system including:
[1727] (Claim 2)
[1728] 2. The system according to claim 1, wherein the artwork data is an image file, and further comprising means for checking the state of the user's sensory organs and converting the image file into audio data.
[1729] (Claim 3)
[1730] 2. The system of claim 1, wherein the work data is an audio file, and further comprising means for checking the state of the user's sensory organs and converting the audio file into a vibration pattern or a visual waveform.
[1731] "Application Example 1"
[1732] (Claim 1)
[1733] A means for acquiring work data;
[1734] A means to call a generation AI that converts the acquired artwork data into a different format based on the state of the user's sensory organs;
[1735] means for providing the converted data to a user;
[1736] a means for receiving feedback from users;
[1737] a means for analyzing the feedback;
[1738] A method to call the generation AI again based on the analysis results and perform additional conversions,
[1739] a means for providing the analysis results to a user;
[1740] a means for setting the state of the sensory organs;
[1741] means for storing sensory organ state information;
[1742] a means for generating a prompt sentence for converting the artwork data into a format corresponding to the state of the sensory organs;
[1743] a means for using a generative AI to convert the work data into a corresponding format using the prompt sentence;
[1744] A system including:
[1745] (Claim 2)
[1746] 2. The system according to claim 1, wherein the work data is an image file, and further comprising means for converting the image file into audio data.
[1747] (Claim 3)
[1748] 10. The system of claim 1, wherein the production data is an audio file and the system includes means for converting the audio file into a vibration pattern or a visual waveform.
[1749] "Example 2: Combining Emotion Engines"
[1750] (Claim 1)
[1751] A means for acquiring work data;
[1752] A means to call a generation AI that converts the acquired artwork data into a different format based on the state of the user's sensory organs;
[1753] a means for recognizing a user's emotions in real time and acquiring emotion data;
[1754] a means for adjusting the converted data based on the acquired emotion data;
[1755] means for providing the adjusted transformation data to a user;
[1756] a means for receiving feedback from the user and collecting emotional data of the user during the process;
[1757] a means for analyzing the feedback and sentiment data;
[1758] A method to call the generation AI again based on the analysis results and emotion data to perform additional conversions,
[1759] A means of providing analysis results to users
[1760] A system including:
[1761] (Claim 2)
[1762] 2. The system according to claim 1, wherein the work data is an image file, and further comprising means for converting the image file into audio data.
[1763] (Claim 3)
[1764] 10. The system of claim 1, wherein the production data is an audio file and the system includes means for converting the audio file into a vibration pattern or a visual waveform.
[1765] "Application example 2 when combining emotion engines"
[1766] (Claim 1)
[1767] A means for users to upload their work data to a device at a physical store,
[1768] A means to call a generation AI that converts the acquired artwork data into a different format based on the state of the user's sensory organs;
[1769] means for providing the converted data to a user;
[1770] means for activating an emotion engine that acquires emotion data from a user's facial expression and voice;
[1771] means for adjusting the conversion data in response to the emotion data;
[1772] a means for receiving feedback from users;
[1773] a means for analyzing the feedback and sentiment data;
[1774] A method to call the generation AI again based on the analysis results and perform additional conversions,
[1775] A means of providing analysis results to users
[1776] A system including:
[1777] (Claim 2)
[1778] 2. The system according to claim 1, wherein the work data is an image file, and further comprising means for converting the image file into audio data.
[1779] (Claim 3)
[1780] 10. The system of claim 1, wherein the production data is an audio file and the system includes means for converting the audio file into a vibration pattern or a visual waveform. [Explanation of symbols]
[1781] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for acquiring work data; A means to call a generation AI that converts the acquired artwork data into a different format based on the state of the user's sensory organs; means for providing the converted data to a user; a means for receiving feedback from users; a means for analyzing the feedback; A method to call the generation AI again based on the analysis results and perform additional conversions, A means of providing analysis results to users A system including:
2. 2. The system according to claim 1, wherein the work data is an image file, and further comprising means for converting the image file into audio data.
3. 2. The system of claim 1, wherein the production data is an audio file and further comprising means for converting the audio file into a vibration pattern or a visual waveform.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A