system
A system that processes user voice input and behavior history to automate application suggestions and image editing addresses inefficiencies in smartphone tasks, enhancing user convenience and quality of life.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Modern smartphone users face cumbersome daily tasks due to inefficient management of applications and image editing, and there is a lack of quick and accurate information provision tailored to individual needs.
A system that converts user voice input into text data, analyzes user behavior history for application suggestions, and automatically edits images, enhancing convenience through natural language processing, automated application suggestions, and image editing.
Enables efficient completion of daily tasks by providing high-quality image editing and personalized application suggestions based on user behavior, improving user convenience and quality of life.
Smart Images

Figure 2026073347000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003] <000**********
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Modern smartphone users have problems that the operation of daily tasks is cumbersome, especially spending a lot of time on the management of many applications and image editing. Furthermore, it is difficult to provide quick and accurate information and functions according to individual needs. It is necessary to solve such problems more efficiently and effectively.
Means for Solving the Problems
[0005] This invention provides a natural language processing means that converts user voice input into text data and analyzes its content, thereby realizing a system that enables voice operation. It also includes a suggestion generation means that records and analyzes the user's behavior history to propose the most suitable application for the user. Furthermore, by including an image editing means capable of automatically editing image data, the system provides high-quality image editing results without direct user intervention.
[0006] "Audio data" refers to audio information spoken by a user, recorded as a digital signal.
[0007] "Text data" refers to character information obtained by processing audio data, and is the subject of natural language processing.
[0008] "Natural language processing means" refers to a device or software that includes technology for converting speech data into text data, analyzing its content, and interpreting its meaning.
[0009] "Analysis means" refers to a device or software that has the function of understanding user requests and commands from text data and generating an appropriate response.
[0010] "Communication means" refers to the technology or device used to send and receive data between a terminal and a server.
[0011] A "proposal generation means" is a device or software that has the function of selecting and presenting optimal or useful applications or information to the user based on the user's behavior history.
[0012] "Image data" refers to visual information represented in digital format as a still image.
[0013] "Image editing means" refers to a device or program that has techniques for changing or improving the visual characteristics of image data. [Brief explanation of the drawing]
[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [[ID=:32]] [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention provides a system incorporating various functions to improve user convenience. The system primarily performs natural language processing for handling voice data, application suggestions based on user behavior history, and automatic image data editing. The detailed operation of each function is described below.
[0036] Voice control using natural language processing
[0037] When a user speaks into their smartphone using natural language, the device converts the audio into digital data and sends it to a server. The server uses advanced natural language processing technology to convert the audio into text data and analyze its content. For example, if a user asks, "What time is the next meeting?", the server refers to calendar information and derives the answer. Then, it converts the analysis result back into audio and outputs it from the device in a format that is easy for the user to understand.
[0038] Automated application suggestions
[0039] The device records the user's application usage history and periodically sends it to the server. Based on this history, the server analyzes the user's behavior patterns and predicts the next most suitable application. For example, if a user has a habit of opening the weather app at 7 AM every day, the server instructs the device to display the weather app on the home screen in the morning. In this way, the system helps save the user time.
[0040] Automated image editing
[0041] Images taken by the user are sent to the server via the device. The server utilizes image editing algorithms to adjust the brightness, contrast, and color of the images, automating high-quality editing. For example, a photo of a cloudy beach taken during a trip can be transformed into a bright and vibrant image. The edited image is then sent back from the server to the device and immediately saved to the gallery.
[0042] This system aims to improve the quality of life by enabling users to complete daily tasks more efficiently.
[0043] The following describes the processing flow.
[0044] Step 1:
[0045] The user speaks into their smartphone, saying, "Tell me my next appointment." The device acquires the voice data through the microphone and temporarily stores it locally.
[0046] Step 2:
[0047] The terminal sends the stored audio data to the server using a communication module. The server then passes the received audio data to a natural language processing engine.
[0048] Step 3:
[0049] The server uses a natural language processing engine to convert the audio data into text data. This text data then serves as input data for analyzing the user's intent.
[0050] Step 4:
[0051] The server parses the converted text data and retrieves the user's next appointment by referencing the calendar database. After generating an appropriate response, it formats the response back into text format.
[0052] Step 5:
[0053] The server sends the generated text response to a text-to-speech module, which converts it back into audio data.
[0054] Step 6:
[0055] The server transmits converted audio data to the terminal. The terminal plays the received data to the user through its speaker. This allows the user to hear their next appointment.
[0056] (Example 1)
[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0058] Modern information systems demand efficient task management and information access from users. However, conventional systems have struggled to centrally and automatically perform diverse tasks such as voice control, application program suggestions, and image processing. This has resulted in a lack of convenience for users and a failure to improve their quality of life.
[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0060] In this invention, the server includes information processing means for acquiring voice information and converting it into text information, recommendation generation means for recording the user's usage history and suggesting appropriate application programs to the user based on that history, and image processing means for receiving image information and processing it automatically. This makes it possible for the user to efficiently manage a variety of tasks and improve convenience.
[0061] "Audio information" refers to the data format obtained from audio signals, and represents acoustic information.
[0062] "Information processing means" refers to a method or device for acquiring audio information and converting it into text information.
[0063] "Textual information" refers to data expressed as text, representing the result of conversion from audio information.
[0064] "Analysis means" refers to a method or device for identifying the user's intent from textual information and generating an appropriate response.
[0065] "Response" refers to response information generated by the analysis tool that corresponds to the user's intent.
[0066] "Audio output means" refers to a method or device for reproducing the response generated by the analysis means as audio.
[0067] "User usage history" refers to a record of operations and activities performed by the user within the system.
[0068] "Recommendation generation means" refers to a method or apparatus for suggesting an appropriate application program based on the user's usage history.
[0069] "Image information" refers to visual data and includes image data expressed in digital format.
[0070] "Image processing means" refers to a method or device for receiving image information and automatically editing or adjusting it.
[0071] "Communication device" refers to a network connection device used to transmit voice and text information to a server.
[0072] A "remote device" refers to an external computing resource used to process data received from a user device and transmit the results.
[0073] This invention is a system that enables voice information processing, automatic suggestion of application programs, and image information processing, with the aim of improving user convenience. The system consists mainly of a terminal and a server, each performing a specific function and operating through mutual communication.
[0074] Speech information processing
[0075] When a user speaks into the device, the device converts the voice information into a digital format. Commonly available voice recognition software can be used for speech recognition. The digital voice information is transmitted to a server via a communication device. The server uses advanced natural language processing software to convert the voice information into text and analyze the user's intent. The analyzed information is then converted into an appropriate response by a response generation algorithm and transmitted back to the device as voice information via a speech synthesizer.
[0076] Automated application suggestions
[0077] The terminal continuously records the user's application program usage history. The recorded history data is sent to a server, where machine learning algorithms analyze the user's behavior patterns. Based on this analysis, the server recommends suitable application programs for the user. The recommended application programs are displayed on the user's screen via the terminal.
[0078] Image information processing
[0079] When a user takes a photo with their device, the image information is sent to a server via a cloud service. The server automatically adjusts the brightness, contrast, and color using image processing algorithms. The edited image is then sent back to the device and immediately saved to the gallery.
[0080] Examples of specific cases and prompt statements
[0081] For example, if a user asks "What's on my schedule tomorrow?", the system will respond with calendar information via voice. By entering a prompt such as "Please analyze the user's statement to determine the next appointment and provide a voice response," the system can generate an appropriate response.
[0082] This invention is expected to enable users to perform everyday tasks more efficiently, thereby improving their quality of life.
[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0084] Step 1:
[0085] The user gives voice commands to the device. The device uses a microphone to convert the voice information into a digital format and acquires the digital signal as data. The input is a voice signal, and the output is digital voice data. This digital voice data is sent to the server via the internet. For example, the user might say, "Tell me my next appointment."
[0086] Step 2:
[0087] The server converts the received digital audio data into text using speech recognition technology. Next, it uses a natural language processing algorithm to analyze the user's intent from the text. The input is digital audio data, and the output is the analyzed text. Specifically, the server extracts the text "Tell me my next appointment" and recognizes that this indicates the user's intention to check their calendar.
[0088] Step 3:
[0089] The server generates an appropriate response by applying a response generation algorithm based on the analyzed text information. The generated response is converted into speech data using a speech synthesis engine. The input is the analyzed text information, and the output is the generated speech data. Specifically, the server generates the response "Tomorrow's schedule is a meeting at 10 AM" and converts it into speech data.
[0090] Step 4:
[0091] The server sends a synthesized speech response to the terminal. The terminal plays the received audio data through its speaker and outputs it in a format audible to the user. The input is the generated audio data, and the output is the playback of the audio. Specifically, the terminal delivers the audio message "Tomorrow's schedule includes a meeting at 10:00 AM" to the user.
[0092] Step 5:
[0093] The terminal records the user's application program usage history and sends this history to the server. The server analyzes the usage history using a machine learning algorithm and determines which application program should be recommended next. The input is the usage history data, and the output is the recommended application. Specifically, the server generates an instruction such as, "Since you use the weather app every morning, we recommend launching it tomorrow morning."
[0094] Step 6:
[0095] The user sends images taken with their device to the server via the cloud. The server applies image processing algorithms to automatically edit the images. The input is the captured image, and the output is the processed image data. Specifically, the server performs a process such as "brightening a photo of a cloudy sky and improving its quality."
[0096] Step 7:
[0097] The server sends the processed image back to the terminal. The terminal saves the received image to its gallery and notifies the user so they can confirm it. The input is the processed image data, and the output is the image saved in the gallery. Specifically, the terminal saves the edited image and notifies the user that "image editing is complete."
[0098] (Application Example 1)
[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0100] Modern consumers, when purchasing goods in physical stores, demand quick access to product information and personalized product recommendations based on their individual preferences. However, conventional systems have struggled to efficiently achieve this, resulting in lower customer satisfaction. This invention aims to solve these problems and provide a richer shopping experience.
[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0102] In this invention, the server includes a natural language processing device that acquires voice information and converts it into text information, an analysis device that analyzes the user's purpose from the text information and generates a response, and a suggestion generation device that records the user's behavior history and makes suggestions. As a result, when users search for products in a store, they can obtain product information by voice operation, and furthermore, products are recommended individually based on their past purchase history, enabling an efficient and personalized shopping experience.
[0103] "Audio information" refers to information acquired as audio data, specifically the content of the user's speech.
[0104] "Textual information" refers to text data obtained by converting audio information using natural language processing.
[0105] A "natural language processing device" is a device that analyzes spoken information and converts it into written information.
[0106] An "analysis device" is a device that reads the user's purpose and intention from textual information and generates an appropriate response.
[0107] A "suggestion generation device" is a device that suggests the next appropriate application or product based on the user's behavioral history.
[0108] "Image information" refers to visual data, including still images and moving images, that is acquired or processed in digital format.
[0109] An "image editing device" is a device that receives image information and automatically adjusts brightness, color tone, and other parameters.
[0110] A "location identification device" is a device that provides location information of products within a physical store, assisting users in accessing those products.
[0111] A "recommendation device" is a device that recommends suitable products or services based on the user's purchase history and preferences.
[0112] A "communication device" is a device used to send and receive data between a server and a device, and it exchanges information via a network such as the internet.
[0113] This invention is an information processing system for improving the customer purchasing experience in physical stores, and has the following specific configuration and operation.
[0114] First, the server acquires voice information and converts it into text information using a natural language processing unit. This process utilizes the Google® Cloud Speech-to-Text API, which uses speech recognition technology to transcribe customer voice commands into text. Next, the transcribed text information is processed by an analysis unit to identify the user's purpose and request. The analysis unit uses the Google Cloud Natural Language API, which understands the context and generates an appropriate response.
[0115] The generated response is presented to the user via the device in the form of audio or text. This response includes information and guidance about products in the store.
[0116] Furthermore, the suggestion generation device utilizes the user's behavioral history to recommend products suitable for the user. Additionally, the location identification device searches for products within the store and provides customers with visual or audio guidance. This incorporates location services using Bluetooth and Wi-Fi to accurately determine the placement of products within the store.
[0117] The image editing device automatically adjusts the brightness and color tone when the user takes a photo of a product. This process utilizes OpenCV, a Python library, to enhance the appeal of the photo.
[0118] In a physical store setting, a customer might ask their smart glasses, "What's on sale next?" The system analyzes the question and displays a list of sale items on the glasses' screen. Furthermore, if the system analyzes that the user frequently purchases chocolate, it might recommend a newly released chocolate product as the next featured item.
[0119] An example of a prompt to input into the generative AI model would be: "Assume a user is looking for a specific product in a store, then simulate a conversation that identifies the location of this product and presents relevant information based on the user's past interests in that product." This prompt allows the system to dynamically provide the user with specific and relevant information.
[0120] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0121] Step 1:
[0122] When a user asks a product question by voice into the device, the device acquires that voice information through its microphone. This voice information is stored on the device as audio data in a digital format.
[0123] Step 2:
[0124] The device converts the acquired audio information into text information. During this process, it uses speech recognition technology to send data to the Google Cloud Speech-to-Text API, which analyzes the audio data and generates text data. The output is the text information corresponding to the original audio information.
[0125] Step 3:
[0126] The server sends the received text information to a parser for analysis to understand the user's intent. Specifically, it uses the Google Cloud Natural Language API to grasp the context and meaning and generate an appropriate response to the user's question. The output is response data based on the analyzed information.
[0127] Step 4:
[0128] The server sends the generated response data to the terminal. The terminal receives the response data and presents it to the user as text or audio. Information based on the user's questions is provided visually or audibly.
[0129] Step 5:
[0130] Simultaneously, a suggestion generator operates using the user's behavioral history data. The server analyzes the user's past behavior and, based on that behavior, analyzes products the user is likely to purchase next. This data processing is performed using machine learning algorithms. The output is product information recommended to the user.
[0131] Step 6:
[0132] When a user takes a picture of a product using their device's camera, the image information is saved on the device and sent to the image editing system on the server. Image editing processing is performed using OpenCV, and the brightness and contrast of the image are automatically adjusted. The output is the edited image data.
[0133] Step 7:
[0134] The edited image data is sent back to the device. Users can review the edited image and, if necessary, post reviews or share it on social media. This entire process improves the user's purchasing experience.
[0135] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0136] This invention is a system that combines a natural language processing function for handling user voice data with an emotion recognition function for recognizing user emotions. This system understands the user's emotional state and selects and provides appropriate responses and content accordingly, thereby realizing a more personalized service.
[0137] Voice input and emotion recognition
[0138] The device captures the user's voice and sends the audio data to the server. The server converts the received audio data into text data using a natural language processing engine and analyzes its content. At the same time, it uses emotion recognition to analyze supplementary information such as intonation, speed, and intensity of the voice and identify the user's emotions. For example, if the user says, "I was very busy and tired today," the emotion engine will read emotions such as fatigue and stress from this voice.
[0139] Generating responses and proposals
[0140] The server generates a response based on the analyzed text and sentiment data. If the text analysis suggests schedule management and the sentiment analysis indicates user fatigue, the server converts the response into speech, such as, "Let's make sure you get enough rest today. Shall I remind you of your schedule?"
[0141] Furthermore, it records user emotional data and suggests appropriate applications and content based on that data. For example, it can suggest relaxing music content to users who are experiencing high levels of fatigue.
[0142] Image editing function
[0143] Images taken by the user with their device are sent to a server, which automatically edits them using image editing tools. The editing process may include adding appropriate filters and effects, taking emotional data into consideration. For example, if a user edits an image that appears depressed, a brightening of the tone may be applied.
[0144] This system aims to improve the user experience by instantly providing appropriate responses that meet the user's needs and emotions.
[0145] The following describes the processing flow.
[0146] Step 1:
[0147] The user speaks into their smartphone, saying, "Today was a long day." The device collects the voice data and sends it to a server.
[0148] Step 2:
[0149] The server converts the received audio data into text data using a natural language processing engine. During this process, it extracts audio features and passes them to an emotion recognition engine.
[0150] Step 3:
[0151] The server's emotion recognition engine analyzes the intonation and speed of the speech and determines that the user is experiencing fatigue. This result is then sent as feedback to the natural language processing engine.
[0152] Step 4:
[0153] The server analyzes text data to understand the user's intent. Simultaneously, it determines an appropriate response to the user based on sentiment data.
[0154] Step 5:
[0155] The server generates the response, "Why not listen to some music to relax today?", and then converts it into audio data using a text-to-speech module.
[0156] Step 6:
[0157] The server sends the converted audio data to the terminal. The terminal plays this audio for the user and presents the suggested content.
[0158] Step 7:
[0159] When a user receives a suggestion for relaxing music and selects a piece of content, the device sends that information to the server and opens the appropriate application. The server then suggests additional content as needed.
[0160] (Example 2)
[0161] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0162] Modern information processing systems are required to quickly and accurately analyze user intent based on voice data and provide personalized responses that understand emotions. However, conventional systems have insufficient accuracy in voice recognition and emotion recognition, and lack the ability to utilize various data to improve the user experience. Furthermore, it has been difficult to perform content suggestions and image editing processing based on user emotions in real time. Against this backdrop, the present invention aims to realize the provision of comprehensive services based on aggregated user behavior history and emotion data.
[0163] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0164] In this invention, the server includes language processing means for acquiring voice information and converting it into text information, information analysis means for analyzing the user's intent from the text information and generating an appropriate response, and emotion recognition means for analyzing the intonation and speed of the voice to identify emotions. This makes it possible to accurately analyze the user's intent and emotions from their voice and provide responses, services, and product suggestions based on that analysis in real time.
[0165] "Voice information" refers to a data format in which the content spoken by a user is recorded and processed as a digital signal.
[0166] "Textual information" refers to string data converted from audio information, which is further analyzed by language processing tools.
[0167] "Language processing means" refers to a system or method for converting speech information into text information and analyzing its content.
[0168] "Information analysis means" refers to technologies that analyze the user's intent and context based on textual information and generate the optimal response.
[0169] "Emotion recognition methods" refer to methods that analyze non-verbal elements such as intonation, speed, and intensity of speech to identify the user's emotional state.
[0170] "Conversion means" refers to a system that converts the response generated based on the analysis results back into an audio format and delivers it to the user.
[0171] "Suggestion generation means" refers to technology that suggests appropriate services and content to users based on recorded user behavior history and emotional data.
[0172] "Image processing means" refers to a system that automatically edits received image information and, in some cases, applies filters or effects that reflect emotional data.
[0173] A description of embodiments for carrying out the present invention will be provided.
[0174] The system of this invention has the function of performing natural language processing using the user's voice information and recognizing their emotional state. The system mainly consists of a terminal used by the user and a server that processes its data.
[0175] First, when the user speaks into the device, the device captures the audio information. The device has a communication function that uses a microphone to capture high-quality audio and transmits the information to a remote server. Specific examples include smartphones and tablets.
[0176] When the server receives voice information from a terminal, it converts it into text using speech recognition technology. Google's voice services or open-source speech recognition engines may be used for this purpose. The converted text information is then analyzed by a natural language processing engine to identify the user's intent. Next, an emotion recognition engine analyzes the intonation and speed of the voice to reveal the user's emotions. Tools specifically designed for emotion analysis are used in these processes.
[0177] Next, the server generates a response to the user based on the analysis data obtained. The generated response is converted back into speech format using speech synthesis and fed back to the user from the terminal. The speech synthesis technology used here includes commercial or open-source speech synthesis engines. For example, if the user says, "I'm very busy and tired today," the server can generate a response such as, "Please take a good rest. Shall I remind you of your appointment?"
[0178] Additionally, the system includes a feature that accumulates user behavior history and emotional data, and suggests content based on this data. When a user sends an image they have taken to the server, the server automatically processes the image using its image editing function. This function utilizes image editing software, applying filters according to the user's emotional state.
[0179] Because the responses and suggestions generated by this system are based on user data, it can provide a personalized experience.
[0180] For example, one could use the following prompt on a generative AI model: "Analyze the user's intent and emotions from their voice input and provide rest-related suggestions based on that." This would allow the system to provide the most appropriate response within the user's context.
[0181] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0182] Step 1:
[0183] The user provides voice input to the device. The device captures this voice using its microphone and saves it as digital audio data. The input is analog audio data, and the output is digital audio data. The device improves the sound quality of this data through noise reduction and sampling.
[0184] Step 2:
[0185] The terminal sends the captured digital audio data to the server. The server receives this data and converts it into text using a speech recognition engine. The input is digital audio data, and the output is text. In this conversion process, the server uses a speech recognition API to convert audio to text.
[0186] Step 3:
[0187] The server passes textual information to a natural language processing engine for analysis. This analysis identifies the user's intent and requirements. The input is textual information, and the output is analyzed intent data. The server uses a text analysis algorithm to extract intent from grammar and keywords.
[0188] Step 4:
[0189] Simultaneously, the server uses an emotion recognition engine that analyzes emotions from the intonation and speed of the audio data. The input is audio data, and the output is emotion data. The server applies signal processing technology to identify the tone and tempo of the voice as an emotion pattern.
[0190] Step 5:
[0191] The server generates an appropriate response using a generative AI model based on the analyzed intent and sentiment data. The input is intent and sentiment data, and the output is the response text. The generative AI model takes the prompt sentence as a variable and calculates the response.
[0192] Step 6:
[0193] The generated response text is converted into speech through a speech synthesis engine and sent back to the terminal. The input is the response text, and the output is synthesized speech data. The server uses speech synthesis technology to generate natural-sounding speech.
[0194] Step 7:
[0195] Ultimately, the terminal outputs the audio data received from the server to the user. The input is synthesized speech data, and the output is an audio response to the user. The terminal plays the audio message using its speaker.
[0196] (Application Example 2)
[0197] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0198] Traditional content delivery systems provide uniform content without considering user emotions, resulting in a limited user experience and a lack of individually personalized services. Furthermore, they lack responses and content tailored to the user's emotional state, highlighting the need for higher-quality user satisfaction.
[0199] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0200] In this invention, the server includes natural language processing means for acquiring voice data and converting it into text data; analysis means for analyzing the user's intent and emotions from the text data and voice data and generating appropriate responses and content; and content recommendation means for presenting content appropriate to the user based on the analyzed emotions. This makes it possible to provide personalized responses and content that correspond to the user's emotional state.
[0201] "Natural language processing means" refers to devices or software that convert speech data into text data, and is a function used to analyze the content of a user's speech.
[0202] "Analysis means" refers to devices or programs that analyze the user's intent and emotions from text data and audio data, and generate appropriate responses and content based on this analysis.
[0203] "Content recommendation methods" refer to functions and system configurations that present users with the most suitable content based on analyzed emotions.
[0204] "Suggestion generation means" refers to a function that suggests appropriate applications and services based on the user's behavioral history and emotional data.
[0205] "Image editing means" refers to devices or software that automatically edit image data received from users and apply filters and effects that correspond to emotions.
[0206] "Communication means" refers to devices and technologies that have the function of transmitting voice data, text data, and sentiment data to a server via a communication device.
[0207] "Behavioral analysis means" refers to technologies and devices for aggregating user behavior history and emotional data over a certain period and analyzing the results of that aggregation.
[0208] The system that implements this application works by directly using the user's voice data and converting it into text data using natural language processing technology. The server analyzes the voice input using dedicated natural language processing software and extracts the user's intent and emotions from the text data. For this process, libraries such as "speech_recognition" and emotion analysis engines such as "DeepAffects" can be used as cloud services.
[0209] Furthermore, the server suggests the most suitable content to the user based on the results of emotion recognition. For example, if emotion analysis indicates that the user is tired, relaxing music or videos will be recommended preferentially. To achieve this, a content recommendation engine works to select content that matches the user's needs. The "content_recommendation" module efficiently handles this.
[0210] A specific scenario would be that when a user voice-inputs "I'm very tired today," the system analyzes this to recognize the user's fatigue level and selects relaxing videos or music. A possible prompt message would be something like, "The user may need rest. Please suggest video content related to relaxation and well-being."
[0211] The device collects and transmits voice data, and then transfers the data to the server via communication means. Through this process, data analysis and content provision are carried out quickly, thereby improving the user experience.
[0212] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0213] Step 1:
[0214] The device acquires audio data from the user. This audio data is collected as an audio waveform, capturing the words the user speaks directly. For example, the device uses a microphone to capture ambient sound and converts that data into a digital signal.
[0215] Step 2:
[0216] The terminal transmits the acquired audio data to the server. Here, data is transmitted in real time over the network using communication methods, preparing the server for audio analysis. The input is audio data in binary format, and the output is data transfer over the network.
[0217] Step 3:
[0218] The server converts the received audio data into text data using a natural language processing engine. This process uses the "speech_recognition" library to perform speech recognition and convert the audio waveform into corresponding sentences. The input is audio data, and the output is text data.
[0219] Step 4:
[0220] The server performs emotion recognition based on text and audio data. This process analyzes specific keywords, speech intonation, and speed, and uses an emotion recognition engine to identify the user's emotions. The output is data indicating the user's emotions.
[0221] Step 5:
[0222] The server uses the analysis results to generate appropriate responses and content. It utilizes a generative AI model to construct responses tailored to the user's intent and a content recommendation engine to select relevant videos and music. The output consists of a response message and recommended content.
[0223] Step 6:
[0224] The server formats the generated response as audio data and sends it to the terminal. At this stage, text-to-speech conversion occurs and is returned to the terminal via the network. The output is response data in audio format.
[0225] Step 7:
[0226] The terminal plays the audio data received from the server and provides it to the user. Here, it uses its speaker to output an audio response, allowing the user to learn about the suggested content. The process concludes when audio playback is complete.
[0227] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0228] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0229] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0230] [Second Embodiment]
[0231] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0232] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0233] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0234] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0235] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0236] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0237] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0238] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0239] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0240] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0241] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0242] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0243] This invention provides a system incorporating various functions to improve user convenience. The system primarily performs natural language processing for handling voice data, application suggestions based on user behavior history, and automatic image data editing. The detailed operation of each function is described below.
[0244] Voice control using natural language processing
[0245] When a user speaks into their smartphone using natural language, the device converts the audio into digital data and sends it to a server. The server uses advanced natural language processing technology to convert the audio into text data and analyze its content. For example, if a user asks, "What time is the next meeting?", the server refers to calendar information and derives the answer. Then, it converts the analysis result back into audio and outputs it from the device in a format that is easy for the user to understand.
[0246] Automated application suggestions
[0247] The device records the user's application usage history and periodically sends it to the server. Based on this history, the server analyzes the user's behavior patterns and predicts the next most suitable application. For example, if a user has a habit of opening the weather app at 7 AM every day, the server instructs the device to display the weather app on the home screen in the morning. In this way, the system helps save the user time.
[0248] Automated image editing
[0249] Images taken by the user are sent to the server via the device. The server utilizes image editing algorithms to adjust the brightness, contrast, and color of the images, automating high-quality editing. For example, a photo of a cloudy beach taken during a trip can be transformed into a bright and vibrant image. The edited image is then sent back from the server to the device and immediately saved to the gallery.
[0250] This system aims to improve the quality of life by enabling users to complete daily tasks more efficiently.
[0251] The following describes the processing flow.
[0252] Step 1:
[0253] The user speaks into their smartphone, saying, "Tell me my next appointment." The device acquires the voice data through the microphone and temporarily stores it locally.
[0254] Step 2:
[0255] The terminal sends the stored audio data to the server using a communication module. The server then passes the received audio data to a natural language processing engine.
[0256] Step 3:
[0257] The server uses a natural language processing engine to convert the audio data into text data. This text data then serves as input data for analyzing the user's intent.
[0258] Step 4:
[0259] The server parses the converted text data and retrieves the user's next appointment by referencing the calendar database. After generating an appropriate response, it formats the response back into text format.
[0260] Step 5:
[0261] The server sends the generated text response to a text-to-speech module, which converts it back into audio data.
[0262] Step 6:
[0263] The server transmits converted audio data to the terminal. The terminal plays the received data to the user through its speaker. This allows the user to hear their next appointment.
[0264] (Example 1)
[0265] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0266] Modern information systems demand efficient task management and information access from users. However, conventional systems have struggled to centrally and automatically perform diverse tasks such as voice control, application program suggestions, and image processing. This has resulted in a lack of convenience for users and a failure to improve their quality of life.
[0267] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0268] In this invention, the server includes information processing means for acquiring voice information and converting it into text information, recommendation generation means for recording the user's usage history and suggesting appropriate application programs to the user based on that history, and image processing means for receiving image information and processing it automatically. This makes it possible for the user to efficiently manage a variety of tasks and improve convenience.
[0269] "Audio information" refers to the data format obtained from audio signals, and represents acoustic information.
[0270] "Information processing means" refers to a method or device for acquiring audio information and converting it into text information.
[0271] "Textual information" refers to data expressed as text, representing the result of conversion from audio information.
[0272] "Analysis means" refers to a method or device for identifying the user's intent from textual information and generating an appropriate response.
[0273] "Response" refers to response information generated by the analysis tool that corresponds to the user's intent.
[0274] "Audio output means" refers to a method or device for reproducing the response generated by the analysis means as audio.
[0275] "User usage history" refers to a record of operations and activities performed by the user within the system.
[0276] "Recommendation generation means" refers to a method or apparatus for suggesting an appropriate application program based on the user's usage history.
[0277] "Image information" refers to visual data and includes image data expressed in digital format.
[0278] "Image processing means" refers to a method or device for receiving image information and automatically editing or adjusting it.
[0279] "Communication device" refers to a network connection device used to transmit voice and text information to a server.
[0280] A "remote device" refers to an external computing resource used to process data received from a user device and transmit the results.
[0281] This invention is a system that enables voice information processing, automatic suggestion of application programs, and image information processing, with the aim of improving user convenience. The system consists mainly of a terminal and a server, each performing a specific function and operating through mutual communication.
[0282] Speech information processing
[0283] When a user speaks into the device, the device converts the voice information into a digital format. Commonly available voice recognition software can be used for speech recognition. The digital voice information is transmitted to a server via a communication device. The server uses advanced natural language processing software to convert the voice information into text and analyze the user's intent. The analyzed information is then converted into an appropriate response by a response generation algorithm and transmitted back to the device as voice information via a speech synthesizer.
[0284] Automatic Recommendation of Applications
[0285] The terminal continuously records the user's application usage history. The recorded history data is sent to the server, and the user's behavior pattern is analyzed by a machine learning algorithm. Based on this analysis result, the server recommends an application suitable for the user. The recommended application is displayed on the user's screen through the terminal.
[0286] Processing of Image Information
[0287] When the user takes a photo with the terminal, the image information is sent to the server via the cloud service. The server automatically adjusts the brightness, contrast, and color using an image processing algorithm. The edited image is sent back to the terminal and immediately saved in the gallery.
[0288] Examples of Specific Examples and Prompt Sentences
[0289] For example, when the user asks "What are my plans for tomorrow?" verbally, the system responds with the calendar information verbally. By inputting, for example, "Analyze the next schedule from the user's speech and answer verbally." as a prompt sentence, it becomes possible for the system to generate an appropriate response.
[0290] According to this invention, it is expected that the user can efficiently carry out daily tasks and the quality of life will be improved.
[0291] The flow of the specific process in Example 1 will be described using FIG. 11.
[0292] Step 1:
[0293] The user gives voice commands to the device. The device uses a microphone to convert the voice information into a digital format and acquires the digital signal as data. The input is a voice signal, and the output is digital voice data. This digital voice data is sent to the server via the internet. For example, the user might say, "Tell me my next appointment."
[0294] Step 2:
[0295] The server converts the received digital audio data into text using speech recognition technology. Next, it uses a natural language processing algorithm to analyze the user's intent from the text. The input is digital audio data, and the output is the analyzed text. Specifically, the server extracts the text "Tell me my next appointment" and recognizes that this indicates the user's intention to check their calendar.
[0296] Step 3:
[0297] The server generates an appropriate response by applying a response generation algorithm based on the analyzed text information. The generated response is converted into speech data using a speech synthesis engine. The input is the analyzed text information, and the output is the generated speech data. Specifically, the server generates the response "Tomorrow's schedule is a meeting at 10 AM" and converts it into speech data.
[0298] Step 4:
[0299] The server sends a synthesized speech response to the terminal. The terminal plays the received audio data through its speaker and outputs it in a format audible to the user. The input is the generated audio data, and the output is the playback of the audio. Specifically, the terminal delivers the audio message "Tomorrow's schedule includes a meeting at 10:00 AM" to the user.
[0300] Step 5:
[0301] The terminal records the user's application usage history and sends the history to the server. The server analyzes the usage history using a machine learning algorithm and then determines the application program to be recommended next. The input is the usage history data, and the output is the recommended application. As a specific operation, the server generates an instruction such as "Since the weather application is used every morning, it is recommended to start it the next morning."
[0302] Step 6:
[0303] The user sends the image taken by the terminal to the server via the cloud. The server applies an image processing algorithm to perform automatic editing of the image. The input is the captured image, and the output is the processed image data. As a specific operation, the server performs a process such as "Adjust the photo of the cloudy sky to be brighter and of high quality."
[0304] Step 7:
[0305] The server sends the processed image back to the terminal. The terminal saves the received image in the gallery and notifies the user so that the user can view it. The input is the processed image data, and the output is the image saved in the gallery. As a specific operation, the terminal saves the edited image and notifies the user with "The editing of the image has been completed."
[0306] (Application Example 1)
[0307] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0308] When modern consumers purchase products at physical stores, they seek quick access to product information and product recommendations based on their individual preferences. However, it has been difficult to efficiently achieve this with conventional systems, which has not led to an improvement in customer satisfaction. The present invention aims to solve these problems and provide a richer purchasing experience.
[0309] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0310] In this invention, the server includes a natural language processing device that acquires voice information and converts it into text information, an analysis device that analyzes the user's purpose from the text information and generates a response, and a suggestion generation device that records the user's behavior history and makes suggestions. As a result, when users search for products in a store, they can obtain product information by voice operation, and furthermore, products are recommended individually based on their past purchase history, enabling an efficient and personalized shopping experience.
[0311] "Audio information" refers to information acquired as audio data, specifically the content of the user's speech.
[0312] "Textual information" refers to text data obtained by converting audio information using natural language processing.
[0313] A "natural language processing device" is a device that analyzes spoken information and converts it into written information.
[0314] An "analysis device" is a device that reads the user's purpose and intention from textual information and generates an appropriate response.
[0315] A "suggestion generation device" is a device that suggests the next appropriate application or product based on the user's behavioral history.
[0316] "Image information" refers to visual data, including still images and moving images, that is acquired or processed in digital format.
[0317] An "image editing device" is a device that receives image information and automatically adjusts brightness, color tone, and other parameters.
[0318] A "location identification device" is a device that provides location information of products within a physical store, assisting users in accessing those products.
[0319] A "recommendation device" is a device that recommends suitable products or services based on the user's purchase history and preferences.
[0320] A "communication device" is a device used to send and receive data between a server and a device, and it exchanges information via a network such as the internet.
[0321] This invention is an information processing system for improving the customer purchasing experience in physical stores, and has the following specific configuration and operation.
[0322] First, the server acquires voice information and converts it into text using a natural language processing unit. This process utilizes the Google Cloud Speech-to-Text API, which uses speech recognition technology to transcribe customer voice commands into text. Next, the transcribed text is processed by an analyzer to identify the user's purpose and request. The analyzer uses the Google Cloud Natural Language API, which understands the context and generates an appropriate response.
[0323] The generated response is presented to the user via the device in the form of audio or text. This response includes information and guidance about products in the store.
[0324] Furthermore, the suggestion generation device utilizes the user's behavioral history to recommend products suitable for the user. Additionally, the location identification device searches for products within the store and provides customers with visual or audio guidance. This incorporates location services using Bluetooth and Wi-Fi to accurately determine the placement of products within the store.
[0325] The image editing device automatically adjusts the brightness and color tone when the user takes a photo of a product. This process utilizes OpenCV, a Python library, to enhance the appeal of the photo.
[0326] In a physical store setting, a customer might ask their smart glasses, "What's on sale next?" The system analyzes the question and displays a list of sale items on the glasses' screen. Furthermore, if the system analyzes that the user frequently purchases chocolate, it might recommend a newly released chocolate product as the next featured item.
[0327] An example of a prompt to input into the generative AI model would be: "Assume a user is looking for a specific product in a store, then simulate a conversation that identifies the location of this product and presents relevant information based on the user's past interests in that product." This prompt allows the system to dynamically provide the user with specific and relevant information.
[0328] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0329] Step 1:
[0330] When a user asks a product question by voice into the device, the device acquires that voice information through its microphone. This voice information is stored on the device as audio data in a digital format.
[0331] Step 2:
[0332] The device converts the acquired audio information into text information. During this process, it uses speech recognition technology to send data to the Google Cloud Speech-to-Text API, which analyzes the audio data and generates text data. The output is the text information corresponding to the original audio information.
[0333] Step 3:
[0334] The server sends the received text information to a parser for analysis to understand the user's intent. Specifically, it uses the Google Cloud Natural Language API to grasp the context and meaning and generate an appropriate response to the user's question. The output is response data based on the analyzed information.
[0335] Step 4:
[0336] The server sends the generated response data to the terminal. The terminal receives the response data and presents it to the user as text or audio. Information based on the user's questions is provided visually or audibly.
[0337] Step 5:
[0338] Simultaneously, a suggestion generator operates using the user's behavioral history data. The server analyzes the user's past behavior and, based on that behavior, analyzes products the user is likely to purchase next. This data processing is performed using machine learning algorithms. The output is product information recommended to the user.
[0339] Step 6:
[0340] When a user takes a picture of a product using their device's camera, the image information is saved on the device and sent to the image editing system on the server. Image editing processing is performed using OpenCV, and the brightness and contrast of the image are automatically adjusted. The output is the edited image data.
[0341] Step 7:
[0342] The edited image data is sent back to the device. Users can review the edited image and, if necessary, post reviews or share it on social media. This entire process improves the user's purchasing experience.
[0343] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0344] This invention is a system that combines a natural language processing function for handling user voice data with an emotion recognition function for recognizing user emotions. This system understands the user's emotional state and selects and provides appropriate responses and content accordingly, thereby realizing a more personalized service.
[0345] Voice input and emotion recognition
[0346] The device captures the user's voice and sends the audio data to the server. The server converts the received audio data into text data using a natural language processing engine and analyzes its content. At the same time, it uses emotion recognition to analyze supplementary information such as intonation, speed, and intensity of the voice and identify the user's emotions. For example, if the user says, "I was very busy and tired today," the emotion engine will read emotions such as fatigue and stress from this voice.
[0347] Generating responses and proposals
[0348] The server generates a response based on the analyzed text and sentiment data. If the text analysis suggests schedule management and the sentiment analysis indicates user fatigue, the server converts the response into speech, such as, "Let's make sure you get enough rest today. Shall I remind you of your schedule?"
[0349] Furthermore, it records user emotional data and suggests appropriate applications and content based on that data. For example, it can suggest relaxing music content to users who are experiencing high levels of fatigue.
[0350] Image editing function
[0351] Images taken by the user with their device are sent to a server, which automatically edits them using image editing tools. The editing process may include adding appropriate filters and effects, taking emotional data into consideration. For example, if a user edits an image that appears depressed, a brightening of the tone may be applied.
[0352] This system aims to improve the user experience by instantly providing appropriate responses that meet the user's needs and emotions.
[0353] The following describes the processing flow.
[0354] Step 1:
[0355] The user speaks into their smartphone, saying, "Today was a long day." The device collects the voice data and sends it to a server.
[0356] Step 2:
[0357] The server converts the received audio data into text data using a natural language processing engine. During this process, it extracts audio features and passes them to an emotion recognition engine.
[0358] Step 3:
[0359] The server's emotion recognition engine analyzes the intonation and speed of the speech and determines that the user is experiencing fatigue. This result is then sent as feedback to the natural language processing engine.
[0360] Step 4:
[0361] The server analyzes text data to understand the user's intent. Simultaneously, it determines an appropriate response to the user based on sentiment data.
[0362] Step 5:
[0363] The server generates the response, "Why not listen to some music to relax today?", and then converts it into audio data using a text-to-speech module.
[0364] Step 6:
[0365] The server sends the converted audio data to the terminal. The terminal plays this audio for the user and presents the suggested content.
[0366] Step 7:
[0367] When a user receives a suggestion for relaxing music and selects a piece of content, the device sends that information to the server and opens the appropriate application. The server then suggests additional content as needed.
[0368] (Example 2)
[0369] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0370] Modern information processing systems are required to quickly and accurately analyze user intent based on voice data and provide personalized responses that understand emotions. However, conventional systems have insufficient accuracy in voice recognition and emotion recognition, and lack the ability to utilize various data to improve the user experience. Furthermore, it has been difficult to perform content suggestions and image editing processing based on user emotions in real time. Against this backdrop, the present invention aims to realize the provision of comprehensive services based on aggregated user behavior history and emotion data.
[0371] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0372] In this invention, the server includes language processing means for acquiring voice information and converting it into text information, information analysis means for analyzing the user's intent from the text information and generating an appropriate response, and emotion recognition means for analyzing the intonation and speed of the voice to identify emotions. This makes it possible to accurately analyze the user's intent and emotions from their voice and provide responses, services, and product suggestions based on that analysis in real time.
[0373] "Voice information" refers to a data format in which the content spoken by a user is recorded and processed as a digital signal.
[0374] "Textual information" refers to string data converted from audio information, which is further analyzed by language processing tools.
[0375] "Language processing means" refers to a system or method for converting speech information into text information and analyzing its content.
[0376] "Information analysis means" refers to technologies that analyze the user's intent and context based on textual information and generate the optimal response.
[0377] "Emotion recognition methods" refer to methods that analyze non-verbal elements such as intonation, speed, and intensity of speech to identify the user's emotional state.
[0378] "Conversion means" refers to a system that converts the response generated based on the analysis results back into an audio format and delivers it to the user.
[0379] "Suggestion generation means" refers to technology that suggests appropriate services and content to users based on recorded user behavior history and emotional data.
[0380] "Image processing means" refers to a system that automatically edits received image information and, in some cases, applies filters or effects that reflect emotional data.
[0381] A description of embodiments for carrying out the present invention will be provided.
[0382] The system of this invention has the function of performing natural language processing using the user's voice information and recognizing their emotional state. The system mainly consists of a terminal used by the user and a server that processes its data.
[0383] First, when the user speaks into the device, the device captures the audio information. The device has a communication function that uses a microphone to capture high-quality audio and transmits the information to a remote server. Specific examples include smartphones and tablets.
[0384] When the server receives voice information from a terminal, it converts it into text using speech recognition technology. Google's voice services or open-source speech recognition engines may be used for this purpose. The converted text information is then analyzed by a natural language processing engine to identify the user's intent. Next, an emotion recognition engine analyzes the intonation and speed of the voice to reveal the user's emotions. Tools specifically designed for emotion analysis are used in these processes.
[0385] Next, the server generates a response to the user based on the analysis data obtained. The generated response is converted back into speech format using speech synthesis and fed back to the user from the terminal. The speech synthesis technology used here includes commercial or open-source speech synthesis engines. For example, if the user says, "I'm very busy and tired today," the server can generate a response such as, "Please take a good rest. Shall I remind you of your appointment?"
[0386] Additionally, the system includes a feature that accumulates user behavior history and emotional data, and suggests content based on this data. When a user sends an image they have taken to the server, the server automatically processes the image using its image editing function. This function utilizes image editing software, applying filters according to the user's emotional state.
[0387] Because the responses and suggestions generated by this system are based on user data, it can provide a personalized experience.
[0388] For example, one could use the following prompt on a generative AI model: "Analyze the user's intent and emotions from their voice input and provide rest-related suggestions based on that." This would allow the system to provide the most appropriate response within the user's context.
[0389] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0390] Step 1:
[0391] The user provides voice input to the device. The device captures this voice using its microphone and saves it as digital audio data. The input is analog audio data, and the output is digital audio data. The device improves the sound quality of this data through noise reduction and sampling.
[0392] Step 2:
[0393] The terminal sends the captured digital audio data to the server. The server receives this data and converts it into text using a speech recognition engine. The input is digital audio data, and the output is text. In this conversion process, the server uses a speech recognition API to convert audio to text.
[0394] Step 3:
[0395] The server passes textual information to a natural language processing engine for analysis. This analysis identifies the user's intent and requirements. The input is textual information, and the output is analyzed intent data. The server uses a text analysis algorithm to extract intent from grammar and keywords.
[0396] Step 4:
[0397] Simultaneously, the server uses an emotion recognition engine that analyzes emotions from the intonation and speed of the audio data. The input is audio data, and the output is emotion data. The server applies signal processing technology to identify the tone and tempo of the voice as an emotion pattern.
[0398] Step 5:
[0399] The server generates an appropriate response using a generative AI model based on the analyzed intent and sentiment data. The input is intent and sentiment data, and the output is the response text. The generative AI model takes the prompt sentence as a variable and calculates the response.
[0400] Step 6:
[0401] The generated response text is converted into speech through a speech synthesis engine and sent back to the terminal. The input is the response text, and the output is synthesized speech data. The server uses speech synthesis technology to generate natural-sounding speech.
[0402] Step 7:
[0403] Ultimately, the terminal outputs the audio data received from the server to the user. The input is synthesized speech data, and the output is an audio response to the user. The terminal plays the audio message using its speaker.
[0404] (Application Example 2)
[0405] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0406] Traditional content delivery systems provide uniform content without considering user emotions, resulting in a limited user experience and a lack of individually personalized services. Furthermore, they lack responses and content tailored to the user's emotional state, highlighting the need for higher-quality user satisfaction.
[0407] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0408] In this invention, the server includes natural language processing means for acquiring voice data and converting it into text data; analysis means for analyzing the user's intent and emotions from the text data and voice data and generating appropriate responses and content; and content recommendation means for presenting content appropriate to the user based on the analyzed emotions. This makes it possible to provide personalized responses and content that correspond to the user's emotional state.
[0409] "Natural language processing means" refers to devices or software that convert speech data into text data, and is a function used to analyze the content of a user's speech.
[0410] "Analysis means" refers to devices or programs that analyze the user's intent and emotions from text data and audio data, and generate appropriate responses and content based on this analysis.
[0411] "Content recommendation methods" refer to functions and system configurations that present users with the most suitable content based on analyzed emotions.
[0412] "Suggestion generation means" refers to a function that suggests appropriate applications and services based on the user's behavioral history and emotional data.
[0413] "Image editing means" refers to devices or software that automatically edit image data received from users and apply filters and effects that correspond to emotions.
[0414] "Communication means" refers to devices and technologies that have the function of transmitting voice data, text data, and sentiment data to a server via a communication device.
[0415] "Behavioral analysis means" refers to technologies and devices for aggregating user behavior history and emotional data over a certain period and analyzing the results of that aggregation.
[0416] The system that implements this application works by directly using the user's voice data and converting it into text data using natural language processing technology. The server analyzes the voice input using dedicated natural language processing software and extracts the user's intent and emotions from the text data. For this process, libraries such as "speech_recognition" and emotion analysis engines such as "DeepAffects" can be used as cloud services.
[0417] Furthermore, the server suggests the most suitable content to the user based on the results of emotion recognition. For example, if emotion analysis indicates that the user is tired, relaxing music or videos will be recommended preferentially. To achieve this, a content recommendation engine works to select content that matches the user's needs. The "content_recommendation" module efficiently handles this.
[0418] A specific scenario would be that when a user voice-inputs "I'm very tired today," the system analyzes this to recognize the user's fatigue level and selects relaxing videos or music. A possible prompt message would be something like, "The user may need rest. Please suggest video content related to relaxation and well-being."
[0419] The device collects and transmits voice data, and then transfers the data to the server via communication means. Through this process, data analysis and content provision are carried out quickly, thereby improving the user experience.
[0420] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0421] Step 1:
[0422] The device acquires audio data from the user. This audio data is collected as an audio waveform, capturing the words the user speaks directly. For example, the device uses a microphone to capture ambient sound and converts that data into a digital signal.
[0423] Step 2:
[0424] The terminal transmits the acquired audio data to the server. Here, data is transmitted in real time over the network using communication methods, preparing the server for audio analysis. The input is audio data in binary format, and the output is data transfer over the network.
[0425] Step 3:
[0426] The server converts the received audio data into text data using a natural language processing engine. This process uses the "speech_recognition" library to perform speech recognition and convert the audio waveform into corresponding sentences. The input is audio data, and the output is text data.
[0427] Step 4:
[0428] The server performs emotion recognition based on text and audio data. This process analyzes specific keywords, speech intonation, and speed, and uses an emotion recognition engine to identify the user's emotions. The output is data indicating the user's emotions.
[0429] Step 5:
[0430] The server uses the analysis results to generate appropriate responses and content. It utilizes a generative AI model to construct responses tailored to the user's intent and a content recommendation engine to select relevant videos and music. The output consists of a response message and recommended content.
[0431] Step 6:
[0432] The server formats the generated response as audio data and sends it to the terminal. At this stage, text-to-speech conversion occurs and is returned to the terminal via the network. The output is response data in audio format.
[0433] Step 7:
[0434] The terminal plays the audio data received from the server and provides it to the user. Here, it uses its speaker to output an audio response, allowing the user to learn about the suggested content. The process concludes when audio playback is complete.
[0435] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0436] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0437] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0438] [Third Embodiment]
[0439] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0440] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0441] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0442] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0443] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0444] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0445] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0446] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0447] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0448] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0449] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0450] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0451] This invention provides a system incorporating various functions to improve user convenience. The system primarily performs natural language processing for handling voice data, application suggestions based on user behavior history, and automatic image data editing. The detailed operation of each function is described below.
[0452] Voice control using natural language processing
[0453] When a user speaks into their smartphone using natural language, the device converts the audio into digital data and sends it to a server. The server uses advanced natural language processing technology to convert the audio into text data and analyze its content. For example, if a user asks, "What time is the next meeting?", the server refers to calendar information and derives the answer. Then, it converts the analysis result back into audio and outputs it from the device in a format that is easy for the user to understand.
[0454] Automated application suggestions
[0455] The device records the user's application usage history and periodically sends it to the server. Based on this history, the server analyzes the user's behavior patterns and predicts the next most suitable application. For example, if a user has a habit of opening the weather app at 7 AM every day, the server instructs the device to display the weather app on the home screen in the morning. In this way, the system helps save the user time.
[0456] Automated image editing
[0457] Images taken by the user are sent to the server via the device. The server utilizes image editing algorithms to adjust the brightness, contrast, and color of the images, automating high-quality editing. For example, a photo of a cloudy beach taken during a trip can be transformed into a bright and vibrant image. The edited image is then sent back from the server to the device and immediately saved to the gallery.
[0458] This system aims to improve the quality of life by enabling users to complete daily tasks more efficiently.
[0459] The following describes the processing flow.
[0460] Step 1:
[0461] The user speaks into their smartphone, saying, "Tell me my next appointment." The device acquires the voice data through the microphone and temporarily stores it locally.
[0462] Step 2:
[0463] The terminal sends the stored audio data to the server using a communication module. The server then passes the received audio data to a natural language processing engine.
[0464] Step 3:
[0465] The server uses a natural language processing engine to convert the audio data into text data. This text data then serves as input data for analyzing the user's intent.
[0466] Step 4:
[0467] The server parses the converted text data and retrieves the user's next appointment by referencing the calendar database. After generating an appropriate response, it formats the response back into text format.
[0468] Step 5:
[0469] The server sends the generated text response to a text-to-speech module, which converts it back into audio data.
[0470] Step 6:
[0471] The server transmits converted audio data to the terminal. The terminal plays the received data to the user through its speaker. This allows the user to hear their next appointment.
[0472] (Example 1)
[0473] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0474] Modern information systems demand efficient task management and information access from users. However, conventional systems have struggled to centrally and automatically perform diverse tasks such as voice control, application program suggestions, and image processing. This has resulted in a lack of convenience for users and a failure to improve their quality of life.
[0475] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0476] In this invention, the server includes information processing means for acquiring voice information and converting it into text information, recommendation generation means for recording the user's usage history and suggesting appropriate application programs to the user based on that history, and image processing means for receiving image information and processing it automatically. This makes it possible for the user to efficiently manage a variety of tasks and improve convenience.
[0477] "Audio information" refers to the data format obtained from audio signals, and represents acoustic information.
[0478] "Information processing means" refers to a method or device for acquiring audio information and converting it into text information.
[0479] "Textual information" refers to data expressed as text, representing the result of conversion from audio information.
[0480] "Analysis means" refers to a method or device for identifying the user's intent from textual information and generating an appropriate response.
[0481] "Response" refers to response information generated by the analysis tool that corresponds to the user's intent.
[0482] "Audio output means" refers to a method or device for reproducing the response generated by the analysis means as audio.
[0483] "User usage history" refers to a record of operations and activities performed by the user within the system.
[0484] "Recommendation generation means" refers to a method or apparatus for suggesting an appropriate application program based on the user's usage history.
[0485] "Image information" refers to visual data and includes image data expressed in digital format.
[0486] "Image processing means" refers to a method or device for receiving image information and automatically editing or adjusting it.
[0487] "Communication device" refers to a network connection device used to transmit voice and text information to a server.
[0488] A "remote device" refers to an external computing resource used to process data received from a user device and transmit the results.
[0489] This invention is a system that enables voice information processing, automatic suggestion of application programs, and image information processing, with the aim of improving user convenience. The system consists mainly of a terminal and a server, each performing a specific function and operating through mutual communication.
[0490] Speech information processing
[0491] When a user speaks into the device, the device converts the voice information into a digital format. Commonly available voice recognition software can be used for speech recognition. The digital voice information is transmitted to a server via a communication device. The server uses advanced natural language processing software to convert the voice information into text and analyze the user's intent. The analyzed information is then converted into an appropriate response by a response generation algorithm and transmitted back to the device as voice information via a speech synthesizer.
[0492] Automated application suggestions
[0493] The terminal continuously records the user's application program usage history. The recorded history data is sent to a server, where machine learning algorithms analyze the user's behavior patterns. Based on this analysis, the server recommends suitable application programs for the user. The recommended application programs are displayed on the user's screen via the terminal.
[0494] Image information processing
[0495] When a user takes a photo with their device, the image information is sent to a server via a cloud service. The server automatically adjusts the brightness, contrast, and color using image processing algorithms. The edited image is then sent back to the device and immediately saved to the gallery.
[0496] Examples of specific cases and prompt statements
[0497] For example, if a user asks "What's on my schedule tomorrow?", the system will respond with calendar information via voice. By entering a prompt such as "Please analyze the user's statement to determine the next appointment and provide a voice response," the system can generate an appropriate response.
[0498] This invention is expected to enable users to perform everyday tasks more efficiently, thereby improving their quality of life.
[0499] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0500] Step 1:
[0501] The user gives voice commands to the device. The device uses a microphone to convert the voice information into a digital format and acquires the digital signal as data. The input is a voice signal, and the output is digital voice data. This digital voice data is sent to the server via the internet. For example, the user might say, "Tell me my next appointment."
[0502] Step 2:
[0503] The server converts the received digital audio data into text using speech recognition technology. Next, it uses a natural language processing algorithm to analyze the user's intent from the text. The input is digital audio data, and the output is the analyzed text. Specifically, the server extracts the text "Tell me my next appointment" and recognizes that this indicates the user's intention to check their calendar.
[0504] Step 3:
[0505] The server generates an appropriate response by applying a response generation algorithm based on the analyzed text information. The generated response is converted into speech data using a speech synthesis engine. The input is the analyzed text information, and the output is the generated speech data. Specifically, the server generates the response "Tomorrow's schedule is a meeting at 10 AM" and converts it into speech data.
[0506] Step 4:
[0507] The server sends a synthesized speech response to the terminal. The terminal plays the received audio data through its speaker and outputs it in a format audible to the user. The input is the generated audio data, and the output is the playback of the audio. Specifically, the terminal delivers the audio message "Tomorrow's schedule includes a meeting at 10:00 AM" to the user.
[0508] Step 5:
[0509] The terminal records the user's application program usage history and sends this history to the server. The server analyzes the usage history using a machine learning algorithm and determines which application program should be recommended next. The input is the usage history data, and the output is the recommended application. Specifically, the server generates an instruction such as, "Since you use the weather app every morning, we recommend launching it tomorrow morning."
[0510] Step 6:
[0511] The user sends images taken with their device to the server via the cloud. The server applies image processing algorithms to automatically edit the images. The input is the captured image, and the output is the processed image data. Specifically, the server performs a process such as "brightening a photo of a cloudy sky and improving its quality."
[0512] Step 7:
[0513] The server sends the processed image back to the terminal. The terminal saves the received image to its gallery and notifies the user so they can confirm it. The input is the processed image data, and the output is the image saved in the gallery. Specifically, the terminal saves the edited image and notifies the user that "image editing is complete."
[0514] (Application Example 1)
[0515] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0516] Modern consumers, when purchasing goods in physical stores, demand quick access to product information and personalized product recommendations based on their individual preferences. However, conventional systems have struggled to efficiently achieve this, resulting in lower customer satisfaction. This invention aims to solve these problems and provide a richer shopping experience.
[0517] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0518] In this invention, the server includes a natural language processing device that acquires voice information and converts it into text information, an analysis device that analyzes the user's purpose from the text information and generates a response, and a suggestion generation device that records the user's behavior history and makes suggestions. As a result, when users search for products in a store, they can obtain product information by voice operation, and furthermore, products are recommended individually based on their past purchase history, enabling an efficient and personalized shopping experience.
[0519] "Audio information" refers to information acquired as audio data, specifically the content of the user's speech.
[0520] "Textual information" refers to text data obtained by converting audio information using natural language processing.
[0521] A "natural language processing device" is a device that analyzes spoken information and converts it into written information.
[0522] An "analysis device" is a device that reads the user's purpose and intention from textual information and generates an appropriate response.
[0523] A "suggestion generation device" is a device that suggests the next appropriate application or product based on the user's behavioral history.
[0524] "Image information" refers to visual data, including still images and moving images, that is acquired or processed in digital format.
[0525] An "image editing device" is a device that receives image information and automatically adjusts brightness, color tone, and other parameters.
[0526] A "location identification device" is a device that provides location information of products within a physical store, assisting users in accessing those products.
[0527] A "recommendation device" is a device that recommends suitable products or services based on the user's purchase history and preferences.
[0528] A "communication device" is a device used to send and receive data between a server and a device, and it exchanges information via a network such as the internet.
[0529] This invention is an information processing system for improving the customer purchasing experience in physical stores, and has the following specific configuration and operation.
[0530] First, the server acquires voice information and converts it into text using a natural language processing unit. This process utilizes the Google Cloud Speech-to-Text API, which uses speech recognition technology to transcribe customer voice commands into text. Next, the transcribed text is processed by an analyzer to identify the user's purpose and request. The analyzer uses the Google Cloud Natural Language API, which understands the context and generates an appropriate response.
[0531] The generated response is presented to the user via the device in the form of audio or text. This response includes information and guidance about products in the store.
[0532] Furthermore, the suggestion generation device utilizes the user's behavioral history to recommend products suitable for the user. Additionally, the location identification device searches for products within the store and provides customers with visual or audio guidance. This incorporates location services using Bluetooth and Wi-Fi to accurately determine the placement of products within the store.
[0533] The image editing device automatically adjusts the brightness and color tone when the user takes a photo of a product. This process utilizes OpenCV, a Python library, to enhance the appeal of the photo.
[0534] In a physical store setting, a customer might ask their smart glasses, "What's on sale next?" The system analyzes the question and displays a list of sale items on the glasses' screen. Furthermore, if the system analyzes that the user frequently purchases chocolate, it might recommend a newly released chocolate product as the next featured item.
[0535] An example of a prompt to input into the generative AI model would be: "Assume a user is looking for a specific product in a store, then simulate a conversation that identifies the location of this product and presents relevant information based on the user's past interests in that product." This prompt allows the system to dynamically provide the user with specific and relevant information.
[0536] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0537] Step 1:
[0538] When a user asks a product question by voice into the device, the device acquires that voice information through its microphone. This voice information is stored on the device as audio data in a digital format.
[0539] Step 2:
[0540] The device converts the acquired audio information into text information. During this process, it uses speech recognition technology to send data to the Google Cloud Speech-to-Text API, which analyzes the audio data and generates text data. The output is the text information corresponding to the original audio information.
[0541] Step 3:
[0542] The server sends the received text information to a parser for analysis to understand the user's intent. Specifically, it uses the Google Cloud Natural Language API to grasp the context and meaning and generate an appropriate response to the user's question. The output is response data based on the analyzed information.
[0543] Step 4:
[0544] The server sends the generated response data to the terminal. The terminal receives the response data and presents it to the user as text or audio. Information based on the user's questions is provided visually or audibly.
[0545] Step 5:
[0546] Simultaneously, a suggestion generator operates using the user's behavioral history data. The server analyzes the user's past behavior and, based on that behavior, analyzes products the user is likely to purchase next. This data processing is performed using machine learning algorithms. The output is product information recommended to the user.
[0547] Step 6:
[0548] When a user takes a picture of a product using their device's camera, the image information is saved on the device and sent to the image editing system on the server. Image editing processing is performed using OpenCV, and the brightness and contrast of the image are automatically adjusted. The output is the edited image data.
[0549] Step 7:
[0550] The edited image data is sent back to the device. Users can review the edited image and, if necessary, post reviews or share it on social media. This entire process improves the user's purchasing experience.
[0551] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0552] This invention is a system that combines a natural language processing function for handling user voice data with an emotion recognition function for recognizing user emotions. This system understands the user's emotional state and selects and provides appropriate responses and content accordingly, thereby realizing a more personalized service.
[0553] Voice input and emotion recognition
[0554] The device captures the user's voice and sends the audio data to the server. The server converts the received audio data into text data using a natural language processing engine and analyzes its content. At the same time, it uses emotion recognition to analyze supplementary information such as intonation, speed, and intensity of the voice and identify the user's emotions. For example, if the user says, "I was very busy and tired today," the emotion engine will read emotions such as fatigue and stress from this voice.
[0555] Generating responses and proposals
[0556] The server generates a response based on the analyzed text and sentiment data. If the text analysis suggests schedule management and the sentiment analysis indicates user fatigue, the server converts the response into speech, such as, "Let's make sure you get enough rest today. Shall I remind you of your schedule?"
[0557] Furthermore, it records user emotional data and suggests appropriate applications and content based on that data. For example, it can suggest relaxing music content to users who are experiencing high levels of fatigue.
[0558] Image editing function
[0559] Images taken by the user with their device are sent to a server, which automatically edits them using image editing tools. The editing process may include adding appropriate filters and effects, taking emotional data into consideration. For example, if a user edits an image that appears depressed, a brightening of the tone may be applied.
[0560] This system aims to improve the user experience by instantly providing appropriate responses that meet the user's needs and emotions.
[0561] The following describes the processing flow.
[0562] Step 1:
[0563] The user speaks into their smartphone, saying, "Today was a long day." The device collects the voice data and sends it to a server.
[0564] Step 2:
[0565] The server converts the received audio data into text data using a natural language processing engine. During this process, it extracts audio features and passes them to an emotion recognition engine.
[0566] Step 3:
[0567] The server's emotion recognition engine analyzes the intonation and speed of the speech and determines that the user is experiencing fatigue. This result is then sent as feedback to the natural language processing engine.
[0568] Step 4:
[0569] The server analyzes text data to understand the user's intent. Simultaneously, it determines an appropriate response to the user based on sentiment data.
[0570] Step 5:
[0571] The server generates the response, "Why not listen to some music to relax today?", and then converts it into audio data using a text-to-speech module.
[0572] Step 6:
[0573] The server sends the converted audio data to the terminal. The terminal plays this audio for the user and presents the suggested content.
[0574] Step 7:
[0575] When a user receives a suggestion for relaxing music and selects a piece of content, the device sends that information to the server and opens the appropriate application. The server then suggests additional content as needed.
[0576] (Example 2)
[0577] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0578] Modern information processing systems are required to quickly and accurately analyze user intent based on voice data and provide personalized responses that understand emotions. However, conventional systems have insufficient accuracy in voice recognition and emotion recognition, and lack the ability to utilize various data to improve the user experience. Furthermore, it has been difficult to perform content suggestions and image editing processing based on user emotions in real time. Against this backdrop, the present invention aims to realize the provision of comprehensive services based on aggregated user behavior history and emotion data.
[0579] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0580] In this invention, the server includes language processing means for acquiring voice information and converting it into text information, information analysis means for analyzing the user's intent from the text information and generating an appropriate response, and emotion recognition means for analyzing the intonation and speed of the voice to identify emotions. This makes it possible to accurately analyze the user's intent and emotions from their voice and provide responses, services, and product suggestions based on that analysis in real time.
[0581] "Voice information" refers to a data format in which the content spoken by a user is recorded and processed as a digital signal.
[0582] "Textual information" refers to string data converted from audio information, which is further analyzed by language processing tools.
[0583] "Language processing means" refers to a system or method for converting speech information into text information and analyzing its content.
[0584] "Information analysis means" refers to technologies that analyze the user's intent and context based on textual information and generate the optimal response.
[0585] "Emotion recognition methods" refer to methods that analyze non-verbal elements such as intonation, speed, and intensity of speech to identify the user's emotional state.
[0586] "Conversion means" refers to a system that converts the response generated based on the analysis results back into an audio format and delivers it to the user.
[0587] "Suggestion generation means" refers to technology that suggests appropriate services and content to users based on recorded user behavior history and emotional data.
[0588] "Image processing means" refers to a system that automatically edits received image information and, in some cases, applies filters or effects that reflect emotional data.
[0589] A description of embodiments for carrying out the present invention will be provided.
[0590] The system of this invention has the function of performing natural language processing using the user's voice information and recognizing their emotional state. The system mainly consists of a terminal used by the user and a server that processes its data.
[0591] First, when the user speaks into the device, the device captures the audio information. The device has a communication function that uses a microphone to capture high-quality audio and transmits the information to a remote server. Specific examples include smartphones and tablets.
[0592] When the server receives voice information from a terminal, it converts it into text using speech recognition technology. Google's voice services or open-source speech recognition engines may be used for this purpose. The converted text information is then analyzed by a natural language processing engine to identify the user's intent. Next, an emotion recognition engine analyzes the intonation and speed of the voice to reveal the user's emotions. Tools specifically designed for emotion analysis are used in these processes.
[0593] Next, the server generates a response to the user based on the analysis data obtained. The generated response is converted back into speech format using speech synthesis and fed back to the user from the terminal. The speech synthesis technology used here includes commercial or open-source speech synthesis engines. For example, if the user says, "I'm very busy and tired today," the server can generate a response such as, "Please take a good rest. Shall I remind you of your appointment?"
[0594] Additionally, the system includes a feature that accumulates user behavior history and emotional data, and suggests content based on this data. When a user sends an image they have taken to the server, the server automatically processes the image using its image editing function. This function utilizes image editing software, applying filters according to the user's emotional state.
[0595] Because the responses and suggestions generated by this system are based on user data, it can provide a personalized experience.
[0596] For example, one could use the following prompt on a generative AI model: "Analyze the user's intent and emotions from their voice input and provide rest-related suggestions based on that." This would allow the system to provide the most appropriate response within the user's context.
[0597] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0598] Step 1:
[0599] The user provides voice input to the device. The device captures this voice using its microphone and saves it as digital audio data. The input is analog audio data, and the output is digital audio data. The device improves the sound quality of this data through noise reduction and sampling.
[0600] Step 2:
[0601] The terminal sends the captured digital audio data to the server. The server receives this data and converts it into text using a speech recognition engine. The input is digital audio data, and the output is text. In this conversion process, the server uses a speech recognition API to convert audio to text.
[0602] Step 3:
[0603] The server passes textual information to a natural language processing engine for analysis. This analysis identifies the user's intent and requirements. The input is textual information, and the output is analyzed intent data. The server uses a text analysis algorithm to extract intent from grammar and keywords.
[0604] Step 4:
[0605] Simultaneously, the server uses an emotion recognition engine that analyzes emotions from the intonation and speed of the audio data. The input is audio data, and the output is emotion data. The server applies signal processing technology to identify the tone and tempo of the voice as an emotion pattern.
[0606] Step 5:
[0607] The server generates an appropriate response using a generative AI model based on the analyzed intent and sentiment data. The input is intent and sentiment data, and the output is the response text. The generative AI model takes the prompt sentence as a variable and calculates the response.
[0608] Step 6:
[0609] The generated response text is converted into speech through a speech synthesis engine and sent back to the terminal. The input is the response text, and the output is synthesized speech data. The server uses speech synthesis technology to generate natural-sounding speech.
[0610] Step 7:
[0611] Ultimately, the terminal outputs the audio data received from the server to the user. The input is synthesized speech data, and the output is an audio response to the user. The terminal plays the audio message using its speaker.
[0612] (Application Example 2)
[0613] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0614] Traditional content delivery systems provide uniform content without considering user emotions, resulting in a limited user experience and a lack of individually personalized services. Furthermore, they lack responses and content tailored to the user's emotional state, highlighting the need for higher-quality user satisfaction.
[0615] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0616] In this invention, the server includes natural language processing means for acquiring voice data and converting it into text data; analysis means for analyzing the user's intent and emotions from the text data and voice data and generating appropriate responses and content; and content recommendation means for presenting content appropriate to the user based on the analyzed emotions. This makes it possible to provide personalized responses and content that correspond to the user's emotional state.
[0617] "Natural language processing means" refers to devices or software that convert speech data into text data, and is a function used to analyze the content of a user's speech.
[0618] "Analysis means" refers to devices or programs that analyze the user's intent and emotions from text data and audio data, and generate appropriate responses and content based on this analysis.
[0619] "Content recommendation methods" refer to functions and system configurations that present users with the most suitable content based on analyzed emotions.
[0620] "Suggestion generation means" refers to a function that suggests appropriate applications and services based on the user's behavioral history and emotional data.
[0621] "Image editing means" refers to devices or software that automatically edit image data received from users and apply filters and effects that correspond to emotions.
[0622] "Communication means" refers to devices and technologies that have the function of transmitting voice data, text data, and sentiment data to a server via a communication device.
[0623] "Behavioral analysis means" refers to technologies and devices for aggregating user behavior history and emotional data over a certain period and analyzing the results of that aggregation.
[0624] The system that implements this application works by directly using the user's voice data and converting it into text data using natural language processing technology. The server analyzes the voice input using dedicated natural language processing software and extracts the user's intent and emotions from the text data. For this process, libraries such as "speech_recognition" and emotion analysis engines such as "DeepAffects" can be used as cloud services.
[0625] Furthermore, the server suggests the most suitable content to the user based on the results of emotion recognition. For example, if emotion analysis indicates that the user is tired, relaxing music or videos will be recommended preferentially. To achieve this, a content recommendation engine works to select content that matches the user's needs. The "content_recommendation" module efficiently handles this.
[0626] A specific scenario would be that when a user voice-inputs "I'm very tired today," the system analyzes this to recognize the user's fatigue level and selects relaxing videos or music. A possible prompt message would be something like, "The user may need rest. Please suggest video content related to relaxation and well-being."
[0627] The device collects and transmits voice data, and then transfers the data to the server via communication means. Through this process, data analysis and content provision are carried out quickly, thereby improving the user experience.
[0628] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0629] Step 1:
[0630] The device acquires audio data from the user. This audio data is collected as an audio waveform, capturing the words the user speaks directly. For example, the device uses a microphone to capture ambient sound and converts that data into a digital signal.
[0631] Step 2:
[0632] The terminal transmits the acquired audio data to the server. Here, data is transmitted in real time over the network using communication methods, preparing the server for audio analysis. The input is audio data in binary format, and the output is data transfer over the network.
[0633] Step 3:
[0634] The server converts the received audio data into text data using a natural language processing engine. This process uses the "speech_recognition" library to perform speech recognition and convert the audio waveform into corresponding sentences. The input is audio data, and the output is text data.
[0635] Step 4:
[0636] The server performs emotion recognition based on text and audio data. This process analyzes specific keywords, speech intonation, and speed, and uses an emotion recognition engine to identify the user's emotions. The output is data indicating the user's emotions.
[0637] Step 5:
[0638] The server uses the analysis results to generate appropriate responses and content. It utilizes a generative AI model to construct responses tailored to the user's intent and a content recommendation engine to select relevant videos and music. The output consists of a response message and recommended content.
[0639] Step 6:
[0640] The server formats the generated response as audio data and sends it to the terminal. At this stage, text-to-speech conversion occurs and is returned to the terminal via the network. The output is response data in audio format.
[0641] Step 7:
[0642] The terminal plays the audio data received from the server and provides it to the user. Here, it uses its speaker to output an audio response, allowing the user to learn about the suggested content. The process concludes when audio playback is complete.
[0643] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0644] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0645] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0646] [Fourth Embodiment]
[0647] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0648] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0649] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0650] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0651] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0652] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0653] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0654] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0655] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0656] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0657] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0658] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0659] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0660] This invention provides a system incorporating various functions to improve user convenience. The system primarily performs natural language processing for handling voice data, application suggestions based on user behavior history, and automatic image data editing. The detailed operation of each function is described below.
[0661] Voice control using natural language processing
[0662] When a user speaks into their smartphone using natural language, the device converts the audio into digital data and sends it to a server. The server uses advanced natural language processing technology to convert the audio into text data and analyze its content. For example, if a user asks, "What time is the next meeting?", the server refers to calendar information and derives the answer. Then, it converts the analysis result back into audio and outputs it from the device in a format that is easy for the user to understand.
[0663] Automated application suggestions
[0664] The device records the user's application usage history and periodically sends it to the server. Based on this history, the server analyzes the user's behavior patterns and predicts the next most suitable application. For example, if a user has a habit of opening the weather app at 7 AM every day, the server instructs the device to display the weather app on the home screen in the morning. In this way, the system helps save the user time.
[0665] Automated image editing
[0666] Images taken by the user are sent to the server via the device. The server utilizes image editing algorithms to adjust the brightness, contrast, and color of the images, automating high-quality editing. For example, a photo of a cloudy beach taken during a trip can be transformed into a bright and vibrant image. The edited image is then sent back from the server to the device and immediately saved to the gallery.
[0667] This system aims to improve the quality of life by enabling users to complete daily tasks more efficiently.
[0668] The following describes the processing flow.
[0669] Step 1:
[0670] The user speaks into their smartphone, saying, "Tell me my next appointment." The device acquires the voice data through the microphone and temporarily stores it locally.
[0671] Step 2:
[0672] The terminal sends the stored audio data to the server using a communication module. The server then passes the received audio data to a natural language processing engine.
[0673] Step 3:
[0674] The server uses a natural language processing engine to convert the audio data into text data. This text data then serves as input data for analyzing the user's intent.
[0675] Step 4:
[0676] The server parses the converted text data and retrieves the user's next appointment by referencing the calendar database. After generating an appropriate response, it formats the response back into text format.
[0677] Step 5:
[0678] The server sends the generated text response to a text-to-speech module, which converts it back into audio data.
[0679] Step 6:
[0680] The server transmits converted audio data to the terminal. The terminal plays the received data to the user through its speaker. This allows the user to hear their next appointment.
[0681] (Example 1)
[0682] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0683] Modern information systems demand efficient task management and information access from users. However, conventional systems have struggled to centrally and automatically perform diverse tasks such as voice control, application program suggestions, and image processing. This has resulted in a lack of convenience for users and a failure to improve their quality of life.
[0684] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0685] In this invention, the server includes information processing means for acquiring voice information and converting it into text information, recommendation generation means for recording the user's usage history and suggesting appropriate application programs to the user based on that history, and image processing means for receiving image information and processing it automatically. This makes it possible for the user to efficiently manage a variety of tasks and improve convenience.
[0686] "Audio information" refers to the data format obtained from audio signals, and represents acoustic information.
[0687] "Information processing means" refers to a method or device for acquiring audio information and converting it into text information.
[0688] "Textual information" refers to data expressed as text, representing the result of conversion from audio information.
[0689] "Analysis means" refers to a method or device for identifying the user's intent from textual information and generating an appropriate response.
[0690] "Response" refers to response information generated by the analysis tool that corresponds to the user's intent.
[0691] "Audio output means" refers to a method or device for reproducing the response generated by the analysis means as audio.
[0692] "User usage history" refers to a record of operations and activities performed by the user within the system.
[0693] "Recommendation generation means" refers to a method or apparatus for suggesting an appropriate application program based on the user's usage history.
[0694] "Image information" refers to visual data and includes image data expressed in digital format.
[0695] "Image processing means" refers to a method or device for receiving image information and automatically editing or adjusting it.
[0696] "Communication device" refers to a network connection device used to transmit voice and text information to a server.
[0697] A "remote device" refers to an external computing resource used to process data received from a user device and transmit the results.
[0698] This invention is a system that enables voice information processing, automatic suggestion of application programs, and image information processing, with the aim of improving user convenience. The system consists mainly of a terminal and a server, each performing a specific function and operating through mutual communication.
[0699] Speech information processing
[0700] When a user speaks into the device, the device converts the voice information into a digital format. Commonly available voice recognition software can be used for speech recognition. The digital voice information is transmitted to a server via a communication device. The server uses advanced natural language processing software to convert the voice information into text and analyze the user's intent. The analyzed information is then converted into an appropriate response by a response generation algorithm and transmitted back to the device as voice information via a speech synthesizer.
[0701] Automated application suggestions
[0702] The terminal continuously records the user's application program usage history. The recorded history data is sent to a server, where machine learning algorithms analyze the user's behavior patterns. Based on this analysis, the server recommends suitable application programs for the user. The recommended application programs are displayed on the user's screen via the terminal.
[0703] Image information processing
[0704] When a user takes a photo with their device, the image information is sent to a server via a cloud service. The server automatically adjusts the brightness, contrast, and color using image processing algorithms. The edited image is then sent back to the device and immediately saved to the gallery.
[0705] Examples of specific cases and prompt statements
[0706] For example, if a user asks "What's on my schedule tomorrow?", the system will respond with calendar information via voice. By entering a prompt such as "Please analyze the user's statement to determine the next appointment and provide a voice response," the system can generate an appropriate response.
[0707] This invention is expected to enable users to perform everyday tasks more efficiently, thereby improving their quality of life.
[0708] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0709] Step 1:
[0710] The user gives voice commands to the device. The device uses a microphone to convert the voice information into a digital format and acquires the digital signal as data. The input is a voice signal, and the output is digital voice data. This digital voice data is sent to the server via the internet. For example, the user might say, "Tell me my next appointment."
[0711] Step 2:
[0712] The server converts the received digital audio data into text using speech recognition technology. Next, it uses a natural language processing algorithm to analyze the user's intent from the text. The input is digital audio data, and the output is the analyzed text. Specifically, the server extracts the text "Tell me my next appointment" and recognizes that this indicates the user's intention to check their calendar.
[0713] Step 3:
[0714] The server generates an appropriate response by applying a response generation algorithm based on the analyzed text information. The generated response is converted into speech data using a speech synthesis engine. The input is the analyzed text information, and the output is the generated speech data. Specifically, the server generates the response "Tomorrow's schedule is a meeting at 10 AM" and converts it into speech data.
[0715] Step 4:
[0716] The server sends a synthesized speech response to the terminal. The terminal plays the received audio data through its speaker and outputs it in a format audible to the user. The input is the generated audio data, and the output is the playback of the audio. Specifically, the terminal delivers the audio message "Tomorrow's schedule includes a meeting at 10:00 AM" to the user.
[0717] Step 5:
[0718] The terminal records the user's application program usage history and sends this history to the server. The server analyzes the usage history using a machine learning algorithm and determines which application program should be recommended next. The input is the usage history data, and the output is the recommended application. Specifically, the server generates an instruction such as, "Since you use the weather app every morning, we recommend launching it tomorrow morning."
[0719] Step 6:
[0720] The user sends images taken with their device to the server via the cloud. The server applies image processing algorithms to automatically edit the images. The input is the captured image, and the output is the processed image data. Specifically, the server performs a process such as "brightening a photo of a cloudy sky and improving its quality."
[0721] Step 7:
[0722] The server sends the processed image back to the terminal. The terminal saves the received image to its gallery and notifies the user so they can confirm it. The input is the processed image data, and the output is the image saved in the gallery. Specifically, the terminal saves the edited image and notifies the user that "image editing is complete."
[0723] (Application Example 1)
[0724] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0725] Modern consumers, when purchasing goods in physical stores, demand quick access to product information and personalized product recommendations based on their individual preferences. However, conventional systems have struggled to efficiently achieve this, resulting in lower customer satisfaction. This invention aims to solve these problems and provide a richer shopping experience.
[0726] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0727] In this invention, the server includes a natural language processing device that acquires voice information and converts it into text information, an analysis device that analyzes the user's purpose from the text information and generates a response, and a suggestion generation device that records the user's behavior history and makes suggestions. As a result, when users search for products in a store, they can obtain product information by voice operation, and furthermore, products are recommended individually based on their past purchase history, enabling an efficient and personalized shopping experience.
[0728] "Audio information" refers to information acquired as audio data, specifically the content of the user's speech.
[0729] "Textual information" refers to text data obtained by converting audio information using natural language processing.
[0730] A "natural language processing device" is a device that analyzes spoken information and converts it into written information.
[0731] An "analysis device" is a device that reads the user's purpose and intention from textual information and generates an appropriate response.
[0732] A "suggestion generation device" is a device that suggests the next appropriate application or product based on the user's behavioral history.
[0733] "Image information" refers to visual data, including still images and moving images, that is acquired or processed in digital format.
[0734] An "image editing device" is a device that receives image information and automatically adjusts brightness, color tone, and other parameters.
[0735] A "location identification device" is a device that provides location information of products within a physical store, assisting users in accessing those products.
[0736] A "recommendation device" is a device that recommends suitable products or services based on the user's purchase history and preferences.
[0737] A "communication device" is a device used to send and receive data between a server and a device, and it exchanges information via a network such as the internet.
[0738] This invention is an information processing system for improving the customer purchasing experience in physical stores, and has the following specific configuration and operation.
[0739] First, the server acquires voice information and converts it into text using a natural language processing unit. This process utilizes the Google Cloud Speech-to-Text API, which uses speech recognition technology to transcribe customer voice commands into text. Next, the transcribed text is processed by an analyzer to identify the user's purpose and request. The analyzer uses the Google Cloud Natural Language API, which understands the context and generates an appropriate response.
[0740] The generated response is presented to the user via the device in the form of audio or text. This response includes information and guidance about products in the store.
[0741] Furthermore, the suggestion generation device utilizes the user's behavioral history to recommend products suitable for the user. Additionally, the location identification device searches for products within the store and provides customers with visual or audio guidance. This incorporates location services using Bluetooth and Wi-Fi to accurately determine the placement of products within the store.
[0742] The image editing device automatically adjusts the brightness and color tone when the user takes a photo of a product. This process utilizes OpenCV, a Python library, to enhance the appeal of the photo.
[0743] In a physical store setting, a customer might ask their smart glasses, "What's on sale next?" The system analyzes the question and displays a list of sale items on the glasses' screen. Furthermore, if the system analyzes that the user frequently purchases chocolate, it might recommend a newly released chocolate product as the next featured item.
[0744] An example of a prompt to input into the generative AI model would be: "Assume a user is looking for a specific product in a store, then simulate a conversation that identifies the location of this product and presents relevant information based on the user's past interests in that product." This prompt allows the system to dynamically provide the user with specific and relevant information.
[0745] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0746] Step 1:
[0747] When a user asks a product question by voice into the device, the device acquires that voice information through its microphone. This voice information is stored on the device as audio data in a digital format.
[0748] Step 2:
[0749] The device converts the acquired audio information into text information. During this process, it uses speech recognition technology to send data to the Google Cloud Speech-to-Text API, which analyzes the audio data and generates text data. The output is the text information corresponding to the original audio information.
[0750] Step 3:
[0751] The server sends the received text information to a parser for analysis to understand the user's intent. Specifically, it uses the Google Cloud Natural Language API to grasp the context and meaning and generate an appropriate response to the user's question. The output is response data based on the analyzed information.
[0752] Step 4:
[0753] The server sends the generated response data to the terminal. The terminal receives the response data and presents it to the user as text or audio. Information based on the user's questions is provided visually or audibly.
[0754] Step 5:
[0755] Simultaneously, a suggestion generator operates using the user's behavioral history data. The server analyzes the user's past behavior and, based on that behavior, analyzes products the user is likely to purchase next. This data processing is performed using machine learning algorithms. The output is product information recommended to the user.
[0756] Step 6:
[0757] When a user takes a picture of a product using their device's camera, the image information is saved on the device and sent to the image editing system on the server. Image editing processing is performed using OpenCV, and the brightness and contrast of the image are automatically adjusted. The output is the edited image data.
[0758] Step 7:
[0759] The edited image data is sent back to the device. Users can review the edited image and, if necessary, post reviews or share it on social media. This entire process improves the user's purchasing experience.
[0760] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0761] This invention is a system that combines a natural language processing function for handling user voice data with an emotion recognition function for recognizing user emotions. This system understands the user's emotional state and selects and provides appropriate responses and content accordingly, thereby realizing a more personalized service.
[0762] Voice input and emotion recognition
[0763] The device captures the user's voice and sends the audio data to the server. The server converts the received audio data into text data using a natural language processing engine and analyzes its content. At the same time, it uses emotion recognition to analyze supplementary information such as intonation, speed, and intensity of the voice and identify the user's emotions. For example, if the user says, "I was very busy and tired today," the emotion engine will read emotions such as fatigue and stress from this voice.
[0764] Generating responses and proposals
[0765] The server generates a response based on the analyzed text and sentiment data. If the text analysis suggests schedule management and the sentiment analysis indicates user fatigue, the server converts the response into speech, such as, "Let's make sure you get enough rest today. Shall I remind you of your schedule?"
[0766] Furthermore, it records user emotional data and suggests appropriate applications and content based on that data. For example, it can suggest relaxing music content to users who are experiencing high levels of fatigue.
[0767] Image editing function
[0768] Images taken by the user with their device are sent to a server, which automatically edits them using image editing tools. The editing process may include adding appropriate filters and effects, taking emotional data into consideration. For example, if a user edits an image that appears depressed, a brightening of the tone may be applied.
[0769] This system aims to improve the user experience by instantly providing appropriate responses that meet the user's needs and emotions.
[0770] The following describes the processing flow.
[0771] Step 1:
[0772] The user speaks into their smartphone, saying, "Today was a long day." The device collects the voice data and sends it to a server.
[0773] Step 2:
[0774] The server converts the received audio data into text data using a natural language processing engine. During this process, it extracts audio features and passes them to an emotion recognition engine.
[0775] Step 3:
[0776] The server's emotion recognition engine analyzes the intonation and speed of the speech and determines that the user is experiencing fatigue. This result is then sent as feedback to the natural language processing engine.
[0777] Step 4:
[0778] The server analyzes text data to understand the user's intent. Simultaneously, it determines an appropriate response to the user based on sentiment data.
[0779] Step 5:
[0780] The server generates the response, "Why not listen to some music to relax today?", and then converts it into audio data using a text-to-speech module.
[0781] Step 6:
[0782] The server sends the converted audio data to the terminal. The terminal plays this audio for the user and presents the suggested content.
[0783] Step 7:
[0784] When a user receives a suggestion for relaxing music and selects a piece of content, the device sends that information to the server and opens the appropriate application. The server then suggests additional content as needed.
[0785] (Example 2)
[0786] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0787] Modern information processing systems are required to quickly and accurately analyze user intent based on voice data and provide personalized responses that understand emotions. However, conventional systems have insufficient accuracy in voice recognition and emotion recognition, and lack the ability to utilize various data to improve the user experience. Furthermore, it has been difficult to perform content suggestions and image editing processing based on user emotions in real time. Against this backdrop, the present invention aims to realize the provision of comprehensive services based on aggregated user behavior history and emotion data.
[0788] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0789] In this invention, the server includes language processing means for acquiring voice information and converting it into text information, information analysis means for analyzing the user's intent from the text information and generating an appropriate response, and emotion recognition means for analyzing the intonation and speed of the voice to identify emotions. This makes it possible to accurately analyze the user's intent and emotions from their voice and provide responses, services, and product suggestions based on that analysis in real time.
[0790] "Voice information" refers to a data format in which the content spoken by a user is recorded and processed as a digital signal.
[0791] "Textual information" refers to string data converted from audio information, which is further analyzed by language processing tools.
[0792] "Language processing means" refers to a system or method for converting speech information into text information and analyzing its content.
[0793] "Information analysis means" refers to technologies that analyze the user's intent and context based on textual information and generate the optimal response.
[0794] "Emotion recognition methods" refer to methods that analyze non-verbal elements such as intonation, speed, and intensity of speech to identify the user's emotional state.
[0795] "Conversion means" refers to a system that converts the response generated based on the analysis results back into an audio format and delivers it to the user.
[0796] "Suggestion generation means" refers to technology that suggests appropriate services and content to users based on recorded user behavior history and emotional data.
[0797] "Image processing means" refers to a system that automatically edits received image information and, in some cases, applies filters or effects that reflect emotional data.
[0798] A description of embodiments for carrying out the present invention will be provided.
[0799] The system of this invention has the function of performing natural language processing using the user's voice information and recognizing their emotional state. The system mainly consists of a terminal used by the user and a server that processes its data.
[0800] First, when the user speaks into the device, the device captures the audio information. The device has a communication function that uses a microphone to capture high-quality audio and transmits the information to a remote server. Specific examples include smartphones and tablets.
[0801] When the server receives voice information from a terminal, it converts it into text using speech recognition technology. Google's voice services or open-source speech recognition engines may be used for this purpose. The converted text information is then analyzed by a natural language processing engine to identify the user's intent. Next, an emotion recognition engine analyzes the intonation and speed of the voice to reveal the user's emotions. Tools specifically designed for emotion analysis are used in these processes.
[0802] Next, the server generates a response to the user based on the analysis data obtained. The generated response is converted back into speech format using speech synthesis and fed back to the user from the terminal. The speech synthesis technology used here includes commercial or open-source speech synthesis engines. For example, if the user says, "I'm very busy and tired today," the server can generate a response such as, "Please take a good rest. Shall I remind you of your appointment?"
[0803] Additionally, the system includes a feature that accumulates user behavior history and emotional data, and suggests content based on this data. When a user sends an image they have taken to the server, the server automatically processes the image using its image editing function. This function utilizes image editing software, applying filters according to the user's emotional state.
[0804] Because the responses and suggestions generated by this system are based on user data, it can provide a personalized experience.
[0805] For example, one could use the following prompt on a generative AI model: "Analyze the user's intent and emotions from their voice input and provide rest-related suggestions based on that." This would allow the system to provide the most appropriate response within the user's context.
[0806] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0807] Step 1:
[0808] The user provides voice input to the device. The device captures this voice using its microphone and saves it as digital audio data. The input is analog audio data, and the output is digital audio data. The device improves the sound quality of this data through noise reduction and sampling.
[0809] Step 2:
[0810] The terminal sends the captured digital audio data to the server. The server receives this data and converts it into text using a speech recognition engine. The input is digital audio data, and the output is text. In this conversion process, the server uses a speech recognition API to convert audio to text.
[0811] Step 3:
[0812] The server passes textual information to a natural language processing engine for analysis. This analysis identifies the user's intent and requirements. The input is textual information, and the output is analyzed intent data. The server uses a text analysis algorithm to extract intent from grammar and keywords.
[0813] Step 4:
[0814] Simultaneously, the server uses an emotion recognition engine that analyzes emotions from the intonation and speed of the audio data. The input is audio data, and the output is emotion data. The server applies signal processing technology to identify the tone and tempo of the voice as an emotion pattern.
[0815] Step 5:
[0816] The server generates an appropriate response using a generative AI model based on the analyzed intent and sentiment data. The input is intent and sentiment data, and the output is the response text. The generative AI model takes the prompt sentence as a variable and calculates the response.
[0817] Step 6:
[0818] The generated response text is converted into speech through a speech synthesis engine and sent back to the terminal. The input is the response text, and the output is synthesized speech data. The server uses speech synthesis technology to generate natural-sounding speech.
[0819] Step 7:
[0820] Ultimately, the terminal outputs the audio data received from the server to the user. The input is synthesized speech data, and the output is an audio response to the user. The terminal plays the audio message using its speaker.
[0821] (Application Example 2)
[0822] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0823] Traditional content delivery systems provide uniform content without considering user emotions, resulting in a limited user experience and a lack of individually personalized services. Furthermore, they lack responses and content tailored to the user's emotional state, highlighting the need for higher-quality user satisfaction.
[0824] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0825] In this invention, the server includes natural language processing means for acquiring voice data and converting it into text data; analysis means for analyzing the user's intent and emotions from the text data and voice data and generating appropriate responses and content; and content recommendation means for presenting content appropriate to the user based on the analyzed emotions. This makes it possible to provide personalized responses and content that correspond to the user's emotional state.
[0826] "Natural language processing means" refers to devices or software that convert speech data into text data, and is a function used to analyze the content of a user's speech.
[0827] "Analysis means" refers to devices or programs that analyze the user's intent and emotions from text data and audio data, and generate appropriate responses and content based on this analysis.
[0828] "Content recommendation methods" refer to functions and system configurations that present users with the most suitable content based on analyzed emotions.
[0829] "Suggestion generation means" refers to a function that suggests appropriate applications and services based on the user's behavioral history and emotional data.
[0830] "Image editing means" refers to devices or software that automatically edit image data received from users and apply filters and effects that correspond to emotions.
[0831] "Communication means" refers to devices and technologies that have the function of transmitting voice data, text data, and sentiment data to a server via a communication device.
[0832] "Behavioral analysis means" refers to technologies and devices for aggregating user behavior history and emotional data over a certain period and analyzing the results of that aggregation.
[0833] The system that implements this application works by directly using the user's voice data and converting it into text data using natural language processing technology. The server analyzes the voice input using dedicated natural language processing software and extracts the user's intent and emotions from the text data. For this process, libraries such as "speech_recognition" and emotion analysis engines such as "DeepAffects" can be used as cloud services.
[0834] Furthermore, the server suggests the most suitable content to the user based on the results of emotion recognition. For example, if emotion analysis indicates that the user is tired, relaxing music or videos will be recommended preferentially. To achieve this, a content recommendation engine works to select content that matches the user's needs. The "content_recommendation" module efficiently handles this.
[0835] A specific scenario would be that when a user voice-inputs "I'm very tired today," the system analyzes this to recognize the user's fatigue level and selects relaxing videos or music. A possible prompt message would be something like, "The user may need rest. Please suggest video content related to relaxation and well-being."
[0836] The device collects and transmits voice data, and then transfers the data to the server via communication means. Through this process, data analysis and content provision are carried out quickly, thereby improving the user experience.
[0837] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0838] Step 1:
[0839] The device acquires audio data from the user. This audio data is collected as an audio waveform, capturing the words the user speaks directly. For example, the device uses a microphone to capture ambient sound and converts that data into a digital signal.
[0840] Step 2:
[0841] The terminal transmits the acquired audio data to the server. Here, data is transmitted in real time over the network using communication methods, preparing the server for audio analysis. The input is audio data in binary format, and the output is data transfer over the network.
[0842] Step 3:
[0843] The server converts the received audio data into text data using a natural language processing engine. This process uses the "speech_recognition" library to perform speech recognition and convert the audio waveform into corresponding sentences. The input is audio data, and the output is text data.
[0844] Step 4:
[0845] The server performs emotion recognition based on text and audio data. This process analyzes specific keywords, speech intonation, and speed, and uses an emotion recognition engine to identify the user's emotions. The output is data indicating the user's emotions.
[0846] Step 5:
[0847] The server uses the analysis results to generate appropriate responses and content. It utilizes a generative AI model to construct responses tailored to the user's intent and a content recommendation engine to select relevant videos and music. The output consists of a response message and recommended content.
[0848] Step 6:
[0849] The server formats the generated response as audio data and sends it to the terminal. At this stage, text-to-speech conversion occurs and is returned to the terminal via the network. The output is response data in audio format.
[0850] Step 7:
[0851] The terminal plays the audio data received from the server and provides it to the user. Here, it uses its speaker to output an audio response, allowing the user to learn about the suggested content. The process concludes when audio playback is complete.
[0852] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0853] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0854] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0855] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0856] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0857] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0858] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0859] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0860] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0861] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0862] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0863] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0864] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0865] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0866] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0867] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0868] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0869] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0870] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0871] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0872] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0873] The following is further disclosed regarding the embodiments described above.
[0874] (Claim 1)
[0875] A natural language processing means that acquires audio data and converts it into text data,
[0876] An analytical means that analyzes user intent from text data and generates an appropriate response,
[0877] A means for outputting the response generated by the analysis means as audio,
[0878] A suggestion generation means that records the user's behavior history and suggests appropriate applications to the user based on that history,
[0879] An image editing method that receives image data and automatically performs editing,
[0880] A system that includes this.
[0881] (Claim 2)
[0882] The system according to claim 1, comprising communication means for transmitting voice data and text data to a server via a communication device.
[0883] (Claim 3)
[0884] The system according to claim 1, comprising a behavioral analysis means that aggregates a user's behavioral history over a certain period and analyzes the results using an analysis means.
[0885] "Example 1"
[0886] (Claim 1)
[0887] Information processing means that acquires audio information and converts it into text information,
[0888] An analytical means that analyzes the user's intent from textual information and generates an appropriate response,
[0889] A means for outputting the response generated by the analysis means as audio,
[0890] A recommendation generation means that records the user's usage history and proposes an appropriate application program to the user based on that history,
[0891] An image processing method that receives image information and automatically processes it,
[0892] A device that includes this.
[0893] (Claim 2)
[0894] The apparatus according to claim 1, a communication means for transmitting voice information and text information to a remote device via a communication device.
[0895] (Claim 3)
[0896] The apparatus according to claim 1, which collects a user's usage history over a certain period and analyzes the results using an analysis means.
[0897] "Application Example 1"
[0898] (Claim 1)
[0899] A natural language processing device that acquires audio information and converts it into text information,
[0900] An analysis device that analyzes the user's purpose from textual information and generates an appropriate response,
[0901] A device that outputs the response generated by the analysis device as audio,
[0902] A suggestion generation device that records the user's behavior history and suggests appropriate applications to the user based on that history,
[0903] An image editing device that receives image information and automatically performs editing,
[0904] A location identification device for providing product information in the store environment,
[0905] A recommendation device that recommends the next suitable product based on purchase history,
[0906] A system that includes this.
[0907] (Claim 2)
[0908] A communication device for transmitting voice information and text information to a computer via a communication device, according to claim 1.
[0909] (Claim 3)
[0910] The behavioral analysis device according to claim 1, which collects the user's behavioral history over a certain period and analyzes the results using an analysis device.
[0911] "Example 2 of combining an emotion engine"
[0912] (Claim 1)
[0913] A language processing means that acquires audio information and converts it into text information,
[0914] An information analysis means that analyzes the user's intent from textual information and generates an appropriate response,
[0915] An emotion recognition method that identifies emotions by analyzing the intonation and speed of speech,
[0916] A conversion means for outputting the response generated by the analysis means and emotion recognition means as audio,
[0917] A proposal generation means that records the user's behavioral history and emotional data, and proposes appropriate services and content based on that data,
[0918] An image processing means that receives image information and automatically edits it based on emotion data,
[0919] A system that includes this.
[0920] (Claim 2)
[0921] The system according to claim 1, a network communication means that transmits voice information and text information to a server via a communication device for processing.
[0922] (Claim 3)
[0923] The system according to claim 1, which includes a behavioral analysis means that aggregates user behavior history and emotional data over a certain period and analyzes the results using an information analysis means.
[0924] "Application example 2 when combining with an emotional engine"
[0925] (Claim 1)
[0926] A natural language processing means that acquires audio data and converts it into text data,
[0927] An analytical means that analyzes the user's intent and emotions from text data and audio data, and generates appropriate responses and content.
[0928] A means for outputting the response generated by the analysis means as audio,
[0929] A content recommendation system that presents content appropriate to the user based on analyzed emotions,
[0930] A suggestion generation means that records the user's behavior history and emotional data and suggests applications based on that data,
[0931] An image editing method that receives image data and automatically edits it by applying filters according to emotions,
[0932] A system that includes this.
[0933] (Claim 2)
[0934] A communication means for transmitting voice data, text data, and emotion data to a server via a communication device, according to claim 1.
[0935] (Claim 3)
[0936] The system according to claim 1, comprising a behavioral analysis means that aggregates user behavior history and emotional data over a certain period and analyzes the results using an analysis means. [Explanation of Symbols]
[0937] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A natural language processing means that acquires audio data and converts it into text data, An analytical means that analyzes user intent from text data and generates an appropriate response, A means for outputting the response generated by the analysis means as audio, A suggestion generation means that records the user's behavior history and suggests appropriate applications to the user based on that history, An image editing method that receives image data and automatically performs editing, A system that includes this.
2. The system according to claim 1, further comprising communication means for transmitting voice data and text data to a server via a communication device.
3. The system according to claim 1, further comprising behavioral analysis means for aggregating user behavior history over a certain period and analyzing the results using analysis means.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A