system
A system using a multimodal AI model for analyzing user data automates UX improvement, addressing inefficiencies in conventional methods by providing rapid and objective interface design evaluations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-17
- Publication Date
- 2026-04-30
AI Technical Summary
Conventional methods for improving user experience (UX) in websites and applications require significant effort and subjective user testing, making rapid and efficient improvements difficult due to variability in user proficiency and observer judgment.
A system that utilizes a multimodal generation AI model to automatically analyze user operation data, video, and audio data to generate and evaluate different user interfaces, enabling efficient and objective UX improvements.
Enables rapid and efficient UX improvements by automatically generating and evaluating user interface designs based on comprehensive user data analysis, leading to increased user satisfaction.
Smart Images

Figure 2026071583000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] For providers of websites and applications, improving the user experience (UX) is important for enhancing competitiveness. However, conventional UX improvement methods have problems that a great deal of effort is required to conduct user usability tests and identify improvement points, making rapid and efficient improvement difficult. Furthermore, since the results vary depending on the proficiency of test users and observers, it is also difficult to make objective judgments.
Means for Solving the Problems
[0005] This invention provides a system that rapidly generates suggestions for improving the user experience by automatically analyzing user operation data and video and audio data using a multimodal generation AI model. It centrally analyzes data acquired from terminals and automatically generates different user interfaces based on multiple improvement suggestions. Furthermore, it includes a means to propose the optimal UX design by automating evaluation tests using these interfaces and analyzing the results. This system enables more efficient and effective improvement of the user experience.
[0006] "User interaction data" refers to information about user interactions such as clicks, scrolls, and taps on websites and applications.
[0007] "Video and audio data" refers to video data and audio recording data used to record the user's visual and auditory feedback.
[0008] A "multimodal generation AI model" is an artificial intelligence technology that simultaneously analyzes data in different formats, such as text, audio, and video, and combines them to generate information.
[0009] "User experience improvement proposals" are suggestions regarding the interface and functionality of websites and applications, aimed at improving usability and user engagement.
[0010] "Different user interfaces" refer to multiple versions of a user interface that have been newly designed and generated based on improvement proposals.
[0011] "Evaluation testing" is a testing process that involves collecting user evaluations and results from users who use the generated user interface. [Brief explanation of the drawing]
[0012] [Figure 1]This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0013] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] This invention provides a system for improving user experience (UX). This system efficiently collects user operation data, video, and audio data, and automatically generates improvement suggestions by analyzing them using a multimodal generation AI model. A specific embodiment of this system is described below.
[0034] First, the device collects data in real time as the user interacts with websites and applications. This includes basic actions such as clicking, scrolling, and tapping, as well as the user's facial expressions and voice feedback.
[0035] Next, the server receives the collected data and performs preprocessing such as noise reduction and data format conversion. The preprocessed data is then sent to a multimodal generative AI model, where it is analyzed using deep learning techniques. In this process, patterns of user behavior and areas for improvement are discovered, and specific UX improvement proposals are automatically generated.
[0036] The generated improvement proposals are automatically designed as different user interfaces for each server. This creates various versions of the interface design, preparing them for evaluation testing.
[0037] Next, the terminal conducts evaluation tests with users using the previously generated interface design. User operation data and feedback for each interface are collected again, and the server analyzes the results in detail. The optimal interface design is identified and compiled into a report and proposal.
[0038] Users participate in the system's improvement cycle by experiencing an interface based on the proposed optimal UX and providing feedback. By repeating this process, the UX continuously improves, ultimately leading to increased user satisfaction.
[0039] As a concrete example, if users on an e-commerce site find it difficult to find products, this system analyzes user browsing data and voice feedback to generate suggestions for improving the product search function. A new search interface is automatically generated, evaluated by several users, and the most effective interface is adopted based on the results. This entire process is largely automated, enabling efficient UX improvement.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The device collects user operation data (clicks, scrolls, taps) in real time. Furthermore, it uses the device's camera and microphone to capture the user's facial expressions and voice feedback.
[0043] Step 2:
[0044] The server receives operation, video, and audio data transmitted from the terminal. It then performs preprocessing to remove noise and convert the data format to one suitable for analysis.
[0045] Step 3:
[0046] The server inputs pre-processed data into a multimodal generation AI model. This AI model identifies UX issues faced by the user by analyzing user behavior patterns and feedback from voice.
[0047] Step 4:
[0048] The server automatically generates specific UX improvement proposals using AI based on identified problems. In this process, it utilizes insights gained from user behavior patterns.
[0049] Step 5:
[0050] The server automatically constructs different user interface designs based on the generated improvement suggestions. These designs include new design elements and improvements.
[0051] Step 6:
[0052] The terminal provides each generated interface design to the target users and conducts evaluation tests. During this process, user operation data and feedback are collected again.
[0053] Step 7:
[0054] The server analyzes the results of evaluation tests and assesses user reactions and effectiveness for each interface version. This helps identify the optimal interface design.
[0055] Step 8:
[0056] Users participate in the continuous improvement process by experiencing the optimized interface and providing further feedback. This feedback will be used in the next improvement cycle.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] Traditional methods for improving user experience lacked mechanisms to fully utilize user behavior data, and analyzing feedback based on users' intuition and emotions was difficult. As a result, effective user interface design and improvement were not achieved to the extent expected.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] In this invention, the server includes means for collecting information on user operations using an information processing device, means for acquiring image and sound data, and means for analyzing the acquired data using a multimodal generation artificial intelligence model and automatically generating improvement proposals regarding the user experience. This makes it possible to automatically generate highly accurate interface improvement proposals that comprehensively utilize a wide range of user feedback.
[0062] An "information processing device" is a device used to acquire, analyze, and manage information related to user operations.
[0063] "Information related to operation" refers to operational data that a user uses when interacting with a digital interface, including data such as clicks, scrolls, and taps.
[0064] "Image and audio data" refers to data that records the user's visual and auditory feedback, and is collected using cameras and microphones.
[0065] A "multimodal generation artificial intelligence model" is an artificial intelligence technology used to integrate and analyze multiple data modalities to extract insights.
[0066] "User interface" refers to the design elements of screens and operating systems that allow users to interact with a system.
[0067] An "evaluation test" is a test conducted to verify how effective the generated user interface is for the user.
[0068] This invention relates to an information system for improving user experience (UX). This system efficiently collects information, images, and audio data related to user operations and automatically generates improvement suggestions by analyzing them using a multi-modal generation artificial intelligence model.
[0069] The device collects information in real time as the user interacts with websites and applications. This information includes basic actions such as clicks, scrolls, and taps, as well as facial expression and audio data collected using the camera and microphone. Specifically, this objective is achieved by utilizing UI tracking tools and speech recognition software.
[0070] The server receives data sent from the terminal. This data is first preprocessed by denoising and converting the data format. The preprocessed data is then input into a multimodal generation artificial intelligence model. The AI model analyzes this data to identify user behavior patterns and areas for UX improvement, and automatically generates improvement suggestions. Machine learning frameworks such as TENSORFLOW® and PyTorch are used to build the AI model.
[0071] Users experience the optimized interface based on the proposed improvements and provide feedback. This feedback will be used in the next cycle to contribute to further improvements in the user experience (UX).
[0072] For example, when addressing the difficulties users experience when searching for products on an e-commerce site, this system uses user interaction data and voice feedback to generate suggestions for improving the product search function. New search interfaces are automatically generated, evaluated by a small number of users, and the optimal interface is selected, thus promoting automatic and efficient system improvement.
[0073] An example of a prompt message might be: "Analyze the difficulties users encounter when searching for a specific product category and suggest a more intuitive search interface."
[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0075] Step 1:
[0076] The device collects data as users interact with websites and applications. As input, user actions such as clicks, scrolls, and taps are captured in real time. This data is initially converted into a digital format using UI tracking tools (e.g., event logging). The output is a dataset recording user actions.
[0077] Step 2:
[0078] The device collects facial expression data and acoustic feedback. Inputs include video and audio data acquired using a camera and microphone. This data is then processed by image recognition and speech recognition software to extract changes in the user's facial expressions and voice tone. The output is attribute data representing the results of the facial and audio analysis.
[0079] Step 3:
[0080] The server receives all collected data and performs preprocessing. The input consists of operation data, facial expression data, and voice data sent from the terminal. The server removes noise using filtering techniques and standardizes the data format (e.g., data cleaning, normalization). The output is a clean dataset suitable for data analysis.
[0081] Step 4:
[0082] The server inputs pre-processed data into a multi-modal generation AI model. This input consists of cleaned operation data, facial expressions, and voice data. The server feeds this data into the AI model and analyzes the relationships between the data. Using deep learning techniques, it discovers patterns in user behavior and areas for UX improvement. The output is a concrete UX improvement proposal based on the analysis results.
[0083] Step 5:
[0084] The server automatically designs different user interfaces based on improvement suggestions generated by the AI model. The input is the specification of the improvement suggestion. The server uses design software to generate multiple versions of the interface (e.g., automatic generation of UI design). The output is a test-ready interface design.
[0085] Step 6:
[0086] The device undergoes evaluation testing with users using the designed interface. The input is the generated interface design. User interaction data and feedback are collected again (e.g., interaction log recording, surveys). The output is the evaluation data obtained from the test.
[0087] Step 7:
[0088] The server analyzes evaluation data in detail to identify the optimal interface design. The input is the results of evaluation tests. The server uses statistical analysis methods and machine learning models to evaluate the data and select the most effective interface design. The output is a report of the optimized interface, compiled as a proposal.
[0089] (Application Example 1)
[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0091] In modern e-commerce, improving the user experience is a crucial element for gaining a competitive edge. However, it is technically challenging to grasp user interaction data and emotional responses in real time and instantly optimize the interface based on that data. Therefore, more advanced data analysis technologies and responsive systems are required to rapidly and continuously improve user satisfaction.
[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0093] In this invention, the server includes means for collecting user operation information in real time, means for acquiring and analyzing video and audio information using a smart device, and means for automatically generating improvement suggestions using multimodal generation AI. This enables rapid evaluation and suggestion of the user experience based on the user's emotional and operational responses.
[0094] "User interaction information" refers to data related to user interactions such as clicks, scrolls, and taps on websites and applications.
[0095] "Video information" refers to digital image and video data acquired through a camera, capturing the user's facial expressions, posture, and other details during operation.
[0096] "Voice information" refers to digital audio data related to the voice and voice feedback that a user makes while operating the device.
[0097] A "multimodal generation AI" is an artificial intelligence model that combines and analyzes multiple different data formats and modalities (e.g., video and audio) to gain insights.
[0098] An "improvement suggestion" is a proposal for changes or optimizations to the user interface that is automatically generated by the system with the aim of improving the user experience.
[0099] "User information display" refers to the visual presentation methods and interaction designs for information that users can see on the screen.
[0100] An "evaluation experiment" is a testing process that involves analyzing feedback obtained from actual users who test the system's generated user information display.
[0101] A "smart device" refers to a mobile device, such as a mobile terminal or wearable device, that has internet connectivity and advanced computing capabilities.
[0102] "Real-time" refers to a processing method that instantly acquires user behavior and feedback, and quickly analyzes and responds to it.
[0103] To realize this invention, a system is needed that collects user operation information and emotional feedback in real time, analyzes them, and generates appropriate improvement suggestions. This system primarily involves three parties: the server, the terminal, and the user.
[0104] First, let's explain the role of the terminal. The terminal is designed to collect user operation information in real time, and simultaneously acquires video and audio information through the smart device. In this process, the terminal efficiently stores data using OpenCV for image processing and libraries for audio processing. For example, the camera on the terminal captures the user's facial expressions, and the microphone acquires audio feedback.
[0105] Next, the server receives this acquired data and performs preprocessing and analysis. Here, noise reduction and data formatting are performed, and the data is analyzed using a multimodal generative AI utilizing deep learning frameworks such as TensorFlow or PyTorch. The generative AI model fuses multiple data sources, automatically identifies UX patterns that can be improved, and quickly generates improvement suggestions. As a concrete example, prompts such as "Please suggest a more user-friendly layout by making the hierarchy of this category shallower" can be generated, enabling the suggestion of a new UI.
[0106] Finally, let's discuss the user's role. Users experience the proposed new user interface and provide feedback. This feedback is then sent back to the server, and the AI goes through an iterative optimization process to provide the optimal UX. Through this iterative process, the UX is continuously improved, and it is expected that a high level of user satisfaction will ultimately be achieved.
[0107] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0108] Step 1:
[0109] The device collects user action information in real time. Input consists of user actions such as clicks, scrolls, and taps. The device acquires this information as a data stream and stores it in a database as initial processing. Output is the action data formatted for analysis by the server.
[0110] Step 2:
[0111] The device acquires video and audio information using a smart device. The input consists of the user's facial expressions and voice captured through the device's camera and microphone. Features are extracted from the image data using OpenCV, and the audio data is acoustically preprocessed to convert it into a format suitable for analysis. The output consists of processed video and audio information for analysis.
[0112] Step 3:
[0113] The server receives processed operation data, video information, and audio information, and performs noise reduction and data format conversion. The input consists of data in multiple formats sent from the terminal, and the server integrates these to build a consistent, analyzable dataset. The output is integrated data in a state suitable for analysis.
[0114] Step 4:
[0115] The server feeds the integrated data into a multimodal generative AI model. The input is the data integrated in the previous step. The server applies the generative AI model using TensorFlow or PyTorch, analyzing user behavior and emotional responses through deep learning. The output is the discovery of UX patterns and the automatic generation of improvement suggestions.
[0116] Step 5:
[0117] The server automatically generates a new user interface based on the proposed improvements. The input is the improvement suggestions obtained in the previous step, and the server prompts the AI to create various UX design options. The output is the new user interface to be tested.
[0118] Step 6:
[0119] The user experiences the new user interface and provides feedback to the server. Inputs are interactions with the generated interface and user feedback. Outputs are evaluation data returned to the server, which the system uses for further optimization.
[0120] Step 7:
[0121] The server analyzes user feedback and adjusts the generated AI model as needed to provide the optimal UX. The input is evaluation feedback collected from users. Based on this data, the server performs pattern analysis to further improve the overall user experience of the system. The output is an optimized UX suggestion.
[0122] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0123] The present invention aims to enhance user experience (UX) through a system that incorporates an emotion engine. This system can generate more accurate UX improvement proposals by considering not only user behavior data but also emotions. A specific embodiment of this system is described below.
[0124] First, the device acquires user operation data, video, and audio data in real time. This data includes basic information about the user's interactions, as well as data related to emotions, such as facial expressions and speech.
[0125] Next, the server receives the collected data and performs preprocessing such as noise reduction and data transformation. This data is then fed into the emotion engine, which recognizes emotional states (e.g., joy, anxiety, frustration). The recognized emotion data, along with other operational data, is input into a multimodal generative AI model.
[0126] This AI model performs a detailed analysis that takes user emotions into account, identifying obstacles and areas for improvement in the user experience. Based on this, a process automatically generates more targeted UX improvement suggestions. These suggestions include elements that improve the quality of the user's perceived experience itself.
[0127] The generated improvement suggestions are automatically designed as different user interfaces for each server. These interfaces incorporate elements designed to improve user emotional responses. Based on this, a new interface design is generated, and evaluation tests are conducted using it.
[0128] The terminal provides users with an interface design specifically created for evaluation testing, and collects user interaction data, sentiment data, and feedback. The server then analyzes this data in detail to identify the optimal interface design.
[0129] As a concrete example, in an educational application, if a user expresses frustration with a particular learning content, this system identifies that emotion and generates design improvement suggestions to reduce the stress experienced with that content. New content layouts and supplementary learning support are designed, and an interface that considers the user's emotions is provided, leading to improved learning efficiency. This entire process is automated, enabling efficient UX improvement.
[0130] The following describes the processing flow.
[0131] Step 1:
[0132] The device begins collecting data about the user's interactions. Specifically, in addition to operation data such as clicks, scrolls, and taps, it also uses the camera and microphone to capture facial expressions and voice data.
[0133] Step 2:
[0134] The server receives operation data, video data, and audio data transmitted from the terminal. After receiving the data, it cleanses it, removes noise, and converts it into a format suitable for analysis.
[0135] Step 3:
[0136] The server inputs the pre-processed data into the emotion engine. The emotion engine recognizes the user's emotions from facial and voice data and generates emotion data such as joy, anxiety, and frustration.
[0137] Step 4:
[0138] The server inputs emotional and operational data into a multimodal generation AI model for analysis. This automatically generates specific UX improvement suggestions to enhance the user experience. Because emotional data is included, the improvement suggestions take into account the user's emotional responses.
[0139] Step 5:
[0140] The server automatically generates different user interface designs based on the proposed improvements. At this stage, emotionally sensitive elements are incorporated into the design.
[0141] Step 6:
[0142] The device presents the created user interface to the user as an evaluation test. It then collects user operation data and changes in emotions.
[0143] Step 7:
[0144] The server analyzes data obtained from evaluation tests to assess user reactions and the effectiveness of each user interface. This identifies the most emotionally resonant and optimal interface.
[0145] Step 8:
[0146] Users experience the optimized user interface and provide feedback again. Based on this feedback, a new improvement cycle is initiated.
[0147] (Example 2)
[0148] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0149] Traditional user experience improvement systems focus solely on user interaction data and do not take emotional information into account. This has resulted in a failure to adequately reflect the dissatisfaction and stress that users may be experiencing. Furthermore, the generated improvement suggestions are not always optimal in response to user emotions, making it difficult to effectively improve user experience satisfaction.
[0150] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0151] In this invention, the server includes means for collecting user behavior information and emotional information in real time, means for pre-processing the collected data by removing noise and performing data transformation, and means for recognizing the emotional state using an emotion engine. This makes it possible to generate suggestions for improving the user experience that take emotions into account and to design an optimal user interface based on these suggestions.
[0152] "User action information" refers to a series of operation records and interaction data generated when a user operates the system.
[0153] "Emotional information" refers to data that indicates the emotional state obtained through the user's facial expressions, tone of voice, and other nonverbal cues.
[0154] "Noise reduction" refers to the process of removing unnecessary information and errors before data analysis, preparing the data for analysis.
[0155] "Data transformation preprocessing" refers to a series of processes for converting collected data into a format that can be analyzed.
[0156] An "emotion engine" refers to a specific algorithm or model used to recognize a user's emotional state.
[0157] A "multimodal generation AI model" refers to an artificial intelligence model that combines and analyzes multiple different types of data to generate useful output.
[0158] "User experience improvement proposals" refer to suggested changes and optimizations based on user interactions, aimed at providing a better user experience.
[0159] A "user interface" refers to a set of visual or manipulative elements that enable a system and a user to interact directly.
[0160] "Evaluation testing" refers to a series of tests and evaluation procedures conducted to confirm the effectiveness of a new user interface or function.
[0161] This system was developed with the aim of improving the user experience, and it involves the coordinated functioning of multiple hardware and software components.
[0162] The device plays a role in collecting user operation data, video, and audio data in real time. This collection uses the camera and microphone built into the device. This data records how the user is using the system and reflects their emotional state at the time. For example, it captures the facial expressions and tone of voice a user may have while using an online educational app.
[0163] The server receives data sent from the terminal and first performs noise reduction and data transformation. This makes the data suitable for analysis and removes unnecessary information. Next, this data is analyzed by the emotion engine to identify the user's emotional state. Based on this emotional data, it is integrated with the operation data and input into the multimodal generation AI model.
[0164] This AI model has the ability to analyze recognized emotions and interaction information and generate suggestions for improving the user experience. These suggestions include proposals to enhance the quality of the experience based on the user's emotions. For example, in an educational app, if a user experiences stress or confusion during learning activities, appropriate support features or interface changes will be suggested.
[0165] The generated improvement proposals are then automatically regenerated by the server as a new user interface. Here, elements that improve user emotional responses are incorporated into the interface. Furthermore, evaluation tests are conducted using the new interface design to confirm its effectiveness.
[0166] As a concrete example, a prompt message might be input into the generative AI model in the form of, "Based on the emotions the user expressed regarding a specific issue, suggest which parts should be improved and how." This aims to improve the quality of the user experience.
[0167] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0168] Step 1:
[0169] The device collects user operation data, video, and audio data in real time. Input here is video and audio from the camera and microphone connected to the device. From this, information related to emotions, such as the specific content of operations, facial expressions, and tone of voice, is obtained. For example, if a user is watching a video on a smartphone app, their facial expressions and the tempo of their operations are recorded. This data is then sent to a server as output.
[0170] Step 2:
[0171] The server receives data sent from the terminal and performs preprocessing. The input is the raw data sent by the terminal. Specifically, it performs noise reduction and converts audio and video data into a format suitable for emotion recognition. For example, it removes noise from audio data and smooths facial expression data frame by frame. As output, the preprocessed data is passed to the emotion engine.
[0172] Step 3:
[0173] The server uses pre-processed data to perform analysis with its emotion engine. The input is data that has been de-noised and transformed. The emotion engine analyzes this data to identify emotional states such as "joy," "anxiety," and "frustration." For example, it estimates what emotions the user is experiencing by analyzing subtle facial muscle movements and vocal intonation patterns. The analysis results are output as emotion data, and the process proceeds to the next step.
[0174] Step 4:
[0175] The server inputs emotional and operational data into a multimodal generation AI model. This input consists of emotional data obtained in the previous step and initial operational data. This AI model comprehensively analyzes various data to identify areas for improvement in the user experience. For example, it can be used to identify the cause of discomfort a user experiences with a particular interface. The analysis results are output as suggestions for improving the generated user experience.
[0176] Step 5:
[0177] The server automatically generates a new user interface design based on the generated improvement suggestions. The input is the improvement suggestions output by the AI model. In this step, the user interface is designed with emotionally sensitive elements in mind, for example, generating a design that enhances guidance and help in stressful situations. The output is a new user interface, which then proceeds to evaluation testing.
[0178] Step 6:
[0179] The terminal conducts evaluation tests on users using the generated user interface. The input is the newly designed interface design. Users interact with this interface and provide further data through their experience. Specifically, feedback is obtained regarding the usability and effectiveness of the interface design. The output is the operation data and feedback, which are used for further analysis on the server.
[0180] Step 7:
[0181] The server analyzes data obtained from evaluation tests to identify the optimal user interface. The input is the data obtained as a result of the evaluation tests. This data is analyzed in detail to confirm whether the initial improvement proposals were actually effective, and feedback for redesign is provided as needed. For example, in a learning app, user stress metrics are monitored to evaluate interface improvements. Ultimately, the optimal design is determined.
[0182] (Application Example 2)
[0183] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0184] Current interface designs often fail to consider the user's emotions, making it difficult to provide an optimal user interface based on the emotional state of individual users. As a result, stress and frustration accumulate among users when using a site, leading to a decline in the overall quality of the experience.
[0185] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0186] In this invention, the server includes means for collecting user behavior data, means for acquiring visual and auditory data, and means for analyzing the acquired data and recognizing the user's emotional state, equipped with an emotion engine. This makes it possible to provide an individually optimized user interface based on the user's emotional state and improve the quality of the user experience.
[0187] "User behavior data" refers to information about input operations and actions that occur when a user interacts with the system.
[0188] "Visual and auditory data" refers to data related to the user's facial expressions and voice, acquired through cameras and microphones.
[0189] The term "emotion engine" refers to a function that analyzes the user's emotional state based on acquired data and recognizes emotions such as joy, anger, sadness, and happiness.
[0190] A "multimodal generation AI model" refers to an artificial intelligence model that integrates and analyzes multiple different types of data, generating suggestions for improving the user experience while considering user behavior and emotions.
[0191] "User interface" refers to the design elements and layouts that allow users to interact with a computer system.
[0192] An "evaluation test" refers to a test conducted to check the effectiveness of the generated user interface and to consider the need for improvement.
[0193] "Text data" refers to information obtained by converting audio or other data formats into written form.
[0194] The system for implementing this invention mainly consists of a user terminal and a server. The terminal collects user operation data, visual data, and auditory data in real time. Specifically, the terminal's camera captures facial expressions and its microphone captures voice, and this data includes basic information about the user's interactions. This data also serves as foundational data for indicating the user's emotional state.
[0195] The server receives data transmitted from the terminal and performs preprocessing such as noise reduction. The emotion engine then analyzes this data to recognize the user's emotional state. This information is fed into a multimodal generation AI model for detailed analysis regarding improvements to the user experience. This AI model generates suggestions for user interfaces that take the emotional state into account.
[0196] Furthermore, the generated interface proposals are presented to the user as different user interfaces. The device then collects operational and emotional data from the user as they use the new interface, and obtains feedback. This feedback data is analyzed on the server to identify the optimal interface design, ultimately leading to improved user experience (UX).
[0197] For example, if a user on an e-commerce site shows an anxious expression on the product purchase page, the system can detect that emotion and adjust the interface to emphasize review and question features to reassure the user, or even suggest a new interface.
[0198] Examples of prompts for a generative AI model:
[0199] "If a user is showing signs of anxiety, how can we incorporate elements into the layout to provide reassurance?"
[0200] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0201] Step 1:
[0202] The device collects user operation data, visual data, and auditory data in real time. This includes facial expression data captured by the smartphone's camera and audio data recorded by the microphone. This collected data is sent to a server. The input is the user's operations and actions, and the output is the collected raw data.
[0203] Step 2:
[0204] The server receives raw data sent from the terminal and performs preprocessing such as noise reduction and data normalization. This data processing generates clean data that makes it easier for the emotion engine to perform accurate sentiment analysis.
[0205] Step 3:
[0206] The server inputs pre-processed data into the emotion engine and analyzes the user's emotional state. For example, major emotions such as joy and anxiety are identified, and these recognition results are obtained. The input data is pre-processed data, and the output is the identified emotional state.
[0207] Step 4:
[0208] The server integrates various types of information, including recognized emotion data, into a multimodal generation AI model. This AI model generates UX improvement suggestions that take each user's emotions into account. The input is various types of data, including emotion data, and the output is the generated UX improvement suggestions.
[0209] Step 5:
[0210] The server automatically generates a new user interface optimized for the user based on the generated UX improvement suggestions. This designs an interface that enhances the user experience. The input is the UX improvement suggestions, and the output is the new interface design.
[0211] Step 6:
[0212] The terminal provides the user with a new user interface and collects data and emotional responses from the user as they interact with it. This interaction data and feedback data are sent to the server as part of an evaluation test. The input is the user's interaction and feedback, and the output is the evaluation test results.
[0213] Step 7:
[0214] The server analyzes the evaluation test results and identifies the optimal interface design. Based on this analysis, if further UX improvements are needed, the design is refined. The input is the evaluation test results, and the output is the final interface design.
[0215] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0216] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0217] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0218] [Second Embodiment]
[0219] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0220] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0221] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0222] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0223] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0224] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0225] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0226] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0227] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0228] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0229] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0230] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0231] This invention provides a system for improving user experience (UX). This system efficiently collects user operation data, video, and audio data, and automatically generates improvement suggestions by analyzing them using a multimodal generation AI model. A specific embodiment of this system is described below.
[0232] First, the device collects data in real time as the user interacts with websites and applications. This includes basic actions such as clicking, scrolling, and tapping, as well as the user's facial expressions and voice feedback.
[0233] Next, the server receives the collected data and performs preprocessing such as noise reduction and data format conversion. The preprocessed data is then sent to a multimodal generative AI model, where it is analyzed using deep learning techniques. In this process, patterns of user behavior and areas for improvement are discovered, and specific UX improvement proposals are automatically generated.
[0234] The generated improvement proposals are automatically designed as different user interfaces for each server. This creates various versions of the interface design, preparing them for evaluation testing.
[0235] Next, the terminal conducts evaluation tests with users using the previously generated interface design. User operation data and feedback for each interface are collected again, and the server analyzes the results in detail. The optimal interface design is identified and compiled into a report and proposal.
[0236] Users participate in the system's improvement cycle by experiencing an interface based on the proposed optimal UX and providing feedback. By repeating this process, the UX continuously improves, ultimately leading to increased user satisfaction.
[0237] As a concrete example, if users on an e-commerce site find it difficult to find products, this system analyzes user browsing data and voice feedback to generate suggestions for improving the product search function. A new search interface is automatically generated, evaluated by several users, and the most effective interface is adopted based on the results. This entire process is largely automated, enabling efficient UX improvement.
[0238] The following describes the processing flow.
[0239] Step 1:
[0240] The device collects user operation data (clicks, scrolls, taps) in real time. Furthermore, it uses the device's camera and microphone to capture the user's facial expressions and voice feedback.
[0241] Step 2:
[0242] The server receives operation, video, and audio data transmitted from the terminal. It then performs preprocessing to remove noise and convert the data format to one suitable for analysis.
[0243] Step 3:
[0244] The server inputs pre-processed data into a multimodal generation AI model. This AI model identifies UX issues faced by the user by analyzing user behavior patterns and feedback from voice.
[0245] Step 4:
[0246] The server automatically generates specific UX improvement proposals using AI based on identified problems. In this process, it utilizes insights gained from user behavior patterns.
[0247] Step 5:
[0248] The server automatically constructs different user interface designs based on the generated improvement suggestions. These designs include new design elements and improvements.
[0249] Step 6:
[0250] The terminal provides each generated interface design to the target users and conducts evaluation tests. During this process, user operation data and feedback are collected again.
[0251] Step 7:
[0252] The server analyzes the results of evaluation tests and assesses user reactions and effectiveness for each interface version. This helps identify the optimal interface design.
[0253] Step 8:
[0254] Users participate in the continuous improvement process by experiencing the optimized interface and providing further feedback. This feedback will be used in the next improvement cycle.
[0255] (Example 1)
[0256] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0257] Traditional methods for improving user experience lacked mechanisms to fully utilize user behavior data, and analyzing feedback based on users' intuition and emotions was difficult. As a result, effective user interface design and improvement were not achieved to the extent expected.
[0258] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0259] In this invention, the server includes means for collecting information on user operations using an information processing device, means for acquiring image and sound data, and means for analyzing the acquired data using a multimodal generation artificial intelligence model and automatically generating improvement proposals regarding the user experience. This makes it possible to automatically generate highly accurate interface improvement proposals that comprehensively utilize a wide range of user feedback.
[0260] An "information processing device" is a device used to acquire, analyze, and manage information related to user operations.
[0261] "Information related to operation" refers to operational data that a user uses when interacting with a digital interface, including data such as clicks, scrolls, and taps.
[0262] "Image and audio data" refers to data that records the user's visual and auditory feedback, and is collected using cameras and microphones.
[0263] A "multimodal generation artificial intelligence model" is an artificial intelligence technology used to integrate and analyze multiple data modalities to extract insights.
[0264] "User interface" refers to the design elements of screens and operating systems that allow users to interact with a system.
[0265] An "evaluation test" is a test conducted to verify how effective the generated user interface is for the user.
[0266] This invention relates to an information system for improving user experience (UX). This system efficiently collects information, images, and audio data related to user operations and automatically generates improvement suggestions by analyzing them using a multi-modal generation artificial intelligence model.
[0267] The device collects information in real time as the user interacts with websites and applications. This information includes basic actions such as clicks, scrolls, and taps, as well as facial expression and audio data collected using the camera and microphone. Specifically, this objective is achieved by utilizing UI tracking tools and speech recognition software.
[0268] The server receives data sent from the terminal. This data is first preprocessed by denoising and converting the data format. The preprocessed data is then input into a multimodal generation artificial intelligence model. The AI model analyzes this data to identify user behavior patterns and areas for UX improvement, and automatically generates improvement suggestions. Machine learning frameworks such as TensorFlow and PyTorch are used to build the AI model.
[0269] Users experience the optimized interface based on the proposed improvements and provide feedback. This feedback will be used in the next cycle to contribute to further improvements in the user experience (UX).
[0270] For example, when addressing the difficulties users experience when searching for products on an e-commerce site, this system uses user interaction data and voice feedback to generate suggestions for improving the product search function. New search interfaces are automatically generated, evaluated by a small number of users, and the optimal interface is selected, thus promoting automatic and efficient system improvement.
[0271] An example of a prompt message might be: "Analyze the difficulties users encounter when searching for a specific product category and suggest a more intuitive search interface."
[0272] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0273] Step 1:
[0274] The device collects data as users interact with websites and applications. As input, user actions such as clicks, scrolls, and taps are captured in real time. This data is initially converted into a digital format using UI tracking tools (e.g., event logging). The output is a dataset recording user actions.
[0275] Step 2:
[0276] The device collects facial expression data and acoustic feedback. Inputs include video and audio data acquired using a camera and microphone. This data is then processed by image recognition and speech recognition software to extract changes in the user's facial expressions and voice tone. The output is attribute data representing the results of the facial and audio analysis.
[0277] Step 3:
[0278] The server receives all collected data and performs preprocessing. The input consists of operation data, facial expression data, and voice data sent from the terminal. The server removes noise using filtering techniques and standardizes the data format (e.g., data cleaning, normalization). The output is a clean dataset suitable for data analysis.
[0279] Step 4:
[0280] The server inputs pre-processed data into a multi-modal generation AI model. This input consists of cleaned operation data, facial expressions, and voice data. The server feeds this data into the AI model and analyzes the relationships between the data. Using deep learning techniques, it discovers patterns in user behavior and areas for UX improvement. The output is a concrete UX improvement proposal based on the analysis results.
[0281] Step 5:
[0282] The server automatically designs different user interfaces based on improvement suggestions generated by the AI model. The input is the specification of the improvement suggestion. The server uses design software to generate multiple versions of the interface (e.g., automatic generation of UI design). The output is a test-ready interface design.
[0283] Step 6:
[0284] The terminal conducts an evaluation test for users using a designed interface. The input is the generated interface design. The operation data and feedback when the user uses it are collected again (e.g., recording interaction logs, conducting questionnaire surveys). The output is the evaluation data obtained from the test.
[0285] Step 7:
[0286] The server analyzes the evaluation data in detail and identifies the optimal interface design. The input is the result data of the evaluation test. The server evaluates the data using statistical analysis methods and machine learning models, and selects the most effective interface design. The output is a report on the optimized interface summarized as a proposal.
[0287] (Application Example 1)
[0288] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0289] In modern e-commerce, improving the user experience is an important factor for gaining competitiveness. However, it is technically difficult to grasp the user's operation information and emotional reactions in real time and optimize the interface immediately based on them. Therefore, in order to quickly and continuously improve the user satisfaction, more advanced data analysis technologies and responsive systems are required.
[0290] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0291] In this invention, the server includes means for collecting the user's operation information in real time, means for acquiring and analyzing video information and audio information using a smart device, and means for automatically generating improvement plans by utilizing a multimodal generation AI. As a result, it becomes possible to quickly evaluate and propose a user experience based on the user's emotional and operational reactions.
[0292] "User interaction information" refers to data related to user interactions such as clicks, scrolls, and taps on websites and applications.
[0293] "Video information" refers to digital image and video data acquired through a camera, capturing the user's facial expressions, posture, and other details during operation.
[0294] "Voice information" refers to digital audio data related to the voice and voice feedback that a user makes while operating the device.
[0295] A "multimodal generation AI" is an artificial intelligence model that combines and analyzes multiple different data formats and modalities (e.g., video and audio) to gain insights.
[0296] An "improvement suggestion" is a proposal for changes or optimizations to the user interface that is automatically generated by the system with the aim of improving the user experience.
[0297] "User information display" refers to the visual presentation methods and interaction designs for information that users can see on the screen.
[0298] An "evaluation experiment" is a testing process that involves analyzing feedback obtained from actual users who test the system's generated user information display.
[0299] A "smart device" refers to a mobile device, such as a mobile terminal or wearable device, that has internet connectivity and advanced computing capabilities.
[0300] "Real-time" refers to a processing method that instantly acquires user behavior and feedback, and quickly analyzes and responds to it.
[0301] To realize this invention, a system is needed that collects user operation information and emotional feedback in real time, analyzes them, and generates appropriate improvement suggestions. This system primarily involves three parties: the server, the terminal, and the user.
[0302] First, let's explain the role of the terminal. The terminal is designed to collect user operation information in real time, and simultaneously acquires video and audio information through the smart device. In this process, the terminal efficiently stores data using OpenCV for image processing and libraries for audio processing. For example, the camera on the terminal captures the user's facial expressions, and the microphone acquires audio feedback.
[0303] Next, the server receives this acquired data and performs preprocessing and analysis. Here, noise reduction and data formatting are performed, and the data is analyzed using a multimodal generative AI utilizing deep learning frameworks such as TensorFlow or PyTorch. The generative AI model fuses multiple data sources, automatically identifies UX patterns that can be improved, and quickly generates improvement suggestions. As a concrete example, prompts such as "Please suggest a more user-friendly layout by making the hierarchy of this category shallower" can be generated, enabling the suggestion of a new UI.
[0304] Finally, let's discuss the user's role. Users experience the proposed new user interface and provide feedback. This feedback is then sent back to the server, and the AI goes through an iterative optimization process to provide the optimal UX. Through this iterative process, the UX is continuously improved, and it is expected that a high level of user satisfaction will ultimately be achieved.
[0305] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0306] Step 1:
[0307] The terminal collects the user's operation information in real time. The input is the operation actions of the user such as clicks, scrolls, taps, etc. The terminal obtains this information as a data stream and stores it in the database as initial processing. The output is the operation data formatted in a form required by the server for analysis.
[0308] Step 2:
[0309] The terminal uses a smart device to obtain video information and audio information. The input is the user's expressions and voices captured through the terminal's camera and microphone. By using OpenCV to extract features from the image data and preprocessing the audio data acoustically, it is converted into a form suitable for analysis. The output is the video information and audio information processed for analysis.
[0310] Step 3:
[0311] The server receives the processed operation data, video information, and audio information, and performs noise removal and data format conversion. The input is the data in multiple formats sent from the terminal, and the server integrates these to construct a consistent analyzable data set. The output is the integrated data in a state suitable for analysis.
[0312] Step 4:
[0313] The server sends the integrated data into a multimodal generation AI model. The input is the data integrated in the previous step. The server applies the generation AI model using TensorFlow or PyTorch and analyzes the user's behavior and emotional reactions through deep learning. The output is the result of automatically generating the discovery of UX patterns and improvement suggestions.
[0314] Step 5:
[0315] The server automatically generates a new user interface based on the proposed improvements. The input is the improvement suggestions obtained in the previous step, and the server prompts the AI to create various UX design options. The output is the new user interface to be tested.
[0316] Step 6:
[0317] The user experiences the new user interface and provides feedback to the server. Inputs are interactions with the generated interface and user feedback. Outputs are evaluation data returned to the server, which the system uses for further optimization.
[0318] Step 7:
[0319] The server analyzes user feedback and adjusts the generated AI model as needed to provide the optimal UX. The input is evaluation feedback collected from users. Based on this data, the server performs pattern analysis to further improve the overall user experience of the system. The output is an optimized UX suggestion.
[0320] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0321] The present invention aims to enhance user experience (UX) through a system that incorporates an emotion engine. This system can generate more accurate UX improvement proposals by considering not only user behavior data but also emotions. A specific embodiment of this system is described below.
[0322] First, the device acquires user operation data, video, and audio data in real time. This data includes basic information about the user's interactions, as well as data related to emotions, such as facial expressions and speech.
[0323] Next, the server receives the collected data and performs preprocessing such as noise reduction and data transformation. This data is then fed into the emotion engine, which recognizes emotional states (e.g., joy, anxiety, frustration). The recognized emotion data, along with other operational data, is input into a multimodal generative AI model.
[0324] This AI model performs a detailed analysis that takes user emotions into account, identifying obstacles and areas for improvement in the user experience. Based on this, a process automatically generates more targeted UX improvement suggestions. These suggestions include elements that improve the quality of the user's perceived experience itself.
[0325] The generated improvement suggestions are automatically designed as different user interfaces for each server. These interfaces incorporate elements designed to improve user emotional responses. Based on this, a new interface design is generated, and evaluation tests are conducted using it.
[0326] The terminal provides users with an interface design specifically created for evaluation testing, and collects user interaction data, sentiment data, and feedback. The server then analyzes this data in detail to identify the optimal interface design.
[0327] As a concrete example, in an educational application, if a user expresses frustration with a particular learning content, this system identifies that emotion and generates design improvement suggestions to reduce the stress experienced with that content. New content layouts and supplementary learning support are designed, and an interface that considers the user's emotions is provided, leading to improved learning efficiency. This entire process is automated, enabling efficient UX improvement.
[0328] The following describes the processing flow.
[0329] Step 1:
[0330] The device begins collecting data about the user's interactions. Specifically, in addition to operation data such as clicks, scrolls, and taps, it also uses the camera and microphone to capture facial expressions and voice data.
[0331] Step 2:
[0332] The server receives operation data, video data, and audio data transmitted from the terminal. After receiving the data, it cleanses it, removes noise, and converts it into a format suitable for analysis.
[0333] Step 3:
[0334] The server inputs the pre-processed data into the emotion engine. The emotion engine recognizes the user's emotions from facial and voice data and generates emotion data such as joy, anxiety, and frustration.
[0335] Step 4:
[0336] The server inputs emotional and operational data into a multimodal generation AI model for analysis. This automatically generates specific UX improvement suggestions to enhance the user experience. Because emotional data is included, the improvement suggestions take into account the user's emotional responses.
[0337] Step 5:
[0338] The server automatically generates different user interface designs based on the proposed improvements. At this stage, emotionally sensitive elements are incorporated into the design.
[0339] Step 6:
[0340] The device presents the created user interface to the user as an evaluation test. It then collects user operation data and changes in emotions.
[0341] Step 7:
[0342] The server analyzes data obtained from evaluation tests to assess user reactions and the effectiveness of each user interface. This identifies the most emotionally resonant and optimal interface.
[0343] Step 8:
[0344] Users experience the optimized user interface and provide feedback again. Based on this feedback, a new improvement cycle is initiated.
[0345] (Example 2)
[0346] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0347] Traditional user experience improvement systems focus solely on user interaction data and do not take emotional information into account. This has resulted in a failure to adequately reflect the dissatisfaction and stress that users may be experiencing. Furthermore, the generated improvement suggestions are not always optimal in response to user emotions, making it difficult to effectively improve user experience satisfaction.
[0348] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0349] In this invention, the server includes means for collecting user behavior information and emotional information in real time, means for pre-processing the collected data by removing noise and performing data transformation, and means for recognizing the emotional state using an emotion engine. This makes it possible to generate suggestions for improving the user experience that take emotions into account and to design an optimal user interface based on these suggestions.
[0350] "User action information" refers to a series of operation records and interaction data generated when a user operates the system.
[0351] "Emotional information" refers to data that indicates the emotional state obtained through the user's facial expressions, tone of voice, and other nonverbal cues.
[0352] "Noise reduction" refers to the process of removing unnecessary information and errors before data analysis, preparing the data for analysis.
[0353] "Data transformation preprocessing" refers to a series of processes for converting collected data into a format that can be analyzed.
[0354] An "emotion engine" refers to a specific algorithm or model used to recognize a user's emotional state.
[0355] A "multimodal generation AI model" refers to an artificial intelligence model that combines and analyzes multiple different types of data to generate useful output.
[0356] "User experience improvement proposals" refer to suggested changes and optimizations based on user interactions, aimed at providing a better user experience.
[0357] A "user interface" refers to a set of visual or manipulative elements that enable a system and a user to interact directly.
[0358] "Evaluation testing" refers to a series of tests and evaluation procedures conducted to confirm the effectiveness of a new user interface or function.
[0359] This system was developed with the aim of improving the user experience, and it involves the coordinated functioning of multiple hardware and software components.
[0360] The device plays a role in collecting user operation data, video, and audio data in real time. This collection uses the camera and microphone built into the device. This data records how the user is using the system and reflects their emotional state at the time. For example, it captures the facial expressions and tone of voice a user may have while using an online educational app.
[0361] The server receives data sent from the terminal and first performs noise reduction and data transformation. This makes the data suitable for analysis and removes unnecessary information. Next, this data is analyzed by the emotion engine to identify the user's emotional state. Based on this emotional data, it is integrated with the operation data and input into the multimodal generation AI model.
[0362] This AI model has the ability to analyze recognized emotions and interaction information and generate suggestions for improving the user experience. These suggestions include proposals to enhance the quality of the experience based on the user's emotions. For example, in an educational app, if a user experiences stress or confusion during learning activities, appropriate support features or interface changes will be suggested.
[0363] The generated improvement proposals are then automatically regenerated by the server as a new user interface. Here, elements that improve user emotional responses are incorporated into the interface. Furthermore, evaluation tests are conducted using the new interface design to confirm its effectiveness.
[0364] As a concrete example, a prompt message might be input into the generative AI model in the form of, "Based on the emotions the user expressed regarding a specific issue, suggest which parts should be improved and how." This aims to improve the quality of the user experience.
[0365] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0366] Step 1:
[0367] The device collects user operation data, video, and audio data in real time. Input here is video and audio from the camera and microphone connected to the device. From this, information related to emotions, such as the specific content of operations, facial expressions, and tone of voice, is obtained. For example, if a user is watching a video on a smartphone app, their facial expressions and the tempo of their operations are recorded. This data is then sent to a server as output.
[0368] Step 2:
[0369] The server receives data sent from the terminal and performs preprocessing. The input is the raw data sent by the terminal. Specifically, it performs noise reduction and converts audio and video data into a format suitable for emotion recognition. For example, it removes noise from audio data and smooths facial expression data frame by frame. As output, the preprocessed data is passed to the emotion engine.
[0370] Step 3:
[0371] The server uses pre-processed data to perform analysis with its emotion engine. The input is data that has been de-noised and transformed. The emotion engine analyzes this data to identify emotional states such as "joy," "anxiety," and "frustration." For example, it estimates what emotions the user is experiencing by analyzing subtle facial muscle movements and vocal intonation patterns. The analysis results are output as emotion data, and the process proceeds to the next step.
[0372] Step 4:
[0373] The server inputs emotional and operational data into a multimodal generation AI model. This input consists of emotional data obtained in the previous step and initial operational data. This AI model comprehensively analyzes various data to identify areas for improvement in the user experience. For example, it can be used to identify the cause of discomfort a user experiences with a particular interface. The analysis results are output as suggestions for improving the generated user experience.
[0374] Step 5:
[0375] The server automatically generates a new user interface design based on the generated improvement suggestions. The input is the improvement suggestions output by the AI model. In this step, the user interface is designed with emotionally sensitive elements in mind, for example, generating a design that enhances guidance and help in stressful situations. The output is a new user interface, which then proceeds to evaluation testing.
[0376] Step 6:
[0377] The terminal conducts evaluation tests on users using the generated user interface. The input is the newly designed interface design. Users interact with this interface and provide further data through their experience. Specifically, feedback is obtained regarding the usability and effectiveness of the interface design. The output is the operation data and feedback, which are used for further analysis on the server.
[0378] Step 7:
[0379] The server analyzes data obtained from evaluation tests to identify the optimal user interface. The input is the data obtained as a result of the evaluation tests. This data is analyzed in detail to confirm whether the initial improvement proposals were actually effective, and feedback for redesign is provided as needed. For example, in a learning app, user stress metrics are monitored to evaluate interface improvements. Ultimately, the optimal design is determined.
[0380] (Application Example 2)
[0381] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0382] Current interface designs often fail to consider the user's emotions, making it difficult to provide an optimal user interface based on the emotional state of individual users. As a result, stress and frustration accumulate among users when using a site, leading to a decline in the overall quality of the experience.
[0383] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0384] In this invention, the server includes means for collecting user behavior data, means for acquiring visual and auditory data, and means for analyzing the acquired data and recognizing the user's emotional state, equipped with an emotion engine. This makes it possible to provide an individually optimized user interface based on the user's emotional state and improve the quality of the user experience.
[0385] "User behavior data" refers to information about input operations and actions that occur when a user interacts with the system.
[0386] "Visual and auditory data" refers to data related to the user's facial expressions and voice, acquired through cameras and microphones.
[0387] The term "emotion engine" refers to a function that analyzes the user's emotional state based on acquired data and recognizes emotions such as joy, anger, sadness, and happiness.
[0388] A "multimodal generation AI model" refers to an artificial intelligence model that integrates and analyzes multiple different types of data, generating suggestions for improving the user experience while considering user behavior and emotions.
[0389] "User interface" refers to the design elements and layouts that allow users to interact with a computer system.
[0390] An "evaluation test" refers to a test conducted to check the effectiveness of the generated user interface and to consider the need for improvement.
[0391] "Text data" refers to information obtained by converting audio or other data formats into written form.
[0392] The system for implementing this invention mainly consists of a user terminal and a server. The terminal collects user operation data, visual data, and auditory data in real time. Specifically, the terminal's camera captures facial expressions and its microphone captures voice, and this data includes basic information about the user's interactions. This data also serves as foundational data for indicating the user's emotional state.
[0393] The server receives data transmitted from the terminal and performs preprocessing such as noise reduction. The emotion engine then analyzes this data to recognize the user's emotional state. This information is fed into a multimodal generation AI model for detailed analysis regarding improvements to the user experience. This AI model generates suggestions for user interfaces that take the emotional state into account.
[0394] Furthermore, the generated interface proposals are presented to the user as different user interfaces. The device then collects operational and emotional data from the user as they use the new interface, and obtains feedback. This feedback data is analyzed on the server to identify the optimal interface design, ultimately leading to improved user experience (UX).
[0395] For example, if a user on an e-commerce site shows an anxious expression on the product purchase page, the system can detect that emotion and adjust the interface to emphasize review and question features to reassure the user, or even suggest a new interface.
[0396] Examples of prompts for a generative AI model:
[0397] "If a user is showing signs of anxiety, how can we incorporate elements into the layout to provide reassurance?"
[0398] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0399] Step 1:
[0400] The device collects user operation data, visual data, and auditory data in real time. This includes facial expression data captured by the smartphone's camera and audio data recorded by the microphone. This collected data is sent to a server. The input is the user's operations and actions, and the output is the collected raw data.
[0401] Step 2:
[0402] The server receives raw data sent from the terminal and performs preprocessing such as noise reduction and data normalization. This data processing generates clean data that makes it easier for the emotion engine to perform accurate sentiment analysis.
[0403] Step 3:
[0404] The server inputs pre-processed data into the emotion engine and analyzes the user's emotional state. For example, major emotions such as joy and anxiety are identified, and these recognition results are obtained. The input data is pre-processed data, and the output is the identified emotional state.
[0405] Step 4:
[0406] The server integrates various types of information, including recognized emotion data, into a multimodal generation AI model. This AI model generates UX improvement suggestions that take each user's emotions into account. The input is various types of data, including emotion data, and the output is the generated UX improvement suggestions.
[0407] Step 5:
[0408] The server automatically generates a new user interface optimized for the user based on the generated UX improvement suggestions. This designs an interface that enhances the user experience. The input is the UX improvement suggestions, and the output is the new interface design.
[0409] Step 6:
[0410] The terminal provides the user with a new user interface and collects data and emotional responses from the user as they interact with it. This interaction data and feedback data are sent to the server as part of an evaluation test. The input is the user's interaction and feedback, and the output is the evaluation test results.
[0411] Step 7:
[0412] The server analyzes the evaluation test results and identifies the optimal interface design. Based on this analysis, if further UX improvements are needed, the design is refined. The input is the evaluation test results, and the output is the final interface design.
[0413] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0414] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0415] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0416] [Third Embodiment]
[0417] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0418] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0419] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0420] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0421] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0422] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0423] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0424] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0425] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0426] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0427] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0428] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0429] This invention provides a system for improving user experience (UX). This system efficiently collects user operation data, video, and audio data, and automatically generates improvement suggestions by analyzing them using a multimodal generation AI model. A specific embodiment of this system is described below.
[0430] First, the device collects data in real time as the user interacts with websites and applications. This includes basic actions such as clicking, scrolling, and tapping, as well as the user's facial expressions and voice feedback.
[0431] Next, the server receives the collected data and performs preprocessing such as noise reduction and data format conversion. The preprocessed data is then sent to a multimodal generative AI model, where it is analyzed using deep learning techniques. In this process, patterns of user behavior and areas for improvement are discovered, and specific UX improvement proposals are automatically generated.
[0432] The generated improvement proposals are automatically designed as different user interfaces for each server. This creates various versions of the interface design, preparing them for evaluation testing.
[0433] Next, the terminal conducts evaluation tests with users using the previously generated interface design. User operation data and feedback for each interface are collected again, and the server analyzes the results in detail. The optimal interface design is identified and compiled into a report and proposal.
[0434] Users participate in the system's improvement cycle by experiencing an interface based on the proposed optimal UX and providing feedback. By repeating this process, the UX continuously improves, ultimately leading to increased user satisfaction.
[0435] As a concrete example, if users on an e-commerce site find it difficult to find products, this system analyzes user browsing data and voice feedback to generate suggestions for improving the product search function. A new search interface is automatically generated, evaluated by several users, and the most effective interface is adopted based on the results. This entire process is largely automated, enabling efficient UX improvement.
[0436] The following describes the processing flow.
[0437] Step 1:
[0438] The device collects user operation data (clicks, scrolls, taps) in real time. Furthermore, it uses the device's camera and microphone to capture the user's facial expressions and voice feedback.
[0439] Step 2:
[0440] The server receives operation, video, and audio data transmitted from the terminal. It then performs preprocessing to remove noise and convert the data format to one suitable for analysis.
[0441] Step 3:
[0442] The server inputs pre-processed data into a multimodal generation AI model. This AI model identifies UX issues faced by the user by analyzing user behavior patterns and feedback from voice.
[0443] Step 4:
[0444] The server automatically generates specific UX improvement proposals using AI based on identified problems. In this process, it utilizes insights gained from user behavior patterns.
[0445] Step 5:
[0446] The server automatically constructs different user interface designs based on the generated improvement suggestions. These designs include new design elements and improvements.
[0447] Step 6:
[0448] The terminal provides each generated interface design to the target users and conducts evaluation tests. During this process, user operation data and feedback are collected again.
[0449] Step 7:
[0450] The server analyzes the results of evaluation tests and assesses user reactions and effectiveness for each interface version. This helps identify the optimal interface design.
[0451] Step 8:
[0452] Users participate in the continuous improvement process by experiencing the optimized interface and providing further feedback. This feedback will be used in the next improvement cycle.
[0453] (Example 1)
[0454] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0455] Traditional methods for improving user experience lacked mechanisms to fully utilize user behavior data, and analyzing feedback based on users' intuition and emotions was difficult. As a result, effective user interface design and improvement were not achieved to the extent expected.
[0456] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0457] In this invention, the server includes means for collecting information on user operations using an information processing device, means for acquiring image and sound data, and means for analyzing the acquired data using a multimodal generation artificial intelligence model and automatically generating improvement proposals regarding the user experience. This makes it possible to automatically generate highly accurate interface improvement proposals that comprehensively utilize a wide range of user feedback.
[0458] An "information processing device" is a device used to acquire, analyze, and manage information related to user operations.
[0459] "Information related to operation" refers to operational data that a user uses when interacting with a digital interface, including data such as clicks, scrolls, and taps.
[0460] "Image and audio data" refers to data that records the user's visual and auditory feedback, and is collected using cameras and microphones.
[0461] A "multimodal generation artificial intelligence model" is an artificial intelligence technology used to integrate and analyze multiple data modalities to extract insights.
[0462] "User interface" refers to the design elements of screens and operating systems that allow users to interact with a system.
[0463] An "evaluation test" is a test conducted to verify how effective the generated user interface is for the user.
[0464] This invention relates to an information system for improving user experience (UX). This system efficiently collects information, images, and audio data related to user operations and automatically generates improvement suggestions by analyzing them using a multi-modal generation artificial intelligence model.
[0465] The device collects information in real time as the user interacts with websites and applications. This information includes basic actions such as clicks, scrolls, and taps, as well as facial expression and audio data collected using the camera and microphone. Specifically, this objective is achieved by utilizing UI tracking tools and speech recognition software.
[0466] The server receives data sent from the terminal. This data is first preprocessed by denoising and converting the data format. The preprocessed data is then input into a multimodal generation artificial intelligence model. The AI model analyzes this data to identify user behavior patterns and areas for UX improvement, and automatically generates improvement suggestions. Machine learning frameworks such as TensorFlow and PyTorch are used to build the AI model.
[0467] Users experience the optimized interface based on the proposed improvements and provide feedback. This feedback will be used in the next cycle to contribute to further improvements in the user experience (UX).
[0468] For example, when addressing the difficulties users experience when searching for products on an e-commerce site, this system uses user interaction data and voice feedback to generate suggestions for improving the product search function. New search interfaces are automatically generated, evaluated by a small number of users, and the optimal interface is selected, thus promoting automatic and efficient system improvement.
[0469] An example of a prompt message might be: "Analyze the difficulties users encounter when searching for a specific product category and suggest a more intuitive search interface."
[0470] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0471] Step 1:
[0472] The device collects data as users interact with websites and applications. As input, user actions such as clicks, scrolls, and taps are captured in real time. This data is initially converted into a digital format using UI tracking tools (e.g., event logging). The output is a dataset recording user actions.
[0473] Step 2:
[0474] The device collects facial expression data and acoustic feedback. Inputs include video and audio data acquired using a camera and microphone. This data is then processed by image recognition and speech recognition software to extract changes in the user's facial expressions and voice tone. The output is attribute data representing the results of the facial and audio analysis.
[0475] Step 3:
[0476] The server receives all collected data and performs preprocessing. The input consists of operation data, facial expression data, and voice data sent from the terminal. The server removes noise using filtering techniques and standardizes the data format (e.g., data cleaning, normalization). The output is a clean dataset suitable for data analysis.
[0477] Step 4:
[0478] The server inputs pre-processed data into a multi-modal generation AI model. This input consists of cleaned operation data, facial expressions, and voice data. The server feeds this data into the AI model and analyzes the relationships between the data. Using deep learning techniques, it discovers patterns in user behavior and areas for UX improvement. The output is a concrete UX improvement proposal based on the analysis results.
[0479] Step 5:
[0480] The server automatically designs different user interfaces based on improvement suggestions generated by the AI model. The input is the specification of the improvement suggestion. The server uses design software to generate multiple versions of the interface (e.g., automatic generation of UI design). The output is a test-ready interface design.
[0481] Step 6:
[0482] The device undergoes evaluation testing with users using the designed interface. The input is the generated interface design. User interaction data and feedback are collected again (e.g., interaction log recording, surveys). The output is the evaluation data obtained from the test.
[0483] Step 7:
[0484] The server analyzes evaluation data in detail to identify the optimal interface design. The input is the results of evaluation tests. The server uses statistical analysis methods and machine learning models to evaluate the data and select the most effective interface design. The output is a report of the optimized interface, compiled as a proposal.
[0485] (Application Example 1)
[0486] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0487] In modern e-commerce, improving the user experience is a crucial element for gaining a competitive edge. However, it is technically challenging to grasp user interaction data and emotional responses in real time and instantly optimize the interface based on that data. Therefore, more advanced data analysis technologies and responsive systems are required to rapidly and continuously improve user satisfaction.
[0488] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0489] In this invention, the server includes means for collecting user operation information in real time, means for acquiring and analyzing video and audio information using a smart device, and means for automatically generating improvement suggestions using multimodal generation AI. This enables rapid evaluation and suggestion of the user experience based on the user's emotional and operational responses.
[0490] "User interaction information" refers to data related to user interactions such as clicks, scrolls, and taps on websites and applications.
[0491] "Video information" refers to digital image and video data acquired through a camera, capturing the user's facial expressions, posture, and other details during operation.
[0492] "Voice information" refers to digital audio data related to the voice and voice feedback that a user makes while operating the device.
[0493] A "multimodal generation AI" is an artificial intelligence model that combines and analyzes multiple different data formats and modalities (e.g., video and audio) to gain insights.
[0494] An "improvement suggestion" is a proposal for changes or optimizations to the user interface that is automatically generated by the system with the aim of improving the user experience.
[0495] "User information display" refers to the visual presentation methods and interaction designs for information that users can see on the screen.
[0496] An "evaluation experiment" is a testing process that involves analyzing feedback obtained from actual users who test the system's generated user information display.
[0497] A "smart device" refers to a mobile device, such as a mobile terminal or wearable device, that has internet connectivity and advanced computing capabilities.
[0498] "Real-time" refers to a processing method that instantly acquires user behavior and feedback, and quickly analyzes and responds to it.
[0499] To realize this invention, a system is needed that collects user operation information and emotional feedback in real time, analyzes them, and generates appropriate improvement suggestions. This system primarily involves three parties: the server, the terminal, and the user.
[0500] First, let's explain the role of the terminal. The terminal is designed to collect user operation information in real time, and simultaneously acquires video and audio information through the smart device. In this process, the terminal efficiently stores data using OpenCV for image processing and libraries for audio processing. For example, the camera on the terminal captures the user's facial expressions, and the microphone acquires audio feedback.
[0501] Next, the server receives this acquired data and performs preprocessing and analysis. Here, noise reduction and data formatting are performed, and the data is analyzed using a multimodal generative AI utilizing deep learning frameworks such as TensorFlow or PyTorch. The generative AI model fuses multiple data sources, automatically identifies UX patterns that can be improved, and quickly generates improvement suggestions. As a concrete example, prompts such as "Please suggest a more user-friendly layout by making the hierarchy of this category shallower" can be generated, enabling the suggestion of a new UI.
[0502] Finally, let's discuss the user's role. Users experience the proposed new user interface and provide feedback. This feedback is then sent back to the server, and the AI goes through an iterative optimization process to provide the optimal UX. Through this iterative process, the UX is continuously improved, and it is expected that a high level of user satisfaction will ultimately be achieved.
[0503] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0504] Step 1:
[0505] The device collects user action information in real time. Input consists of user actions such as clicks, scrolls, and taps. The device acquires this information as a data stream and stores it in a database as initial processing. Output is the action data formatted for analysis by the server.
[0506] Step 2:
[0507] The device acquires video and audio information using a smart device. The input consists of the user's facial expressions and voice captured through the device's camera and microphone. Features are extracted from the image data using OpenCV, and the audio data is acoustically preprocessed to convert it into a format suitable for analysis. The output consists of processed video and audio information for analysis.
[0508] Step 3:
[0509] The server receives processed operation data, video information, and audio information, and performs noise reduction and data format conversion. The input consists of data in multiple formats sent from the terminal, and the server integrates these to build a consistent, analyzable dataset. The output is integrated data in a state suitable for analysis.
[0510] Step 4:
[0511] The server feeds the integrated data into a multimodal generative AI model. The input is the data integrated in the previous step. The server applies the generative AI model using TensorFlow or PyTorch, analyzing user behavior and emotional responses through deep learning. The output is the discovery of UX patterns and the automatic generation of improvement suggestions.
[0512] Step 5:
[0513] The server automatically generates a new user interface based on the proposed improvements. The input is the improvement suggestions obtained in the previous step, and the server prompts the AI to create various UX design options. The output is the new user interface to be tested.
[0514] Step 6:
[0515] The user experiences the new user interface and provides feedback to the server. Inputs are interactions with the generated interface and user feedback. Outputs are evaluation data returned to the server, which the system uses for further optimization.
[0516] Step 7:
[0517] The server analyzes user feedback and adjusts the generated AI model as needed to provide the optimal UX. The input is evaluation feedback collected from users. Based on this data, the server performs pattern analysis to further improve the overall user experience of the system. The output is an optimized UX suggestion.
[0518] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0519] The present invention aims to enhance user experience (UX) through a system that incorporates an emotion engine. This system can generate more accurate UX improvement proposals by considering not only user behavior data but also emotions. A specific embodiment of this system is described below.
[0520] First, the device acquires user operation data, video, and audio data in real time. This data includes basic information about the user's interactions, as well as data related to emotions, such as facial expressions and speech.
[0521] Next, the server receives the collected data and performs preprocessing such as noise reduction and data transformation. This data is then fed into the emotion engine, which recognizes emotional states (e.g., joy, anxiety, frustration). The recognized emotion data, along with other operational data, is input into a multimodal generative AI model.
[0522] This AI model performs a detailed analysis that takes user emotions into account, identifying obstacles and areas for improvement in the user experience. Based on this, a process automatically generates more targeted UX improvement suggestions. These suggestions include elements that improve the quality of the user's perceived experience itself.
[0523] The generated improvement suggestions are automatically designed as different user interfaces for each server. These interfaces incorporate elements designed to improve user emotional responses. Based on this, a new interface design is generated, and evaluation tests are conducted using it.
[0524] The terminal provides users with an interface design specifically created for evaluation testing, and collects user interaction data, sentiment data, and feedback. The server then analyzes this data in detail to identify the optimal interface design.
[0525] As a concrete example, in an educational application, if a user expresses frustration with a particular learning content, this system identifies that emotion and generates design improvement suggestions to reduce the stress experienced with that content. New content layouts and supplementary learning support are designed, and an interface that considers the user's emotions is provided, leading to improved learning efficiency. This entire process is automated, enabling efficient UX improvement.
[0526] The following describes the processing flow.
[0527] Step 1:
[0528] The device begins collecting data about the user's interactions. Specifically, in addition to operation data such as clicks, scrolls, and taps, it also uses the camera and microphone to capture facial expressions and voice data.
[0529] Step 2:
[0530] The server receives operation data, video data, and audio data transmitted from the terminal. After receiving the data, it cleanses it, removes noise, and converts it into a format suitable for analysis.
[0531] Step 3:
[0532] The server inputs the pre-processed data into the emotion engine. The emotion engine recognizes the user's emotions from facial and voice data and generates emotion data such as joy, anxiety, and frustration.
[0533] Step 4:
[0534] The server inputs emotional and operational data into a multimodal generation AI model for analysis. This automatically generates specific UX improvement suggestions to enhance the user experience. Because emotional data is included, the improvement suggestions take into account the user's emotional responses.
[0535] Step 5:
[0536] The server automatically generates different user interface designs based on the proposed improvements. At this stage, emotionally sensitive elements are incorporated into the design.
[0537] Step 6:
[0538] The device presents the created user interface to the user as an evaluation test. It then collects user operation data and changes in emotions.
[0539] Step 7:
[0540] The server analyzes data obtained from evaluation tests to assess user reactions and the effectiveness of each user interface. This identifies the most emotionally resonant and optimal interface.
[0541] Step 8:
[0542] Users experience the optimized user interface and provide feedback again. Based on this feedback, a new improvement cycle is initiated.
[0543] (Example 2)
[0544] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0545] Traditional user experience improvement systems focus solely on user interaction data and do not take emotional information into account. This has resulted in a failure to adequately reflect the dissatisfaction and stress that users may be experiencing. Furthermore, the generated improvement suggestions are not always optimal in response to user emotions, making it difficult to effectively improve user experience satisfaction.
[0546] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0547] In this invention, the server includes means for collecting user behavior information and emotional information in real time, means for pre-processing the collected data by removing noise and performing data transformation, and means for recognizing the emotional state using an emotion engine. This makes it possible to generate suggestions for improving the user experience that take emotions into account and to design an optimal user interface based on these suggestions.
[0548] "User action information" refers to a series of operation records and interaction data generated when a user operates the system.
[0549] "Emotional information" refers to data that indicates the emotional state obtained through the user's facial expressions, tone of voice, and other nonverbal cues.
[0550] "Noise reduction" refers to the process of removing unnecessary information and errors before data analysis, preparing the data for analysis.
[0551] "Data transformation preprocessing" refers to a series of processes for converting collected data into a format that can be analyzed.
[0552] An "emotion engine" refers to a specific algorithm or model used to recognize a user's emotional state.
[0553] A "multimodal generation AI model" refers to an artificial intelligence model that combines and analyzes multiple different types of data to generate useful output.
[0554] "User experience improvement proposals" refer to suggested changes and optimizations based on user interactions, aimed at providing a better user experience.
[0555] A "user interface" refers to a set of visual or manipulative elements that enable a system and a user to interact directly.
[0556] "Evaluation testing" refers to a series of tests and evaluation procedures conducted to confirm the effectiveness of a new user interface or function.
[0557] This system was developed with the aim of improving the user experience, and it involves the coordinated functioning of multiple hardware and software components.
[0558] The device plays a role in collecting user operation data, video, and audio data in real time. This collection uses the camera and microphone built into the device. This data records how the user is using the system and reflects their emotional state at the time. For example, it captures the facial expressions and tone of voice a user may have while using an online educational app.
[0559] The server receives data sent from the terminal and first performs noise reduction and data transformation. This makes the data suitable for analysis and removes unnecessary information. Next, this data is analyzed by the emotion engine to identify the user's emotional state. Based on this emotional data, it is integrated with the operation data and input into the multimodal generation AI model.
[0560] This AI model has the ability to analyze recognized emotions and interaction information and generate suggestions for improving the user experience. These suggestions include proposals to enhance the quality of the experience based on the user's emotions. For example, in an educational app, if a user experiences stress or confusion during learning activities, appropriate support features or interface changes will be suggested.
[0561] The generated improvement proposals are then automatically regenerated by the server as a new user interface. Here, elements that improve user emotional responses are incorporated into the interface. Furthermore, evaluation tests are conducted using the new interface design to confirm its effectiveness.
[0562] As a concrete example, a prompt message might be input into the generative AI model in the form of, "Based on the emotions the user expressed regarding a specific issue, suggest which parts should be improved and how." This aims to improve the quality of the user experience.
[0563] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0564] Step 1:
[0565] The device collects user operation data, video, and audio data in real time. Input here is video and audio from the camera and microphone connected to the device. From this, information related to emotions, such as the specific content of operations, facial expressions, and tone of voice, is obtained. For example, if a user is watching a video on a smartphone app, their facial expressions and the tempo of their operations are recorded. This data is then sent to a server as output.
[0566] Step 2:
[0567] The server receives data sent from the terminal and performs preprocessing. The input is the raw data sent by the terminal. Specifically, it performs noise reduction and converts audio and video data into a format suitable for emotion recognition. For example, it removes noise from audio data and smooths facial expression data frame by frame. As output, the preprocessed data is passed to the emotion engine.
[0568] Step 3:
[0569] The server uses pre-processed data to perform analysis with its emotion engine. The input is data that has been de-noised and transformed. The emotion engine analyzes this data to identify emotional states such as "joy," "anxiety," and "frustration." For example, it estimates what emotions the user is experiencing by analyzing subtle facial muscle movements and vocal intonation patterns. The analysis results are output as emotion data, and the process proceeds to the next step.
[0570] Step 4:
[0571] The server inputs emotional and operational data into a multimodal generation AI model. This input consists of emotional data obtained in the previous step and initial operational data. This AI model comprehensively analyzes various data to identify areas for improvement in the user experience. For example, it can be used to identify the cause of discomfort a user experiences with a particular interface. The analysis results are output as suggestions for improving the generated user experience.
[0572] Step 5:
[0573] The server automatically generates a new user interface design based on the generated improvement suggestions. The input is the improvement suggestions output by the AI model. In this step, the user interface is designed with emotionally sensitive elements in mind, for example, generating a design that enhances guidance and help in stressful situations. The output is a new user interface, which then proceeds to evaluation testing.
[0574] Step 6:
[0575] The terminal conducts evaluation tests on users using the generated user interface. The input is the newly designed interface design. Users interact with this interface and provide further data through their experience. Specifically, feedback is obtained regarding the usability and effectiveness of the interface design. The output is the operation data and feedback, which are used for further analysis on the server.
[0576] Step 7:
[0577] The server analyzes data obtained from evaluation tests to identify the optimal user interface. The input is the data obtained as a result of the evaluation tests. This data is analyzed in detail to confirm whether the initial improvement proposals were actually effective, and feedback for redesign is provided as needed. For example, in a learning app, user stress metrics are monitored to evaluate interface improvements. Ultimately, the optimal design is determined.
[0578] (Application Example 2)
[0579] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0580] Current interface designs often fail to consider the user's emotions, making it difficult to provide an optimal user interface based on the emotional state of individual users. As a result, stress and frustration accumulate among users when using a site, leading to a decline in the overall quality of the experience.
[0581] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0582] In this invention, the server includes means for collecting user behavior data, means for acquiring visual and auditory data, and means for analyzing the acquired data and recognizing the user's emotional state, equipped with an emotion engine. This makes it possible to provide an individually optimized user interface based on the user's emotional state and improve the quality of the user experience.
[0583] "User behavior data" refers to information about input operations and actions that occur when a user interacts with the system.
[0584] "Visual and auditory data" refers to data related to the user's facial expressions and voice, acquired through cameras and microphones.
[0585] The term "emotion engine" refers to a function that analyzes the user's emotional state based on acquired data and recognizes emotions such as joy, anger, sadness, and happiness.
[0586] A "multimodal generation AI model" refers to an artificial intelligence model that integrates and analyzes multiple different types of data, generating suggestions for improving the user experience while considering user behavior and emotions.
[0587] "User interface" refers to the design elements and layouts that allow users to interact with a computer system.
[0588] An "evaluation test" refers to a test conducted to check the effectiveness of the generated user interface and to consider the need for improvement.
[0589] "Text data" refers to information obtained by converting audio or other data formats into written form.
[0590] The system for implementing this invention mainly consists of a user terminal and a server. The terminal collects user operation data, visual data, and auditory data in real time. Specifically, the terminal's camera captures facial expressions and its microphone captures voice, and this data includes basic information about the user's interactions. This data also serves as foundational data for indicating the user's emotional state.
[0591] The server receives data transmitted from the terminal and performs preprocessing such as noise reduction. The emotion engine then analyzes this data to recognize the user's emotional state. This information is fed into a multimodal generation AI model for detailed analysis regarding improvements to the user experience. This AI model generates suggestions for user interfaces that take the emotional state into account.
[0592] Furthermore, the generated interface proposals are presented to the user as different user interfaces. The device then collects operational and emotional data from the user as they use the new interface, and obtains feedback. This feedback data is analyzed on the server to identify the optimal interface design, ultimately leading to improved user experience (UX).
[0593] For example, if a user on an e-commerce site shows an anxious expression on the product purchase page, the system can detect that emotion and adjust the interface to emphasize review and question features to reassure the user, or even suggest a new interface.
[0594] Examples of prompts for a generative AI model:
[0595] "If a user is showing signs of anxiety, how can we incorporate elements into the layout to provide reassurance?"
[0596] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0597] Step 1:
[0598] The device collects user operation data, visual data, and auditory data in real time. This includes facial expression data captured by the smartphone's camera and audio data recorded by the microphone. This collected data is sent to a server. The input is the user's operations and actions, and the output is the collected raw data.
[0599] Step 2:
[0600] The server receives raw data sent from the terminal and performs preprocessing such as noise reduction and data normalization. This data processing generates clean data that makes it easier for the emotion engine to perform accurate sentiment analysis.
[0601] Step 3:
[0602] The server inputs pre-processed data into the emotion engine and analyzes the user's emotional state. For example, major emotions such as joy and anxiety are identified, and these recognition results are obtained. The input data is pre-processed data, and the output is the identified emotional state.
[0603] Step 4:
[0604] The server integrates various types of information, including recognized emotion data, into a multimodal generation AI model. This AI model generates UX improvement suggestions that take each user's emotions into account. The input is various types of data, including emotion data, and the output is the generated UX improvement suggestions.
[0605] Step 5:
[0606] The server automatically generates a new user interface optimized for the user based on the generated UX improvement suggestions. This designs an interface that enhances the user experience. The input is the UX improvement suggestions, and the output is the new interface design.
[0607] Step 6:
[0608] The terminal provides the user with a new user interface and collects data and emotional responses from the user as they interact with it. This interaction data and feedback data are sent to the server as part of an evaluation test. The input is the user's interaction and feedback, and the output is the evaluation test results.
[0609] Step 7:
[0610] The server analyzes the evaluation test results and identifies the optimal interface design. Based on this analysis, if further UX improvements are needed, the design is refined. The input is the evaluation test results, and the output is the final interface design.
[0611] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0612] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0613] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0614] [Fourth Embodiment]
[0615] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0616] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0617] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0618] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0619] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0620] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0621] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0622] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0623] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0624] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0625] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0626] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0627] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0628] This invention provides a system for improving user experience (UX). This system efficiently collects user operation data, video, and audio data, and automatically generates improvement suggestions by analyzing them using a multimodal generation AI model. A specific embodiment of this system is described below.
[0629] First, the device collects data in real time as the user interacts with websites and applications. This includes basic actions such as clicking, scrolling, and tapping, as well as the user's facial expressions and voice feedback.
[0630] Next, the server receives the collected data and performs preprocessing such as noise reduction and data format conversion. The preprocessed data is then sent to a multimodal generative AI model, where it is analyzed using deep learning techniques. In this process, patterns of user behavior and areas for improvement are discovered, and specific UX improvement proposals are automatically generated.
[0631] The generated improvement proposals are automatically designed as different user interfaces for each server. This creates various versions of the interface design, preparing them for evaluation testing.
[0632] Next, the terminal conducts evaluation tests with users using the previously generated interface design. User operation data and feedback for each interface are collected again, and the server analyzes the results in detail. The optimal interface design is identified and compiled into a report and proposal.
[0633] Users participate in the system's improvement cycle by experiencing an interface based on the proposed optimal UX and providing feedback. By repeating this process, the UX continuously improves, ultimately leading to increased user satisfaction.
[0634] As a concrete example, if users on an e-commerce site find it difficult to find products, this system analyzes user browsing data and voice feedback to generate suggestions for improving the product search function. A new search interface is automatically generated, evaluated by several users, and the most effective interface is adopted based on the results. This entire process is largely automated, enabling efficient UX improvement.
[0635] The following describes the processing flow.
[0636] Step 1:
[0637] The device collects user operation data (clicks, scrolls, taps) in real time. Furthermore, it uses the device's camera and microphone to capture the user's facial expressions and voice feedback.
[0638] Step 2:
[0639] The server receives operation, video, and audio data transmitted from the terminal. It then performs preprocessing to remove noise and convert the data format to one suitable for analysis.
[0640] Step 3:
[0641] The server inputs pre-processed data into a multimodal generation AI model. This AI model identifies UX issues faced by the user by analyzing user behavior patterns and feedback from voice.
[0642] Step 4:
[0643] The server automatically generates specific UX improvement proposals using AI based on identified problems. In this process, it utilizes insights gained from user behavior patterns.
[0644] Step 5:
[0645] The server automatically constructs different user interface designs based on the generated improvement suggestions. These designs include new design elements and improvements.
[0646] Step 6:
[0647] The terminal provides each generated interface design to the target users and conducts evaluation tests. During this process, user operation data and feedback are collected again.
[0648] Step 7:
[0649] The server analyzes the results of evaluation tests and assesses user reactions and effectiveness for each interface version. This helps identify the optimal interface design.
[0650] Step 8:
[0651] Users participate in the continuous improvement process by experiencing the optimized interface and providing further feedback. This feedback will be used in the next improvement cycle.
[0652] (Example 1)
[0653] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0654] Traditional methods for improving user experience lacked mechanisms to fully utilize user behavior data, and analyzing feedback based on users' intuition and emotions was difficult. As a result, effective user interface design and improvement were not achieved to the extent expected.
[0655] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0656] In this invention, the server includes means for collecting information on user operations using an information processing device, means for acquiring image and sound data, and means for analyzing the acquired data using a multimodal generation artificial intelligence model and automatically generating improvement proposals regarding the user experience. This makes it possible to automatically generate highly accurate interface improvement proposals that comprehensively utilize a wide range of user feedback.
[0657] An "information processing device" is a device used to acquire, analyze, and manage information related to user operations.
[0658] "Information related to operation" refers to operational data that a user uses when interacting with a digital interface, including data such as clicks, scrolls, and taps.
[0659] "Image and audio data" refers to data that records the user's visual and auditory feedback, and is collected using cameras and microphones.
[0660] A "multimodal generation artificial intelligence model" is an artificial intelligence technology used to integrate and analyze multiple data modalities to extract insights.
[0661] "User interface" refers to the design elements of screens and operating systems that allow users to interact with a system.
[0662] An "evaluation test" is a test conducted to verify how effective the generated user interface is for the user.
[0663] This invention relates to an information system for improving user experience (UX). This system efficiently collects information, images, and audio data related to user operations and automatically generates improvement suggestions by analyzing them using a multi-modal generation artificial intelligence model.
[0664] The device collects information in real time as the user interacts with websites and applications. This information includes basic actions such as clicks, scrolls, and taps, as well as facial expression and audio data collected using the camera and microphone. Specifically, this objective is achieved by utilizing UI tracking tools and speech recognition software.
[0665] The server receives data sent from the terminal. This data is first preprocessed by denoising and converting the data format. The preprocessed data is then input into a multimodal generation artificial intelligence model. The AI model analyzes this data to identify user behavior patterns and areas for UX improvement, and automatically generates improvement suggestions. Machine learning frameworks such as TensorFlow and PyTorch are used to build the AI model.
[0666] Users experience the optimized interface based on the proposed improvements and provide feedback. This feedback will be used in the next cycle to contribute to further improvements in the user experience (UX).
[0667] For example, when addressing the difficulties users experience when searching for products on an e-commerce site, this system uses user interaction data and voice feedback to generate suggestions for improving the product search function. New search interfaces are automatically generated, evaluated by a small number of users, and the optimal interface is selected, thus promoting automatic and efficient system improvement.
[0668] An example of a prompt message might be: "Analyze the difficulties users encounter when searching for a specific product category and suggest a more intuitive search interface."
[0669] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0670] Step 1:
[0671] The device collects data as users interact with websites and applications. As input, user actions such as clicks, scrolls, and taps are captured in real time. This data is initially converted into a digital format using UI tracking tools (e.g., event logging). The output is a dataset recording user actions.
[0672] Step 2:
[0673] The device collects facial expression data and acoustic feedback. Inputs include video and audio data acquired using a camera and microphone. This data is then processed by image recognition and speech recognition software to extract changes in the user's facial expressions and voice tone. The output is attribute data representing the results of the facial and audio analysis.
[0674] Step 3:
[0675] The server receives all collected data and performs preprocessing. The input consists of operation data, facial expression data, and voice data sent from the terminal. The server removes noise using filtering techniques and standardizes the data format (e.g., data cleaning, normalization). The output is a clean dataset suitable for data analysis.
[0676] Step 4:
[0677] The server inputs pre-processed data into a multi-modal generation AI model. This input consists of cleaned operation data, facial expressions, and voice data. The server feeds this data into the AI model and analyzes the relationships between the data. Using deep learning techniques, it discovers patterns in user behavior and areas for UX improvement. The output is a concrete UX improvement proposal based on the analysis results.
[0678] Step 5:
[0679] The server automatically designs different user interfaces based on improvement suggestions generated by the AI model. The input is the specification of the improvement suggestion. The server uses design software to generate multiple versions of the interface (e.g., automatic generation of UI design). The output is a test-ready interface design.
[0680] Step 6:
[0681] The device undergoes evaluation testing with users using the designed interface. The input is the generated interface design. User interaction data and feedback are collected again (e.g., interaction log recording, surveys). The output is the evaluation data obtained from the test.
[0682] Step 7:
[0683] The server analyzes evaluation data in detail to identify the optimal interface design. The input is the results of evaluation tests. The server uses statistical analysis methods and machine learning models to evaluate the data and select the most effective interface design. The output is a report of the optimized interface, compiled as a proposal.
[0684] (Application Example 1)
[0685] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0686] In modern e-commerce, improving the user experience is a crucial element for gaining a competitive edge. However, it is technically challenging to grasp user interaction data and emotional responses in real time and instantly optimize the interface based on that data. Therefore, more advanced data analysis technologies and responsive systems are required to rapidly and continuously improve user satisfaction.
[0687] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0688] In this invention, the server includes means for collecting user operation information in real time, means for acquiring and analyzing video and audio information using a smart device, and means for automatically generating improvement suggestions using multimodal generation AI. This enables rapid evaluation and suggestion of the user experience based on the user's emotional and operational responses.
[0689] "User interaction information" refers to data related to user interactions such as clicks, scrolls, and taps on websites and applications.
[0690] "Video information" refers to digital image and video data acquired through a camera, capturing the user's facial expressions, posture, and other details during operation.
[0691] "Voice information" refers to digital audio data related to the voice and voice feedback that a user makes while operating the device.
[0692] A "multimodal generation AI" is an artificial intelligence model that combines and analyzes multiple different data formats and modalities (e.g., video and audio) to gain insights.
[0693] An "improvement suggestion" is a proposal for changes or optimizations to the user interface that is automatically generated by the system with the aim of improving the user experience.
[0694] "User information display" refers to the visual presentation methods and interaction designs for information that users can see on the screen.
[0695] An "evaluation experiment" is a testing process that involves analyzing feedback obtained from actual users who test the system's generated user information display.
[0696] A "smart device" refers to a mobile device, such as a mobile terminal or wearable device, that has internet connectivity and advanced computing capabilities.
[0697] "Real-time" refers to a processing method that instantly acquires user behavior and feedback, and quickly analyzes and responds to it.
[0698] To realize this invention, a system is needed that collects user operation information and emotional feedback in real time, analyzes them, and generates appropriate improvement suggestions. This system primarily involves three parties: the server, the terminal, and the user.
[0699] First, let's explain the role of the terminal. The terminal is designed to collect user operation information in real time, and simultaneously acquires video and audio information through the smart device. In this process, the terminal efficiently stores data using OpenCV for image processing and libraries for audio processing. For example, the camera on the terminal captures the user's facial expressions, and the microphone acquires audio feedback.
[0700] Next, the server receives this acquired data and performs preprocessing and analysis. Here, noise reduction and data formatting are performed, and the data is analyzed using a multimodal generative AI utilizing deep learning frameworks such as TensorFlow or PyTorch. The generative AI model fuses multiple data sources, automatically identifies UX patterns that can be improved, and quickly generates improvement suggestions. As a concrete example, prompts such as "Please suggest a more user-friendly layout by making the hierarchy of this category shallower" can be generated, enabling the suggestion of a new UI.
[0701] Finally, let's discuss the user's role. Users experience the proposed new user interface and provide feedback. This feedback is then sent back to the server, and the AI goes through an iterative optimization process to provide the optimal UX. Through this iterative process, the UX is continuously improved, and it is expected that a high level of user satisfaction will ultimately be achieved.
[0702] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0703] Step 1:
[0704] The device collects user action information in real time. Input consists of user actions such as clicks, scrolls, and taps. The device acquires this information as a data stream and stores it in a database as initial processing. Output is the action data formatted for analysis by the server.
[0705] Step 2:
[0706] The device acquires video and audio information using a smart device. The input consists of the user's facial expressions and voice captured through the device's camera and microphone. Features are extracted from the image data using OpenCV, and the audio data is acoustically preprocessed to convert it into a format suitable for analysis. The output consists of processed video and audio information for analysis.
[0707] Step 3:
[0708] The server receives processed operation data, video information, and audio information, and performs noise reduction and data format conversion. The input consists of data in multiple formats sent from the terminal, and the server integrates these to build a consistent, analyzable dataset. The output is integrated data in a state suitable for analysis.
[0709] Step 4:
[0710] The server feeds the integrated data into a multimodal generative AI model. The input is the data integrated in the previous step. The server applies the generative AI model using TensorFlow or PyTorch, analyzing user behavior and emotional responses through deep learning. The output is the discovery of UX patterns and the automatic generation of improvement suggestions.
[0711] Step 5:
[0712] The server automatically generates a new user interface based on the proposed improvements. The input is the improvement suggestions obtained in the previous step, and the server prompts the AI to create various UX design options. The output is the new user interface to be tested.
[0713] Step 6:
[0714] The user experiences the new user interface and provides feedback to the server. Inputs are interactions with the generated interface and user feedback. Outputs are evaluation data returned to the server, which the system uses for further optimization.
[0715] Step 7:
[0716] The server analyzes user feedback and adjusts the generated AI model as needed to provide the optimal UX. The input is evaluation feedback collected from users. Based on this data, the server performs pattern analysis to further improve the overall user experience of the system. The output is an optimized UX suggestion.
[0717] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0718] The present invention aims to enhance user experience (UX) through a system that incorporates an emotion engine. This system can generate more accurate UX improvement proposals by considering not only user behavior data but also emotions. A specific embodiment of this system is described below.
[0719] First, the device acquires user operation data, video, and audio data in real time. This data includes basic information about the user's interactions, as well as data related to emotions, such as facial expressions and speech.
[0720] Next, the server receives the collected data and performs preprocessing such as noise reduction and data transformation. This data is then fed into the emotion engine, which recognizes emotional states (e.g., joy, anxiety, frustration). The recognized emotion data, along with other operational data, is input into a multimodal generative AI model.
[0721] This AI model performs a detailed analysis that takes user emotions into account, identifying obstacles and areas for improvement in the user experience. Based on this, a process automatically generates more targeted UX improvement suggestions. These suggestions include elements that improve the quality of the user's perceived experience itself.
[0722] The generated improvement suggestions are automatically designed as different user interfaces for each server. These interfaces incorporate elements designed to improve user emotional responses. Based on this, a new interface design is generated, and evaluation tests are conducted using it.
[0723] The terminal provides users with an interface design specifically created for evaluation testing, and collects user interaction data, sentiment data, and feedback. The server then analyzes this data in detail to identify the optimal interface design.
[0724] As a concrete example, in an educational application, if a user expresses frustration with a particular learning content, this system identifies that emotion and generates design improvement suggestions to reduce the stress experienced with that content. New content layouts and supplementary learning support are designed, and an interface that considers the user's emotions is provided, leading to improved learning efficiency. This entire process is automated, enabling efficient UX improvement.
[0725] The following describes the processing flow.
[0726] Step 1:
[0727] The device begins collecting data about the user's interactions. Specifically, in addition to operation data such as clicks, scrolls, and taps, it also uses the camera and microphone to capture facial expressions and voice data.
[0728] Step 2:
[0729] The server receives operation data, video data, and audio data transmitted from the terminal. After receiving the data, it cleanses it, removes noise, and converts it into a format suitable for analysis.
[0730] Step 3:
[0731] The server inputs the pre-processed data into the emotion engine. The emotion engine recognizes the user's emotions from facial and voice data and generates emotion data such as joy, anxiety, and frustration.
[0732] Step 4:
[0733] The server inputs emotional and operational data into a multimodal generation AI model for analysis. This automatically generates specific UX improvement suggestions to enhance the user experience. Because emotional data is included, the improvement suggestions take into account the user's emotional responses.
[0734] Step 5:
[0735] The server automatically generates different user interface designs based on the proposed improvements. At this stage, emotionally sensitive elements are incorporated into the design.
[0736] Step 6:
[0737] The device presents the created user interface to the user as an evaluation test. It then collects user operation data and changes in emotions.
[0738] Step 7:
[0739] The server analyzes data obtained from evaluation tests to assess user reactions and the effectiveness of each user interface. This identifies the most emotionally resonant and optimal interface.
[0740] Step 8:
[0741] Users experience the optimized user interface and provide feedback again. Based on this feedback, a new improvement cycle is initiated.
[0742] (Example 2)
[0743] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0744] Traditional user experience improvement systems focus solely on user interaction data and do not take emotional information into account. This has resulted in a failure to adequately reflect the dissatisfaction and stress that users may be experiencing. Furthermore, the generated improvement suggestions are not always optimal in response to user emotions, making it difficult to effectively improve user experience satisfaction.
[0745] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0746] In this invention, the server includes means for collecting user behavior information and emotional information in real time, means for pre-processing the collected data by removing noise and performing data transformation, and means for recognizing the emotional state using an emotion engine. This makes it possible to generate suggestions for improving the user experience that take emotions into account and to design an optimal user interface based on these suggestions.
[0747] "User action information" refers to a series of operation records and interaction data generated when a user operates the system.
[0748] "Emotional information" refers to data that indicates the emotional state obtained through the user's facial expressions, tone of voice, and other nonverbal cues.
[0749] "Noise reduction" refers to the process of removing unnecessary information and errors before data analysis, preparing the data for analysis.
[0750] "Data transformation preprocessing" refers to a series of processes for converting collected data into a format that can be analyzed.
[0751] An "emotion engine" refers to a specific algorithm or model used to recognize a user's emotional state.
[0752] A "multimodal generation AI model" refers to an artificial intelligence model that combines and analyzes multiple different types of data to generate useful output.
[0753] "User experience improvement proposals" refer to suggested changes and optimizations based on user interactions, aimed at providing a better user experience.
[0754] A "user interface" refers to a set of visual or manipulative elements that enable a system and a user to interact directly.
[0755] "Evaluation testing" refers to a series of tests and evaluation procedures conducted to confirm the effectiveness of a new user interface or function.
[0756] This system was developed with the aim of improving the user experience, and it involves the coordinated functioning of multiple hardware and software components.
[0757] The device plays a role in collecting user operation data, video, and audio data in real time. This collection uses the camera and microphone built into the device. This data records how the user is using the system and reflects their emotional state at the time. For example, it captures the facial expressions and tone of voice a user may have while using an online educational app.
[0758] The server receives data sent from the terminal and first performs noise reduction and data transformation. This makes the data suitable for analysis and removes unnecessary information. Next, this data is analyzed by the emotion engine to identify the user's emotional state. Based on this emotional data, it is integrated with the operation data and input into the multimodal generation AI model.
[0759] This AI model has the ability to analyze recognized emotions and interaction information and generate suggestions for improving the user experience. These suggestions include proposals to enhance the quality of the experience based on the user's emotions. For example, in an educational app, if a user experiences stress or confusion during learning activities, appropriate support features or interface changes will be suggested.
[0760] The generated improvement proposals are then automatically regenerated by the server as a new user interface. Here, elements that improve user emotional responses are incorporated into the interface. Furthermore, evaluation tests are conducted using the new interface design to confirm its effectiveness.
[0761] As a concrete example, a prompt message might be input into the generative AI model in the form of, "Based on the emotions the user expressed regarding a specific issue, suggest which parts should be improved and how." This aims to improve the quality of the user experience.
[0762] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0763] Step 1:
[0764] The device collects user operation data, video, and audio data in real time. Input here is video and audio from the camera and microphone connected to the device. From this, information related to emotions, such as the specific content of operations, facial expressions, and tone of voice, is obtained. For example, if a user is watching a video on a smartphone app, their facial expressions and the tempo of their operations are recorded. This data is then sent to a server as output.
[0765] Step 2:
[0766] The server receives data sent from the terminal and performs preprocessing. The input is the raw data sent by the terminal. Specifically, it performs noise reduction and converts audio and video data into a format suitable for emotion recognition. For example, it removes noise from audio data and smooths facial expression data frame by frame. As output, the preprocessed data is passed to the emotion engine.
[0767] Step 3:
[0768] The server uses pre-processed data to perform analysis with its emotion engine. The input is data that has been de-noised and transformed. The emotion engine analyzes this data to identify emotional states such as "joy," "anxiety," and "frustration." For example, it estimates what emotions the user is experiencing by analyzing subtle facial muscle movements and vocal intonation patterns. The analysis results are output as emotion data, and the process proceeds to the next step.
[0769] Step 4:
[0770] The server inputs emotional and operational data into a multimodal generation AI model. This input consists of emotional data obtained in the previous step and initial operational data. This AI model comprehensively analyzes various data to identify areas for improvement in the user experience. For example, it can be used to identify the cause of discomfort a user experiences with a particular interface. The analysis results are output as suggestions for improving the generated user experience.
[0771] Step 5:
[0772] The server automatically generates a new user interface design based on the generated improvement suggestions. The input is the improvement suggestions output by the AI model. In this step, the user interface is designed with emotionally sensitive elements in mind, for example, generating a design that enhances guidance and help in stressful situations. The output is a new user interface, which then proceeds to evaluation testing.
[0773] Step 6:
[0774] The terminal conducts evaluation tests on users using the generated user interface. The input is the newly designed interface design. Users interact with this interface and provide further data through their experience. Specifically, feedback is obtained regarding the usability and effectiveness of the interface design. The output is the operation data and feedback, which are used for further analysis on the server.
[0775] Step 7:
[0776] The server analyzes data obtained from evaluation tests to identify the optimal user interface. The input is the data obtained as a result of the evaluation tests. This data is analyzed in detail to confirm whether the initial improvement proposals were actually effective, and feedback for redesign is provided as needed. For example, in a learning app, user stress metrics are monitored to evaluate interface improvements. Ultimately, the optimal design is determined.
[0777] (Application Example 2)
[0778] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0779] Current interface designs often fail to consider the user's emotions, making it difficult to provide an optimal user interface based on the emotional state of individual users. As a result, stress and frustration accumulate among users when using a site, leading to a decline in the overall quality of the experience.
[0780] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0781] In this invention, the server includes means for collecting user behavior data, means for acquiring visual and auditory data, and means for analyzing the acquired data and recognizing the user's emotional state, equipped with an emotion engine. This makes it possible to provide an individually optimized user interface based on the user's emotional state and improve the quality of the user experience.
[0782] "User behavior data" refers to information about input operations and actions that occur when a user interacts with the system.
[0783] "Visual and auditory data" refers to data related to the user's facial expressions and voice, acquired through cameras and microphones.
[0784] The term "emotion engine" refers to a function that analyzes the user's emotional state based on acquired data and recognizes emotions such as joy, anger, sadness, and happiness.
[0785] A "multimodal generation AI model" refers to an artificial intelligence model that integrates and analyzes multiple different types of data, generating suggestions for improving the user experience while considering user behavior and emotions.
[0786] "User interface" refers to the design elements and layouts that allow users to interact with a computer system.
[0787] An "evaluation test" refers to a test conducted to check the effectiveness of the generated user interface and to consider the need for improvement.
[0788] "Text data" refers to information obtained by converting audio or other data formats into written form.
[0789] The system for implementing this invention mainly consists of a user terminal and a server. The terminal collects user operation data, visual data, and auditory data in real time. Specifically, the terminal's camera captures facial expressions and its microphone captures voice, and this data includes basic information about the user's interactions. This data also serves as foundational data for indicating the user's emotional state.
[0790] The server receives data transmitted from the terminal and performs preprocessing such as noise reduction. The emotion engine then analyzes this data to recognize the user's emotional state. This information is fed into a multimodal generation AI model for detailed analysis regarding improvements to the user experience. This AI model generates suggestions for user interfaces that take the emotional state into account.
[0791] Furthermore, the generated interface proposals are presented to the user as different user interfaces. The device then collects operational and emotional data from the user as they use the new interface, and obtains feedback. This feedback data is analyzed on the server to identify the optimal interface design, ultimately leading to improved user experience (UX).
[0792] For example, if a user on an e-commerce site shows an anxious expression on the product purchase page, the system can detect that emotion and adjust the interface to emphasize review and question features to reassure the user, or even suggest a new interface.
[0793] Examples of prompts for a generative AI model:
[0794] "If a user is showing signs of anxiety, how can we incorporate elements into the layout to provide reassurance?"
[0795] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0796] Step 1:
[0797] The device collects user operation data, visual data, and auditory data in real time. This includes facial expression data captured by the smartphone's camera and audio data recorded by the microphone. This collected data is sent to a server. The input is the user's operations and actions, and the output is the collected raw data.
[0798] Step 2:
[0799] The server receives raw data sent from the terminal and performs preprocessing such as noise reduction and data normalization. This data processing generates clean data that makes it easier for the emotion engine to perform accurate sentiment analysis.
[0800] Step 3:
[0801] The server inputs pre-processed data into the emotion engine and analyzes the user's emotional state. For example, major emotions such as joy and anxiety are identified, and these recognition results are obtained. The input data is pre-processed data, and the output is the identified emotional state.
[0802] Step 4:
[0803] The server integrates various types of information, including recognized emotion data, into a multimodal generation AI model. This AI model generates UX improvement suggestions that take each user's emotions into account. The input is various types of data, including emotion data, and the output is the generated UX improvement suggestions.
[0804] Step 5:
[0805] The server automatically generates a new user interface optimized for the user based on the generated UX improvement suggestions. This designs an interface that enhances the user experience. The input is the UX improvement suggestions, and the output is the new interface design.
[0806] Step 6:
[0807] The terminal provides the user with a new user interface and collects data and emotional responses from the user as they interact with it. This interaction data and feedback data are sent to the server as part of an evaluation test. The input is the user's interaction and feedback, and the output is the evaluation test results.
[0808] Step 7:
[0809] The server analyzes the evaluation test results and identifies the optimal interface design. Based on this analysis, if further UX improvements are needed, the design is refined. The input is the evaluation test results, and the output is the final interface design.
[0810] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0811] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0812] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0813] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0814] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0815] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0816] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0817] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0818] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0819] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0820] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0821] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0822] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0823] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0824] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0825] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0826] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0827] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0828] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0829] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0830] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0831] The following is further disclosed regarding the embodiments described above.
[0832] (Claim 1)
[0833] Means for collecting user operation data,
[0834] Means for acquiring video and audio data,
[0835] A means for analyzing the acquired data using a multimodal generation AI model and automatically generating suggestions for improving the user experience,
[0836] A means for automatically generating different user interfaces based on multiple improvement proposals that have been generated,
[0837] A means for conducting evaluation tests targeting users using the aforementioned different user interfaces and analyzing the results,
[0838] A system that includes this.
[0839] (Claim 2)
[0840] The system according to claim 1, comprising means for observing the user's operating screen and facial expressions to identify points of impairment in the user experience.
[0841] (Claim 3)
[0842] The system according to claim 1, comprising means for converting user voice feedback into text data and performing analysis.
[0843] "Example 1"
[0844] (Claim 1)
[0845] A means used by an information processing device to collect information about user operations,
[0846] Means for acquiring image and sound data,
[0847] A means of automatically generating improvement suggestions regarding the user experience by analyzing acquired data using a multi-modal generation artificial intelligence model,
[0848] A means for automatically generating different user interfaces based on multiple improvement proposals that have been generated,
[0849] A means of conducting evaluation tests targeting users using different user interfaces and analyzing the results,
[0850] Information systems including
[0851] (Claim 2)
[0852] The information system according to claim 1, further comprising means for observing the user's operation screen and facial expressions to identify points of impairment in the user experience.
[0853] (Claim 3)
[0854] The information system according to claim 1, comprising means for converting user feedback regarding sound into text data and performing analysis.
[0855] "Application Example 1"
[0856] (Claim 1)
[0857] A device for collecting user operation information,
[0858] A device for acquiring video and audio information,
[0859] A device that uses a multimodal generation AI to analyze the acquired information and automatically generate suggestions for improving the user experience,
[0860] A device that automatically generates different user information displays based on multiple proposed improvements,
[0861] An apparatus for conducting evaluation experiments targeting users using the aforementioned different user information displays and analyzing the results,
[0862] A device that evaluates users in real time based on facial and voice information acquired from smart devices,
[0863] A system that includes this.
[0864] (Claim 2)
[0865] The system according to claim 1, comprising means for acquiring user facial expression information and voice information in real time and for immediately evaluating and suggesting user experience.
[0866] (Claim 3)
[0867] The system according to claim 1, comprising means for converting user voice feedback into text information and performing analysis.
[0868] "Example 2 of combining an emotion engine"
[0869] (Claim 1)
[0870] A means for collecting user behavior information and emotional information in real time,
[0871] A means of applying preprocessing such as noise reduction and data transformation to the collected data,
[0872] A means of recognizing emotional states using an emotional engine,
[0873] A means of automatically generating suggestions for improving the user experience by analyzing emotion and behavioral information using a multimodal generation AI model,
[0874] A means for automatically generating a user interface that includes elements to improve the user's emotional response based on the generated improvement proposals,
[0875] A means to conduct evaluation tests using the generated user interface and analyze the results in detail,
[0876] A system that includes this.
[0877] (Claim 2)
[0878] The system according to claim 1, comprising means for observing the user's actions and emotions to identify points of disruption in the user experience.
[0879] (Claim 3)
[0880] The system according to claim 1, comprising means for converting user voice feedback into text information and performing analysis.
[0881] "Application example 2 when combining with an emotional engine"
[0882] (Claim 1)
[0883] Means for collecting user behavior data,
[0884] Means for acquiring visual and auditory data,
[0885] A means for having an emotion engine, analyzing the acquired data, and recognizing the user's emotional state,
[0886] A means of automatically generating suggestions for improving the user experience along with recognized emotion data using a multimodal generation AI model,
[0887] A means for automatically generating different user interface screens based on multiple proposed improvements,
[0888] A means for conducting evaluation tests targeting users using the aforementioned different user operation screens and analyzing the results,
[0889] A system that includes this.
[0890] (Claim 2)
[0891] The system according to claim 1, comprising means for observing the emotional expressions of users and identifying points of impairment in the user experience.
[0892] (Claim 3)
[0893] The system according to claim 1, comprising means for converting the user's auditory feedback into text data and performing analysis. [Explanation of symbols]
[0894] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means for collecting user operation data, Means for acquiring video and audio data, A means for analyzing the acquired data using a multimodal generation AI model and automatically generating suggestions for improving the user experience, A means for automatically generating different user interfaces based on multiple improvement proposals that have been generated, A means for conducting evaluation tests targeting users using the aforementioned different user interfaces and analyzing the results, A system that includes this.
2. The system according to claim 1, further comprising means for observing the user's operating screen and facial expressions to identify points of impairment in the user experience.
3. The system according to claim 1, comprising means for converting user voice feedback into text data and performing analysis.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A