system
The system addresses the limitations of one-to-many communication on social platforms by generating a personality-mimicking electronic proxy for efficient, adaptive, and emotionally sensitive interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Existing communication systems on social networking platforms restrict individual flexibility and require significant time investment due to one-to-many conversation formats, limiting personal expression and interaction efficiency.
A system that generates an electronic proxy reflecting a user's personality by analyzing voice and video data, using natural language processing to mimic their characteristics and engage in flexible, autonomous communication, while recording and refining conversation history for improved accuracy.
Enables efficient, personalized communication that maintains individuality and enhances interaction quality on social media by reducing time burden and adapting to user personality and emotional changes.
Smart Images

Figure 2026074878000001 
Figure 2026074878000002 
Figure 2026074878000003
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] To solve the problem that when a user communicates with other users on an SNS platform, it takes a lot of time and individual flexible communication is restricted by being limited to a one-to-many conversation format.
Means for Solving the Problems
[0005] Provided is a system that generates an electronic proxy reflecting the personality of a user and automatically uses it to communicate with other users. Specifically, it collects and analyzes the user's voice and video data, constructs an electronic proxy similar to the user, and analyzes the user's personality using natural language processing to achieve individual and flexible communication.
[0006] A "personality-reflecting electronic proxy" is an electronic agent designed to mimic a user's characteristics and expressive style, and to communicate on their behalf.
[0007] "Means for collecting and analyzing user data" refers to a process or system that acquires information provided by users, such as voice data or text data, and analyzes it to understand the user's personality.
[0008] "Means of managing communication" refer to mechanisms and processes that coordinate and facilitate conversations between multiple electronic proxies.
[0009] "Means for saving conversation history" refers to a system or method for accumulating records of communications conducted by electronic agents so that they can be referenced later.
[0010] "Means of constructing based on audio and video data" refers to technologies and processes for creating electronic surrogates based on a user's audio and video information.
[0011] "Methods of analysis using natural language processing" refer to methods of analyzing a user's personality and characteristics using technologies for understanding, interpreting, and generating language data with a computer. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5]It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the language used in the following description will be explained.
[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface that includes a communication processor, an antenna, and the like. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] This invention relates to a system for generating a proxy that electronically mimics a user's personality and automatically communicates with other users on social networking services (SNS). Specific embodiments are described below.
[0034] First, users connect their social media accounts to the platform and provide data that reflects their personality. This includes text-based conversation logs, audio clips, and video clips. This data forms the basis for capturing user characteristics.
[0035] Next, the terminal formats the received data and sends it to the server. Upon receiving this data, the server uses natural language processing techniques to analyze the user's conversation patterns and unique writing style. Simultaneously, it uses speech synthesis technology to extract features of the user's voice and build a profile to mimic the user's voice.
[0036] Furthermore, video analysis technology is used to generate avatars that reflect non-verbal elements such as the user's gestures and facial expressions. These avatars visually represent the user's personality and are used when the electronic proxy engages in conversation.
[0037] The generated electronic proxy initiates a conversation with other users on the SNS platform via a server. During this process, the electronic proxy generates responses based on the user's previously learned conversational style. While the conversation content is generated in real time, the generation process incorporates individual contextual understanding and adaptive capabilities.
[0038] For example, if a user frequently uses the characteristic daily greeting "Good morning!", the electronic proxy will also use the same greeting at the same time. In addition, the electronic proxy records conversation logs and further refines the user's personality based on the content and frequency of conversations. This allows the proxy to more accurately reflect the user's evolving communication style over time.
[0039] In this way, users can save their time while maintaining their individuality and continuing to have rich interactions with other users on social media. This system makes it possible to further enhance the communication experience that the platform provides.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] Users upload necessary information to the platform, such as their profile information, past conversation history, audio data, and video data.
[0043] Step 2:
[0044] The terminal receives this user data, converts it to the required format, and then securely and efficiently transmits the data to the server.
[0045] Step 3:
[0046] The server analyzes the user data it receives and uses natural language processing techniques to extract characteristic expressions and frequently occurring phrases from the conversation logs.
[0047] Step 4:
[0048] The server uses speech synthesis technology based on the audio data to analyze the characteristics of the user's voice, generate a voice profile, and prepare to reproduce the user's voice.
[0049] Step 5:
[0050] The server analyzes the video data and generates an avatar that reflects the user's gestures and facial expressions. This avatar then acts as an electronic proxy.
[0051] Step 6:
[0052] Based on these analysis results, the server builds individually customized electronic proxies that reflect the user's personality and style.
[0053] Step 7:
[0054] Users choose which users they communicate with on the social networking platform.
[0055] Step 8:
[0056] The server sets up a conversation session with the selected electronic proxy and initiates a dialogue between the electronic proxies.
[0057] Step 9:
[0058] The server generates the conversation in real time, adjusting the response at each step to ensure that the meaning and flow of the conversation are natural.
[0059] Step 10:
[0060] The server automatically saves the conversation history and manages logs so that users can review the content later.
[0061] Step 11:
[0062] Throughout this process, users occasionally provide feedback, which is used to further customize the electronic proxy and improve its accuracy.
[0063] (Example 1)
[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0065] In modern society, while communication on social media is active, a challenge remains in that users find it difficult to maintain their individuality and continue interacting while reducing the time burden. In particular, there is a need for proxy systems that can reflect a user's unique writing style, expression, and non-verbal characteristics such as voice and gestures.
[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0067] In this invention, the server includes means for generating an electronic surrogate that mimics the user's characteristics, means for collecting and analyzing the user's conversation history, voice data, and video data, means for autonomously communicating with other users through the electronic surrogate, and means for recording the generated communication data and learning the user's personality. As a result, the user can communicate smoothly on social networking services while reducing the time burden by using a surrogate that effectively reflects their own personality.
[0068] A "user" is an individual or group that utilizes the system and provides data that reflects their own personality.
[0069] "Characteristics" refer to elements that represent a user's individuality, such as their unique conversation patterns, writing style, tone of voice, and gestures.
[0070] An "electronic proxy" is a virtual entity that mimics the characteristics of a user and engages in autonomous communication with other users on social networking services (SNS).
[0071] "Conversation history" refers to a record of the communication a user has had so far, including text-based logs.
[0072] "Audio data" refers to digital audio files that contain information about the user's voice, and are data that possesses characteristics such as timbre and pitch.
[0073] "Video data" refers to video files that include the user's gestures and facial expressions, and is digital data that represents non-verbal elements.
[0074] "Analyzing" refers to a method of extracting user characteristics and behavioral patterns by examining data in detail.
[0075] "Autonomously" refers to a system that performs functions independently without requiring explicit user intervention.
[0076] This invention is a system that generates an electronic proxy that accurately mimics the user's characteristics and autonomously communicates with other users on social networking services (SNS).
[0077] Data collection and analysis
[0078] Users connect their social media accounts to the system and provide text-based conversation history, audio data, and video data. This forms a foundation for reflecting the user's unique conversation patterns and nonverbal characteristics.
[0079] The terminal converts the data received from the user into a standard format and sends it to the server. At this stage, technologies such as BERT and GPT are used as natural language processing engines to analyze the conversation data. A speech synthesis library is used for feature analysis of the audio data. For video data, for example, OpenCV is used to perform facial recognition and gesture extraction to capture the user's nonverbal characteristics.
[0080] Generation of electronic proxy
[0081] The server generates an electronic avatar that mimics the user's voice and video characteristics based on the analyzed data. It visually constructs an avatar, including the user's facial expressions and gestures, using 3D modeling. It then forms a speech synthesis profile and generates a voice similar to the user's.
[0082] Autonomous communication
[0083] The server enables autonomous communication on social networking services (SNS) through generated electronic proxies. In this process, a generative AI model is used to provide real-time responses. A concrete example of a prompt is, "Generate a response based on a user's frequent conversation about the weather." This prompt allows the electronic proxies to form natural responses such as, "It's sunny today."
[0084] This system allows users to maintain their individuality while continuing online interactions, reducing the time burden involved.
[0085] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0086] Step 1:
[0087] Users connect their social media accounts to the system and provide past conversation history, audio data, and video data as input data. This data is accumulated as a foundation to reflect the user's personality.
[0088] Step 2:
[0089] The terminal processes the input data received from the user into a standard format (e.g., JSON) and sends it to the server. Audio data is converted to an appropriate sample rate, and video data is converted to an appropriate resolution and frame rate. This creates formatted data.
[0090] Step 3:
[0091] The server analyzes the formatted data it receives. First, it uses a natural language processing engine (such as BERT or GPT) to identify the user's conversation patterns and writing style from the conversation history. Next, it performs acoustic analysis to extract voice features (tone and pitch) from the audio data. Furthermore, it uses technologies such as OpenCV to recognize gestures and facial expressions from the video data. This generates a user personality profile.
[0092] Step 4:
[0093] The server generates an electronic surrogate based on the analysis results. Specifically, it uses 3D modeling and speech synthesis to construct an electronic avatar that mimics the user's voice and video characteristics. This is how the electronic surrogate is generated.
[0094] Step 5:
[0095] The server utilizes a generative AI model to communicate with other users on the social networking service (SNS) through an electronic proxy. Specifically, it generates real-time conversations using prompts such as "Generate an appropriate response when the user talks about the weather." Based on the input conversation content, the electronic proxy returns a natural response. In this way, communication on the SNS is managed autonomously.
[0096] Step 6:
[0097] The server records the communications made by the electronic proxy and their outcomes. This record is analyzed and used to further refine the user's personality. This improves the accuracy of the electronic proxy and its reflection of the user's characteristics over time.
[0098] (Application Example 1)
[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0100] In today's digital content world, there is a lack of systems that allow for broad interaction while preserving individual personality. In particular, there is a need for methods that enable content creators to interact with many fans simultaneously, efficiently managing large volumes of communication while maintaining their own unique style.
[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0102] In this invention, the server includes information processing means for generating an electronic surrogate that mimics personality, computation means for collecting and analyzing user records, and means for displaying a virtual character via a smart device. This enables content creators to maintain their individuality while efficiently engaging in real-time, interactive communication with a wide range of fans.
[0103] An "electronic surrogate that mimics personality" is a virtual character created using information processing technology to reproduce the user's personality and expression.
[0104] "Information processing means" refers to a set of hardware and software for collecting, analyzing, and processing digital data.
[0105] "User records" refers to all data that represents the user's behavior and characteristics, including audio, video, and text.
[0106] A "computational means" is a system that has the function of performing calculations and analyses according to a specific purpose based on user records.
[0107] "Communication methods" refer to technologies that enable the transmission and reception of data, and facilitate the exchange of information between devices.
[0108] "Storage means" refers to devices and media used to retain digital data for extended periods.
[0109] A "smart device" is a general term for electronic devices that have internet connectivity and application execution capabilities.
[0110] A "virtual character" is a visually represented surrogate being created using digital technology.
[0111] A system for carrying out this invention has a configuration that generates an electronic surrogate that mimics the personality of a content creator and displays it as a virtual character via a smart device.
[0112] The server first collects and analyzes user records, i.e., audio, video, and text data, using information processing tools. For analysis, it uses natural language processing (NLP) libraries such as NLTK and spaCy, and for voice feature extraction, it uses speech synthesis engines such as Google Cloud Text-to-Speech. This allows the server to understand the user's characteristics and conversational style and generate a virtual character appropriate to their personality.
[0113] The generated virtual character is visualized using avatar generation tools such as Unity3D. This virtual character is displayed in real time via smart devices, specifically smart glasses, and interacts with fans and other users.
[0114] As a concrete example, consider a scenario where a content creator holds a virtual event with fans through smart glasses. In this event, the virtual character would mimic phrases and greetings frequently used by the creator in the past, and generate responses to fans' questions that reflect the creator's style.
[0115] An example of a prompt to input into the generation AI model would be, "Generate a response that greets the audience warmly and in a friendly manner, while maintaining the humorous tone that is characteristic of the creator."
[0116] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0117] Step 1:
[0118] The user inputs their own voice, video, and text data into the device. This data serves as foundational information for mimicking the user's personality. The device collects this data, formats it, and sends it to the server. The input is raw audio and video, and the output is formalized data.
[0119] Step 2:
[0120] The server analyzes the received user data using natural language processing tools. Specifically, it extracts voice characteristics and trains an AI model to generate conversational styles that reflect the creator's individuality. In this process, the user's past conversation logs are used as input, and a profile modeling the user's conversational patterns is obtained as output.
[0121] Step 3:
[0122] The server uses a speech synthesis engine to generate a voice profile to mimic the user's voice. The input is the user's voice data, and the output is the mimicked voice profile. This allows the virtual character to communicate using the user's voice.
[0123] Step 4:
[0124] The server uses an avatar generation tool based on the user's video data to create a virtual character that reflects the user's appearance and gestures. The input is the analyzed video data, and the output is a 3D modeled virtual character.
[0125] Step 5:
[0126] The generated virtual character is displayed through smart glasses on the device. This allows the user to use the virtual character as their proxy and interact with other users in real time. The input is the data of the virtual character, and the output is the visual interface displayed on the smart device.
[0127] Step 6:
[0128] The user's smart glasses control the actions of a virtual character and perform prompt-based conversations on social media and communication platforms on the user's behalf. Here, responses generated by a generative AI model are input and displayed as output in a format that allows for real-time interaction with viewers. An example of a prompt to input to the generative AI model is, "Generate a response that greets the viewers warmly and in a friendly manner, while maintaining the humorous tone that is characteristic of the creator."
[0129] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0130] This invention relates to a system for recognizing a user's emotions and for an electronic surrogate to appropriately adjust communication based on those emotions. The invention generates an electronic surrogate that reflects the user's personality and incorporates an emotion engine to achieve more natural and adaptive communication.
[0131] First, users register on the platform and provide profile information along with audio and video data. This data includes typical conversational scenes and situations involving specific emotions. This data serves as foundational information for accurately understanding the user's emotional state.
[0132] Next, the device converts the provided data into the appropriate format and sends it to the server. The server receives this data and begins the process of analyzing the user's emotional state using an emotion engine. The emotion engine identifies the user's emotional state by analyzing the tone, pitch, and speed of the voice from the audio data, and by analyzing changes in facial expressions from the video data.
[0133] Subsequently, the server constructs an electronic surrogate that reflects the user's personality based on the analysis results. This electronic surrogate lists appropriate responses and attitudes to specific emotions and uses them in various communication situations. This process incorporates dynamic adjustment mechanisms to quickly respond to potential emotional shifts.
[0134] For example, if a user uses an emotion engine to recognize that they are "happy" in a particular conversation, the electronic proxy can respond with appropriate words and facial expressions to express that joy. Conversely, if "anger" or "sadness" is recognized, the electronic proxy will show a corresponding attitude of comfort and empathy.
[0135] Furthermore, the server records the conversation history of the electronic proxy, allowing the user to review the history later and provide feedback. This enables the electronic proxy's responses to become more sophisticated and adapt to changes in the user.
[0136] Through the embodiments described above, users can obtain a more personalized experience and embody communication that resonates with their emotions through AI. This invention enables communication based on the user's unique emotions and elevates interactions on social networking platforms to a new dimension.
[0137] The following describes the processing flow.
[0138] Step 1:
[0139] Users log in to their accounts and upload their own audio and video data to the platform. This data includes scenes that evoke a variety of emotions.
[0140] Step 2:
[0141] The device converts the data received from the user into an appropriate format, compresses and encrypts the data, and then sends it to the server.
[0142] Step 3:
[0143] To analyze the user's emotional state, the server extracts voice tone and pitch from received audio data and detects subtle changes in facial expressions from video data. The emotion engine then generates an emotion label based on this information.
[0144] Step 4:
[0145] The server uses the analyzed emotion data to build an electronic proxy profile that corresponds to the user's emotional response. This allows the electronic proxy to generate responses optimized for the user's emotions.
[0146] Step 5:
[0147] Users select someone they want to communicate with on the SNS platform and start a conversation session.
[0148] Step 6:
[0149] The server automatically sets an appropriate conversation style and response for the electronic proxy based on the user's emotions, and initiates a conversation with the other party's electronic proxy.
[0150] Step 7:
[0151] During the conversation, the server adjusts the electronic proxy's responses in real time and dynamically applies feedback based on the latest emotional state.
[0152] Step 8:
[0153] The server records the conversation history and changes in emotions, and stores the analysis results in a database. Users can review this log later.
[0154] Step 9:
[0155] Later, users will review the provided conversation logs, provide feedback, and improve the electronic proxy response tactics to build a more user-friendly communication experience.
[0156] (Example 2)
[0157] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0158] In today's communication environment, there is a lack of response that responds to users' emotions, making natural communication that reflects individual feelings and situations difficult. Furthermore, flexibility to respond immediately to changes in emotions is also required, but conventional technologies do not adequately address this. Therefore, there is a need to provide adaptive communication methods based on the individuality and emotions of users.
[0159] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0160] In this invention, the server includes means for analyzing the user's emotional state, means for generating an electronic surrogate that reflects the user's personality, and means for incorporating a dynamic adjustment function. This enables personalized responses that are in line with the user's emotions.
[0161] A "user" is an entity that registers with the system and provides audio and video information.
[0162] "Emotional state" refers to the user's mental condition as identified through the analysis of audio and video.
[0163] An "electronic proxy" is a virtual entity created based on personality and emotions, used for communication.
[0164] "Communication" refers to the process of exchanging information with other users through electronic agents.
[0165] "Dynamic adjustment function" refers to a function that automatically optimizes the electronic proxy's response in accordance with changes in emotions and situations.
[0166] A "generative AI model" is an artificial intelligence system that creates electronic surrogates based on the analysis of emotions and personality.
[0167] A "prompt" refers to a text-based input used to give instructions to a generative AI model.
[0168] To implement this invention, the user must first register with the system and provide audio and video data. The user uses a smartphone or computer to record audio using a microphone and transmits the audio data to the system. They also record video of their face using a camera and upload it as video data.
[0169] The terminal converts the audio and video data received from the user into the appropriate format. This process converts audio data to WAV or MP3 format and video data to MP4 format. This ensures the data is suitable for analysis on the server.
[0170] The server receives the converted data and performs analysis using an emotion engine. The emotion engine identifies the user's emotional state by analyzing the tone, pitch, speed, and facial expressions of their voice. For example, if the voice is high-pitched and fast-paced, it is determined that the user is likely excited.
[0171] Based on the analysis results, the server uses a generative AI model to generate an electronic proxy that reflects the user's personality. The electronic proxy lists response patterns to appropriately respond to the user's emotions and uses them during conversations. Dynamic adjustment functions are also incorporated, allowing for immediate response to real-time changes in emotions.
[0172] For example, if a user tells the system, "I have some great news today!", the electronic agent will understand the user's emotion and reply, "That's wonderful! Congratulations!" In this way, communication that is sensitive to the user's feelings is realized.
[0173] An example of a prompt might be, "List the responses from the user's most recent conversation history that elicited the most positive emotions." Based on such prompts, the generative AI model can continuously learn and become capable of performing increasingly accurate dialogues.
[0174] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0175] Step 1:
[0176] Users log in to the system and input audio and video data. They use the microphone and camera on their smartphone or computer to record audio and video, including everyday conversations and specific emotional states, and upload them to the system. The input data consists of raw audio and video files.
[0177] Step 2:
[0178] The terminal converts the raw data received from the user into the appropriate format. Audio data is converted to WAV or MP3 format, and video data is converted to MP4 format. This format conversion makes the data ready for processing on the server. The input is raw data, and the output is data converted to the appropriate format.
[0179] Step 3:
[0180] The server inputs formatted data received from the terminal into the emotion engine for analysis. Here, the user's emotional state is analyzed based on data such as voice tone, pitch, speed, and facial expression changes obtained from the video. For example, voice pitch and speed are considered indicators of emotions such as joy or excitement. The input is formatted audio and video data, and the output is the analyzed emotional state.
[0181] Step 4:
[0182] The server uses a generative AI model to construct an electronic proxy that reflects the user's personality based on the analyzed emotional state. The electronic proxy prepares responses and attitudes that match that emotion and lists them. The generative AI model outputs an appropriate dialogue format that corresponds to the user's emotions and personality based on a given prompt sentence. For example, the prompt might be, "Generate a dialogue format that expresses the user's joy." The input is the emotional state and the prompt sentence, and the output is a response pattern generated based on the emotion.
[0183] Step 5:
[0184] The server interacts with the user using a generated electronic proxy. Here, it utilizes real-time dynamic adjustment capabilities to respond immediately to changes in the user's emotions. For example, if the user is suddenly surprised, the electronic proxy will instantly provide an appropriate response. The input is the user's current emotional state, and the output is the dynamically adjusted response.
[0185] Step 6:
[0186] The server records a history of the interactions that take place. This allows users to review past responses and provide feedback to help improve the system. This interaction history includes the user's emotions and the responses of the electronic proxy. The input is the data of the interactions that took place, and the output is the saved interaction history.
[0187] (Application Example 2)
[0188] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0189] Traditionally, customer service in commercial facilities has been uniform, making it difficult to provide appropriate suggestions and services tailored to the emotional state and preferences of individual customers. Furthermore, there has been a lack of means to recommend the most suitable products and services to meet the diverse needs of customers.
[0190] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0191] In this invention, the server includes means for generating an electronic surrogate that reflects the individual's personality, means for collecting and analyzing user information, and means for analyzing emotions and suggesting products based on that analysis. This makes it possible to provide appropriate product suggestions and services tailored to each customer's emotional state.
[0192] An "electronic proxy that reflects individuality" is a digital entity that takes into account the user's personal characteristics and preferences and communicates on their behalf.
[0193] "User information" refers to audio data, video data, and other related data provided by the user, and serves as the basis for analyzing the user's emotions and characteristics.
[0194] "Means for collecting and analyzing data" refers to functions that receive and analyze data obtained from users in an appropriate format to identify the user's emotional state and characteristics.
[0195] "Means of managing interactions with other users using electronic proxies" refers to functions that perform communication and optimize interactions on behalf of the user.
[0196] "Means for retaining generated communication history" refers to a function for saving records of interactions with the user and keeping them in a state where they can be referenced later.
[0197] "A means of analyzing emotions and proposing products based on them" refers to a function that analyzes the user's emotional state and provides appropriate products or services based on the results.
[0198] This invention aims to build a system for analyzing customer emotions in physical stores and providing optimal product recommendations. Specific embodiments of the present invention are shown below.
[0199] Users first access the platform using smartphones or tablet devices installed within the store. These devices are equipped with voice input capabilities and cameras, making it possible to collect the user's voice and image data.
[0200] The device converts the collected data into an appropriate format and sends it to the server over the network. The server receives each user's data and analyzes it using a speech recognition API (e.g., Google Cloud Speech-to-Text API) and a facial recognition API (e.g., Microsoft® Azure® Face API). This analysis identifies the user's emotional state.
[0201] Based on the analysis results, the server generates an electronic proxy and recommends products and services that match the user's emotions. This process utilizes an emotion engine, which analyzes the user's facial expressions and vocal characteristics.
[0202] For example, if a user is observed looking at products in a store and appears to be deep in thought, their facial expression can be interpreted as "undecided." In this case, the server will display recommended product reviews and suggestions tailored to that user on their device, providing a more engaging shopping experience.
[0203] A concrete example of a prompt might be, "This user does not have a smartphone, but suggest a way to estimate his emotions from his facial expressions and generate a customer service profile." By inputting this prompt into the AI model, the data analysis and product suggestion process is automated.
[0204] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0205] Step 1:
[0206] The device collects audio and image data provided by the user. This collected data is processed in real time within the device and converted into a format that can be analyzed on the server. The input is raw audio and image data, and the output is formatted digital data.
[0207] Step 2:
[0208] The converted digital data is transmitted from the terminal to the server via the network. Specifically, this involves generating data packets and transferring them to the server using the appropriate communication protocol. The input is the formatted digital data, and the output is the transmission of data to the server.
[0209] Step 3:
[0210] The server inputs the received data into the speech recognition API and facial expression recognition API, and analyzes the user's voice tone and facial expressions. Data processing involves extracting features from the voice tone and detecting facial expression features from the image. The input is digital data sent to the server, and the output is the analyzed emotional state data.
[0211] Step 4:
[0212] Based on the analysis results, the server uses a generative AI model to generate product and service suggestions that are appropriate for the user's emotions. A prompt is input to the generative AI model, and the output is the suggested content. This process specifically involves sending a prompt to the generative AI model and receiving the result. The input is the analyzed emotional state data, and the output is the suggested content.
[0213] Step 5:
[0214] The server sends the generated suggestions to the terminal and displays them to the user. This allows the user to check the suggested products and services in the store. The specific operations include data transmission and screen display. The input is the generated suggestions, and the output is information that the user can visually confirm.
[0215] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0216] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0217] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0218] [Second Embodiment]
[0219] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0220] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0221] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0222] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0223] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0224] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0225] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0226] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0227] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0228] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0229] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0230] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0231] This invention relates to a system for generating a proxy that electronically mimics a user's personality and automatically communicates with other users on social networking services (SNS). Specific embodiments are described below.
[0232] First, users connect their social media accounts to the platform and provide data that reflects their personality. This includes text-based conversation logs, audio clips, and video clips. This data forms the basis for capturing user characteristics.
[0233] Next, the terminal formats the received data and sends it to the server. Upon receiving this data, the server uses natural language processing techniques to analyze the user's conversation patterns and unique writing style. Simultaneously, it uses speech synthesis technology to extract features of the user's voice and build a profile to mimic the user's voice.
[0234] Furthermore, video analysis technology is used to generate avatars that reflect non-verbal elements such as the user's gestures and facial expressions. These avatars visually represent the user's personality and are used when the electronic proxy engages in conversation.
[0235] The generated electronic proxy initiates a conversation with other users on the SNS platform via a server. During this process, the electronic proxy generates responses based on the user's previously learned conversational style. While the conversation content is generated in real time, the generation process incorporates individual contextual understanding and adaptive capabilities.
[0236] For example, if a user frequently uses the characteristic daily greeting "Good morning!", the electronic proxy will also use the same greeting at the same time. In addition, the electronic proxy records conversation logs and further refines the user's personality based on the content and frequency of conversations. This allows the proxy to more accurately reflect the user's evolving communication style over time.
[0237] In this way, users can save their time while maintaining their individuality and continuing to have rich interactions with other users on social media. This system makes it possible to further enhance the communication experience that the platform provides.
[0238] The following describes the processing flow.
[0239] Step 1:
[0240] Users upload necessary information to the platform, such as their profile information, past conversation history, audio data, and video data.
[0241] Step 2:
[0242] The terminal receives this user data, converts it to the required format, and then securely and efficiently transmits the data to the server.
[0243] Step 3:
[0244] The server analyzes the user data it receives and uses natural language processing techniques to extract characteristic expressions and frequently occurring phrases from the conversation logs.
[0245] Step 4:
[0246] The server uses speech synthesis technology based on the audio data to analyze the characteristics of the user's voice, generate a voice profile, and prepare to reproduce the user's voice.
[0247] Step 5:
[0248] The server analyzes the video data and generates an avatar that reflects the user's gestures and facial expressions. This avatar then acts as an electronic proxy.
[0249] Step 6:
[0250] Based on these analysis results, the server builds individually customized electronic proxies that reflect the user's personality and style.
[0251] Step 7:
[0252] Users choose which users they communicate with on the social networking platform.
[0253] Step 8:
[0254] The server sets up a conversation session with the selected electronic proxy and initiates a dialogue between the electronic proxies.
[0255] Step 9:
[0256] The server generates the conversation in real time, adjusting the response at each step to ensure that the meaning and flow of the conversation are natural.
[0257] Step 10:
[0258] The server automatically saves the conversation history and manages logs so that users can review the content later.
[0259] Step 11:
[0260] Throughout this process, users occasionally provide feedback, which is used to further customize the electronic proxy and improve its accuracy.
[0261] (Example 1)
[0262] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0263] In modern society, while communication on social media is active, a challenge remains in that users find it difficult to maintain their individuality and continue interacting while reducing the time burden. In particular, there is a need for proxy systems that can reflect a user's unique writing style, expression, and non-verbal characteristics such as voice and gestures.
[0264] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0265] In this invention, the server includes means for generating an electronic surrogate that mimics the user's characteristics, means for collecting and analyzing the user's conversation history, voice data, and video data, means for autonomously communicating with other users through the electronic surrogate, and means for recording the generated communication data and learning the user's personality. As a result, the user can communicate smoothly on social networking services while reducing the time burden by using a surrogate that effectively reflects their own personality.
[0266] A "user" is an individual or group that utilizes the system and provides data that reflects their own personality.
[0267] "Characteristics" refer to elements that represent a user's individuality, such as their unique conversation patterns, writing style, tone of voice, and gestures.
[0268] An "electronic proxy" is a virtual entity that mimics the characteristics of a user and engages in autonomous communication with other users on social networking services (SNS).
[0269] "Conversation history" refers to a record of the communication a user has had so far, including text-based logs.
[0270] "Audio data" refers to digital audio files that contain information about the user's voice, and are data that possesses characteristics such as timbre and pitch.
[0271] "Video data" refers to video files that include the user's gestures and facial expressions, and is digital data that represents non-verbal elements.
[0272] "Analyzing" refers to a method of extracting user characteristics and behavioral patterns by examining data in detail.
[0273] "Autonomously" refers to a system that performs functions independently without requiring explicit user intervention.
[0274] This invention is a system that generates an electronic proxy that accurately mimics the user's characteristics and autonomously communicates with other users on social networking services (SNS).
[0275] Data collection and analysis
[0276] Users connect their social media accounts to the system and provide text-based conversation history, audio data, and video data. This forms a foundation for reflecting the user's unique conversation patterns and nonverbal characteristics.
[0277] The terminal converts the data received from the user into a standard format and sends it to the server. At this stage, technologies such as BERT and GPT are used as natural language processing engines to analyze the conversation data. A speech synthesis library is used for feature analysis of the audio data. For video data, for example, OpenCV is used to perform facial recognition and gesture extraction to capture the user's nonverbal characteristics.
[0278] Generation of electronic proxy
[0279] The server generates an electronic avatar that mimics the user's voice and video characteristics based on the analyzed data. It visually constructs an avatar, including the user's facial expressions and gestures, using 3D modeling. It then forms a speech synthesis profile and generates a voice similar to the user's.
[0280] Autonomous communication
[0281] The server enables autonomous communication on social networking services (SNS) through generated electronic proxies. In this process, a generative AI model is used to provide real-time responses. A concrete example of a prompt is, "Generate a response based on a user's frequent conversation about the weather." This prompt allows the electronic proxies to form natural responses such as, "It's sunny today."
[0282] With this system, the user can maintain online communication while preserving their individuality, and it becomes possible to reduce the time burden.
[0283] The flow of the specific process in Example 1 will be described using FIG. 11.
[0284] Step 1:
[0285] The user connects their SNS account to the system and provides past conversation history, voice data, and video data as input data. This data is accumulated as a basis for reflecting the user's individuality.
[0286] Step 2:
[0287] The terminal processes the input data received from the user into a standard format (e.g., JSON format) and sends it to the server. The voice data is converted to an appropriate sample rate, and the video data is also converted to an appropriate resolution and frame rate. Thereby, the formatted data is created.
[0288] Step 3:
[0289] The server analyzes the received formatted data. First, using a natural language processing engine (e.g., BERT or GPT), the conversation pattern and style of the user are identified from the conversation history. Next, acoustic analysis is performed to extract voice features (tone and pitch) from the voice data. Furthermore, using technologies such as OpenCV, gestures and expressions are recognized from the video data. Thereby, the user's personality profile is generated.
[0290] Step 4:
[0291] The server generates an electronic proxy based on the analysis results. Specifically, using 3D modeling and voice synthesis, an electronic avatar that mimics the voice and video features of the user is constructed. Thereby, the electronic proxy is generated.
[0292] Step 5:
[0293] The server utilizes a generative AI model to communicate with other users on the social networking service (SNS) through an electronic proxy. Specifically, it generates real-time conversations using prompts such as "Generate an appropriate response when the user talks about the weather." Based on the input conversation content, the electronic proxy returns a natural response. In this way, communication on the SNS is managed autonomously.
[0294] Step 6:
[0295] The server records the communications made by the electronic proxy and their outcomes. This record is analyzed and used to further refine the user's personality. This improves the accuracy of the electronic proxy and its reflection of the user's characteristics over time.
[0296] (Application Example 1)
[0297] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0298] In today's digital content world, there is a lack of systems that allow for broad interaction while preserving individual personality. In particular, there is a need for methods that enable content creators to interact with many fans simultaneously, efficiently managing large volumes of communication while maintaining their own unique style.
[0299] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0300] In this invention, the server includes information processing means for generating an electronic agent that mimics individuality, computing means for collecting and analyzing user records, and means for displaying virtual characters via a smart device. As a result, it becomes possible for content creators to efficiently engage in real-time interactive communication with a wide range of fans while maintaining their individuality.
[0301] The "electronic agent that mimics individuality" is a virtual character generated using information processing technology to reproduce the user's personality and expression.
[0302] The "information processing means" is a set of hardware and software for collecting, analyzing, and processing digital data.
[0303] The "user records" refer to all data representing the user's actions and characteristics, such as voice, video, and text.
[0304] The "computing means" is a system with the function of performing calculations and analyses according to specific purposes based on user records.
[0305] The "communication means" is a technology for transmitting and receiving data and enabling information exchange between devices.
[0306] The "storage means" refers to a device or medium for long-term retention of digital data.
[0307] The "smart device" is a general term for electronic devices with Internet connection capabilities and application execution capabilities.
[0308] The "virtual character" is a visually represented proxy entity generated by digital technology.
[0309] The system for implementing this invention has a configuration that generates an electronic agent that mimics the individuality of a content creator and displays it as a virtual character via a smart device.
[0310] The server first collects and analyzes user records, i.e., audio, video, and text data, using information processing tools. For analysis, it uses natural language processing (NLP) libraries such as NLTK and spaCy, and for voice feature extraction, it uses speech synthesis engines such as Google Cloud Text-to-Speech. This allows the server to understand the user's characteristics and conversational style and generate a virtual character appropriate to their personality.
[0311] The generated virtual character is visualized using avatar generation tools such as Unity3D. This virtual character is displayed in real time via smart devices, specifically smart glasses, and interacts with fans and other users.
[0312] As a concrete example, consider a scenario where a content creator holds a virtual event with fans through smart glasses. In this event, the virtual character would mimic phrases and greetings frequently used by the creator in the past, and generate responses to fans' questions that reflect the creator's style.
[0313] An example of a prompt to input into the generation AI model would be, "Generate a response that greets the audience warmly and in a friendly manner, while maintaining the humorous tone that is characteristic of the creator."
[0314] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0315] Step 1:
[0316] The user inputs their own voice, video, and text data into the device. This data serves as foundational information for mimicking the user's personality. The device collects this data, formats it, and sends it to the server. The input is raw audio and video, and the output is formalized data.
[0317] Step 2:
[0318] The server analyzes the received user data using natural language processing tools. Specifically, it extracts voice characteristics and trains an AI model to generate conversational styles that reflect the creator's individuality. In this process, the user's past conversation logs are used as input, and a profile modeling the user's conversational patterns is obtained as output.
[0319] Step 3:
[0320] The server uses a speech synthesis engine to generate a voice profile to mimic the user's voice. The input is the user's voice data, and the output is the mimicked voice profile. This allows the virtual character to communicate using the user's voice.
[0321] Step 4:
[0322] The server uses an avatar generation tool based on the user's video data to create a virtual character that reflects the user's appearance and gestures. The input is the analyzed video data, and the output is a 3D modeled virtual character.
[0323] Step 5:
[0324] The generated virtual character is displayed through smart glasses on the device. This allows the user to use the virtual character as their proxy and interact with other users in real time. The input is the data of the virtual character, and the output is the visual interface displayed on the smart device.
[0325] Step 6:
[0326] The user's smart glasses control the actions of a virtual character and perform prompt-based conversations on social media and communication platforms on the user's behalf. Here, responses generated by a generative AI model are input and displayed as output in a format that allows for real-time interaction with viewers. An example of a prompt to input to the generative AI model is, "Generate a response that greets the viewers warmly and in a friendly manner, while maintaining the humorous tone that is characteristic of the creator."
[0327] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0328] This invention relates to a system for recognizing a user's emotions and for an electronic surrogate to appropriately adjust communication based on those emotions. The invention generates an electronic surrogate that reflects the user's personality and incorporates an emotion engine to achieve more natural and adaptive communication.
[0329] First, users register on the platform and provide profile information along with audio and video data. This data includes typical conversational scenes and situations involving specific emotions. This data serves as foundational information for accurately understanding the user's emotional state.
[0330] Next, the device converts the provided data into the appropriate format and sends it to the server. The server receives this data and begins the process of analyzing the user's emotional state using an emotion engine. The emotion engine identifies the user's emotional state by analyzing the tone, pitch, and speed of the voice from the audio data, and by analyzing changes in facial expressions from the video data.
[0331] Subsequently, the server constructs an electronic surrogate that reflects the user's personality based on the analysis results. This electronic surrogate lists appropriate responses and attitudes to specific emotions and uses them in various communication situations. This process incorporates dynamic adjustment mechanisms to quickly respond to potential emotional shifts.
[0332] For example, if a user uses an emotion engine to recognize that they are "happy" in a particular conversation, the electronic proxy can respond with appropriate words and facial expressions to express that joy. Conversely, if "anger" or "sadness" is recognized, the electronic proxy will show a corresponding attitude of comfort and empathy.
[0333] Furthermore, the server records the conversation history of the electronic proxy, allowing the user to review the history later and provide feedback. This enables the electronic proxy's responses to become more sophisticated and adapt to changes in the user.
[0334] Through the embodiments described above, users can obtain a more personalized experience and embody communication that resonates with their emotions through AI. This invention enables communication based on the user's unique emotions and elevates interactions on social networking platforms to a new dimension.
[0335] The following describes the processing flow.
[0336] Step 1:
[0337] Users log in to their accounts and upload their own audio and video data to the platform. This data includes scenes that evoke a variety of emotions.
[0338] Step 2:
[0339] The device converts the data received from the user into an appropriate format, compresses and encrypts the data, and then sends it to the server.
[0340] Step 3:
[0341] To analyze the user's emotional state, the server extracts voice tone and pitch from received audio data and detects subtle changes in facial expressions from video data. The emotion engine then generates an emotion label based on this information.
[0342] Step 4:
[0343] The server uses the analyzed emotion data to build an electronic proxy profile that corresponds to the user's emotional response. This allows the electronic proxy to generate responses optimized for the user's emotions.
[0344] Step 5:
[0345] Users select someone they want to communicate with on the SNS platform and start a conversation session.
[0346] Step 6:
[0347] The server automatically sets an appropriate conversation style and response for the electronic proxy based on the user's emotions, and initiates a conversation with the other party's electronic proxy.
[0348] Step 7:
[0349] During the conversation, the server adjusts the electronic proxy's responses in real time and dynamically applies feedback based on the latest emotional state.
[0350] Step 8:
[0351] The server records the conversation history and changes in emotions, and stores the analysis results in a database. Users can review this log later.
[0352] Step 9:
[0353] Later, users will review the provided conversation logs, provide feedback, and improve the electronic proxy response tactics to build a more user-friendly communication experience.
[0354] (Example 2)
[0355] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0356] In today's communication environment, there is a lack of response that responds to users' emotions, making natural communication that reflects individual feelings and situations difficult. Furthermore, flexibility to respond immediately to changes in emotions is also required, but conventional technologies do not adequately address this. Therefore, there is a need to provide adaptive communication methods based on the individuality and emotions of users.
[0357] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0358] In this invention, the server includes means for analyzing the user's emotional state, means for generating an electronic surrogate that reflects the user's personality, and means for incorporating a dynamic adjustment function. This enables personalized responses that are in line with the user's emotions.
[0359] A "user" is an entity that registers with the system and provides audio and video information.
[0360] "Emotional state" refers to the user's mental condition as identified through the analysis of audio and video.
[0361] An "electronic proxy" is a virtual entity created based on personality and emotions, used for communication.
[0362] "Communication" refers to the process of exchanging information with other users through electronic agents.
[0363] "Dynamic adjustment function" refers to a function that automatically optimizes the electronic proxy's response in accordance with changes in emotions and situations.
[0364] A "generative AI model" is an artificial intelligence system that creates electronic surrogates based on the analysis of emotions and personality.
[0365] A "prompt" refers to a text-based input used to give instructions to a generative AI model.
[0366] To implement this invention, the user must first register with the system and provide audio and video data. The user uses a smartphone or computer to record audio using a microphone and transmits the audio data to the system. They also record video of their face using a camera and upload it as video data.
[0367] The terminal converts the audio and video data received from the user into the appropriate format. This process converts audio data to WAV or MP3 format and video data to MP4 format. This ensures the data is suitable for analysis on the server.
[0368] The server receives the converted data and performs analysis using an emotion engine. The emotion engine identifies the user's emotional state by analyzing the tone, pitch, speed, and facial expressions of their voice. For example, if the voice is high-pitched and fast-paced, it is determined that the user is likely excited.
[0369] Based on the analysis results, the server uses a generative AI model to generate an electronic proxy that reflects the user's personality. The electronic proxy lists response patterns to appropriately respond to the user's emotions and uses them during conversations. Dynamic adjustment functions are also incorporated, allowing for immediate response to real-time changes in emotions.
[0370] For example, if a user tells the system, "I have some great news today!", the electronic agent will understand the user's emotion and reply, "That's wonderful! Congratulations!" In this way, communication that is sensitive to the user's feelings is realized.
[0371] An example of a prompt might be, "List the responses from the user's most recent conversation history that elicited the most positive emotions." Based on such prompts, the generative AI model can continuously learn and become capable of performing increasingly accurate dialogues.
[0372] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0373] Step 1:
[0374] Users log in to the system and input audio and video data. They use the microphone and camera on their smartphone or computer to record audio and video, including everyday conversations and specific emotional states, and upload them to the system. The input data consists of raw audio and video files.
[0375] Step 2:
[0376] The terminal converts the raw data received from the user into the appropriate format. Audio data is converted to WAV or MP3 format, and video data is converted to MP4 format. This format conversion makes the data ready for processing on the server. The input is raw data, and the output is data converted to the appropriate format.
[0377] Step 3:
[0378] The server inputs formatted data received from the terminal into the emotion engine for analysis. Here, the user's emotional state is analyzed based on data such as voice tone, pitch, speed, and facial expression changes obtained from the video. For example, voice pitch and speed are considered indicators of emotions such as joy or excitement. The input is formatted audio and video data, and the output is the analyzed emotional state.
[0379] Step 4:
[0380] The server uses a generative AI model to construct an electronic proxy that reflects the user's personality based on the analyzed emotional state. The electronic proxy prepares responses and attitudes that match that emotion and lists them. The generative AI model outputs an appropriate dialogue format that corresponds to the user's emotions and personality based on a given prompt sentence. For example, the prompt might be, "Generate a dialogue format that expresses the user's joy." The input is the emotional state and the prompt sentence, and the output is a response pattern generated based on the emotion.
[0381] Step 5:
[0382] The server interacts with the user using a generated electronic proxy. Here, it utilizes real-time dynamic adjustment capabilities to respond immediately to changes in the user's emotions. For example, if the user is suddenly surprised, the electronic proxy will instantly provide an appropriate response. The input is the user's current emotional state, and the output is the dynamically adjusted response.
[0383] Step 6:
[0384] The server records a history of the interactions that take place. This allows users to review past responses and provide feedback to help improve the system. This interaction history includes the user's emotions and the responses of the electronic proxy. The input is the data of the interactions that took place, and the output is the saved interaction history.
[0385] (Application Example 2)
[0386] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0387] Traditionally, customer service in commercial facilities has been uniform, making it difficult to provide appropriate suggestions and services tailored to the emotional state and preferences of individual customers. Furthermore, there has been a lack of means to recommend the most suitable products and services to meet the diverse needs of customers.
[0388] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0389] In this invention, the server includes means for generating an electronic surrogate that reflects the individual's personality, means for collecting and analyzing user information, and means for analyzing emotions and suggesting products based on that analysis. This makes it possible to provide appropriate product suggestions and services tailored to each customer's emotional state.
[0390] An "electronic proxy that reflects individuality" is a digital entity that takes into account the user's personal characteristics and preferences and communicates on their behalf.
[0391] "User information" refers to audio data, video data, and other related data provided by the user, and serves as the basis for analyzing the user's emotions and characteristics.
[0392] "Means for collecting and analyzing data" refers to functions that receive and analyze data obtained from users in an appropriate format to identify the user's emotional state and characteristics.
[0393] "Means of managing interactions with other users using electronic proxies" refers to functions that perform communication and optimize interactions on behalf of the user.
[0394] "Means for retaining generated communication history" refers to a function for saving records of interactions with the user and keeping them in a state where they can be referenced later.
[0395] "A means of analyzing emotions and proposing products based on them" refers to a function that analyzes the user's emotional state and provides appropriate products or services based on the results.
[0396] This invention aims to build a system for analyzing customer emotions in physical stores and providing optimal product recommendations. Specific embodiments of the present invention are shown below.
[0397] Users first access the platform using smartphones or tablet devices installed within the store. These devices are equipped with voice input capabilities and cameras, making it possible to collect the user's voice and image data.
[0398] The device converts the collected data into an appropriate format and sends it to the server over the network. The server receives each user's data and analyzes it using speech recognition APIs (e.g., Google Cloud Speech-to-Text API) and facial recognition APIs (e.g., Microsoft Azure Face API). This analysis identifies the user's emotional state.
[0399] Based on the analysis results, the server generates an electronic proxy and recommends products and services that match the user's emotions. This process utilizes an emotion engine, which analyzes the user's facial expressions and vocal characteristics.
[0400] For example, if a user is observed looking at products in a store and appears to be deep in thought, their facial expression can be interpreted as "undecided." In this case, the server will display recommended product reviews and suggestions tailored to that user on their device, providing a more engaging shopping experience.
[0401] A concrete example of a prompt might be, "This user does not have a smartphone, but suggest a way to estimate his emotions from his facial expressions and generate a customer service profile." By inputting this prompt into the AI model, the data analysis and product suggestion process is automated.
[0402] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0403] Step 1:
[0404] The device collects audio and image data provided by the user. This collected data is processed in real time within the device and converted into a format that can be analyzed on the server. The input is raw audio and image data, and the output is formatted digital data.
[0405] Step 2:
[0406] The converted digital data is transmitted from the terminal to the server via the network. Specifically, this involves generating data packets and transferring them to the server using the appropriate communication protocol. The input is the formatted digital data, and the output is the transmission of data to the server.
[0407] Step 3:
[0408] The server inputs the received data into the speech recognition API and facial expression recognition API, and analyzes the user's voice tone and facial expressions. Data processing involves extracting features from the voice tone and detecting facial expression features from the image. The input is digital data sent to the server, and the output is the analyzed emotional state data.
[0409] Step 4:
[0410] Based on the analysis results, the server uses a generative AI model to generate product and service suggestions that are appropriate for the user's emotions. A prompt is input to the generative AI model, and the output is the suggested content. This process specifically involves sending a prompt to the generative AI model and receiving the result. The input is the analyzed emotional state data, and the output is the suggested content.
[0411] Step 5:
[0412] The server sends the generated suggestions to the terminal and displays them to the user. This allows the user to check the suggested products and services in the store. The specific operations include data transmission and screen display. The input is the generated suggestions, and the output is information that the user can visually confirm.
[0413] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0414] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0415] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0416] [Third Embodiment]
[0417] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0418] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0419] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0420] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0421] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0422] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0423] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0424] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0425] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0426] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0427] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0428] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0429] This invention relates to a system for generating a proxy that electronically mimics a user's personality and automatically communicates with other users on social networking services (SNS). Specific embodiments are described below.
[0430] First, users connect their social media accounts to the platform and provide data that reflects their personality. This includes text-based conversation logs, audio clips, and video clips. This data forms the basis for capturing user characteristics.
[0431] Next, the terminal formats the received data and sends it to the server. Upon receiving this data, the server uses natural language processing techniques to analyze the user's conversation patterns and unique writing style. Simultaneously, it uses speech synthesis technology to extract features of the user's voice and build a profile to mimic the user's voice.
[0432] Furthermore, video analysis technology is used to generate avatars that reflect non-verbal elements such as the user's gestures and facial expressions. These avatars visually represent the user's personality and are used when the electronic proxy engages in conversation.
[0433] The generated electronic proxy initiates a conversation with other users on the SNS platform via a server. During this process, the electronic proxy generates responses based on the user's previously learned conversational style. While the conversation content is generated in real time, the generation process incorporates individual contextual understanding and adaptive capabilities.
[0434] For example, if a user frequently uses the characteristic daily greeting "Good morning!", the electronic proxy will also use the same greeting at the same time. In addition, the electronic proxy records conversation logs and further refines the user's personality based on the content and frequency of conversations. This allows the proxy to more accurately reflect the user's evolving communication style over time.
[0435] In this way, users can save their time while maintaining their individuality and continuing to have rich interactions with other users on social media. This system makes it possible to further enhance the communication experience that the platform provides.
[0436] The following describes the processing flow.
[0437] Step 1:
[0438] Users upload necessary information to the platform, such as their profile information, past conversation history, audio data, and video data.
[0439] Step 2:
[0440] The terminal receives this user data, converts it to the required format, and then securely and efficiently transmits the data to the server.
[0441] Step 3:
[0442] The server analyzes the user data it receives and uses natural language processing techniques to extract characteristic expressions and frequently occurring phrases from the conversation logs.
[0443] Step 4:
[0444] The server uses speech synthesis technology based on the audio data to analyze the characteristics of the user's voice, generate a voice profile, and prepare to reproduce the user's voice.
[0445] Step 5:
[0446] The server analyzes the video data and generates an avatar that reflects the user's gestures and facial expressions. This avatar then acts as an electronic proxy.
[0447] Step 6:
[0448] Based on these analysis results, the server builds individually customized electronic proxies that reflect the user's personality and style.
[0449] Step 7:
[0450] Users choose which users they communicate with on the social networking platform.
[0451] Step 8:
[0452] The server sets up a conversation session with the selected electronic proxy and initiates a dialogue between the electronic proxies.
[0453] Step 9:
[0454] The server generates the conversation in real time, adjusting the response at each step to ensure that the meaning and flow of the conversation are natural.
[0455] Step 10:
[0456] The server automatically saves the conversation history and manages logs so that users can review the content later.
[0457] Step 11:
[0458] Throughout this process, users occasionally provide feedback, which is used to further customize the electronic proxy and improve its accuracy.
[0459] (Example 1)
[0460] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0461] In modern society, while communication on social media is active, a challenge remains in that users find it difficult to maintain their individuality and continue interacting while reducing the time burden. In particular, there is a need for proxy systems that can reflect a user's unique writing style, expression, and non-verbal characteristics such as voice and gestures.
[0462] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0463] In this invention, the server includes means for generating an electronic surrogate that mimics the user's characteristics, means for collecting and analyzing the user's conversation history, voice data, and video data, means for autonomously communicating with other users through the electronic surrogate, and means for recording the generated communication data and learning the user's personality. As a result, the user can communicate smoothly on social networking services while reducing the time burden by using a surrogate that effectively reflects their own personality.
[0464] A "user" is an individual or group that utilizes the system and provides data that reflects their own personality.
[0465] "Characteristics" refer to elements that represent a user's individuality, such as their unique conversation patterns, writing style, tone of voice, and gestures.
[0466] An "electronic proxy" is a virtual entity that mimics the characteristics of a user and engages in autonomous communication with other users on social networking services (SNS).
[0467] "Conversation history" refers to a record of the communication a user has had so far, including text-based logs.
[0468] "Audio data" refers to digital audio files that contain information about the user's voice, and are data that possesses characteristics such as timbre and pitch.
[0469] "Video data" refers to video files that include the user's gestures and facial expressions, and is digital data that represents non-verbal elements.
[0470] "Analyzing" refers to a method of extracting user characteristics and behavioral patterns by examining data in detail.
[0471] "Autonomously" refers to a system that performs functions independently without requiring explicit user intervention.
[0472] This invention is a system that generates an electronic proxy that accurately mimics the user's characteristics and autonomously communicates with other users on social networking services (SNS).
[0473] Data collection and analysis
[0474] Users connect their social media accounts to the system and provide text-based conversation history, audio data, and video data. This forms a foundation for reflecting the user's unique conversation patterns and nonverbal characteristics.
[0475] The terminal converts the data received from the user into a standard format and sends it to the server. At this stage, technologies such as BERT and GPT are used as natural language processing engines to analyze the conversation data. A speech synthesis library is used for feature analysis of the audio data. For video data, for example, OpenCV is used to perform facial recognition and gesture extraction to capture the user's nonverbal characteristics.
[0476] Generation of electronic proxy
[0477] The server generates an electronic avatar that mimics the user's voice and video characteristics based on the analyzed data. It visually constructs an avatar, including the user's facial expressions and gestures, using 3D modeling. It then forms a speech synthesis profile and generates a voice similar to the user's.
[0478] Autonomous communication
[0479] The server enables autonomous communication on social networking services (SNS) through generated electronic proxies. In this process, a generative AI model is used to provide real-time responses. A concrete example of a prompt is, "Generate a response based on a user's frequent conversation about the weather." This prompt allows the electronic proxies to form natural responses such as, "It's sunny today."
[0480] This system allows users to maintain their individuality while continuing online interactions, reducing the time burden involved.
[0481] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0482] Step 1:
[0483] Users connect their social media accounts to the system and provide past conversation history, audio data, and video data as input data. This data is accumulated as a foundation to reflect the user's personality.
[0484] Step 2:
[0485] The terminal processes the input data received from the user into a standard format (e.g., JSON) and sends it to the server. Audio data is converted to an appropriate sample rate, and video data is converted to an appropriate resolution and frame rate. This creates formatted data.
[0486] Step 3:
[0487] The server analyzes the formatted data it receives. First, it uses a natural language processing engine (such as BERT or GPT) to identify the user's conversation patterns and writing style from the conversation history. Next, it performs acoustic analysis to extract voice features (tone and pitch) from the audio data. Furthermore, it uses technologies such as OpenCV to recognize gestures and facial expressions from the video data. This generates a user personality profile.
[0488] Step 4:
[0489] The server generates an electronic surrogate based on the analysis results. Specifically, it uses 3D modeling and speech synthesis to construct an electronic avatar that mimics the user's voice and video characteristics. This is how the electronic surrogate is generated.
[0490] Step 5:
[0491] The server utilizes a generative AI model to communicate with other users on the social networking service (SNS) through an electronic proxy. Specifically, it generates real-time conversations using prompts such as "Generate an appropriate response when the user talks about the weather." Based on the input conversation content, the electronic proxy returns a natural response. In this way, communication on the SNS is managed autonomously.
[0492] Step 6:
[0493] The server records the communications made by the electronic proxy and their outcomes. This record is analyzed and used to further refine the user's personality. This improves the accuracy of the electronic proxy and its reflection of the user's characteristics over time.
[0494] (Application Example 1)
[0495] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0496] In today's digital content world, there is a lack of systems that allow for broad interaction while preserving individual personality. In particular, there is a need for methods that enable content creators to interact with many fans simultaneously, efficiently managing large volumes of communication while maintaining their own unique style.
[0497] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0498] In this invention, the server includes information processing means for generating an electronic surrogate that mimics personality, computation means for collecting and analyzing user records, and means for displaying a virtual character via a smart device. This enables content creators to maintain their individuality while efficiently engaging in real-time, interactive communication with a wide range of fans.
[0499] An "electronic surrogate that mimics personality" is a virtual character created using information processing technology to reproduce the user's personality and expression.
[0500] "Information processing means" refers to a set of hardware and software for collecting, analyzing, and processing digital data.
[0501] "User records" refers to all data that represents the user's behavior and characteristics, including audio, video, and text.
[0502] A "computational means" is a system that has the function of performing calculations and analyses according to a specific purpose based on user records.
[0503] "Communication methods" refer to technologies that enable the transmission and reception of data, and facilitate the exchange of information between devices.
[0504] "Storage means" refers to devices and media used to retain digital data for extended periods.
[0505] A "smart device" is a general term for electronic devices that have internet connectivity and application execution capabilities.
[0506] A "virtual character" is a visually represented surrogate being created using digital technology.
[0507] A system for carrying out this invention has a configuration that generates an electronic surrogate that mimics the personality of a content creator and displays it as a virtual character via a smart device.
[0508] The server first collects and analyzes user records, i.e., audio, video, and text data, using information processing tools. For analysis, it uses natural language processing (NLP) libraries such as NLTK and spaCy, and for voice feature extraction, it uses speech synthesis engines such as Google Cloud Text-to-Speech. This allows the server to understand the user's characteristics and conversational style and generate a virtual character appropriate to their personality.
[0509] The generated virtual character is visualized using avatar generation tools such as Unity3D. This virtual character is displayed in real time via smart devices, specifically smart glasses, and interacts with fans and other users.
[0510] As a concrete example, consider a scenario where a content creator holds a virtual event with fans through smart glasses. In this event, the virtual character would mimic phrases and greetings frequently used by the creator in the past, and generate responses to fans' questions that reflect the creator's style.
[0511] An example of a prompt to input into the generation AI model would be, "Generate a response that greets the audience warmly and in a friendly manner, while maintaining the humorous tone that is characteristic of the creator."
[0512] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0513] Step 1:
[0514] The user inputs their own voice, video, and text data into the device. This data serves as foundational information for mimicking the user's personality. The device collects this data, formats it, and sends it to the server. The input is raw audio and video, and the output is formalized data.
[0515] Step 2:
[0516] The server analyzes the received user data using natural language processing tools. Specifically, it extracts voice characteristics and trains an AI model to generate conversational styles that reflect the creator's individuality. In this process, the user's past conversation logs are used as input, and a profile modeling the user's conversational patterns is obtained as output.
[0517] Step 3:
[0518] The server uses a speech synthesis engine to generate a voice profile to mimic the user's voice. The input is the user's voice data, and the output is the mimicked voice profile. This allows the virtual character to communicate using the user's voice.
[0519] Step 4:
[0520] The server uses an avatar generation tool based on the user's video data to create a virtual character that reflects the user's appearance and gestures. The input is the analyzed video data, and the output is a 3D modeled virtual character.
[0521] Step 5:
[0522] The generated virtual character is displayed through smart glasses on the device. This allows the user to use the virtual character as their proxy and interact with other users in real time. The input is the data of the virtual character, and the output is the visual interface displayed on the smart device.
[0523] Step 6:
[0524] The user's smart glasses control the actions of a virtual character and perform prompt-based conversations on social media and communication platforms on the user's behalf. Here, responses generated by a generative AI model are input and displayed as output in a format that allows for real-time interaction with viewers. An example of a prompt to input to the generative AI model is, "Generate a response that greets the viewers warmly and in a friendly manner, while maintaining the humorous tone that is characteristic of the creator."
[0525] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0526] This invention relates to a system for recognizing a user's emotions and for an electronic surrogate to appropriately adjust communication based on those emotions. The invention generates an electronic surrogate that reflects the user's personality and incorporates an emotion engine to achieve more natural and adaptive communication.
[0527] First, users register on the platform and provide profile information along with audio and video data. This data includes typical conversational scenes and situations involving specific emotions. This data serves as foundational information for accurately understanding the user's emotional state.
[0528] Next, the device converts the provided data into the appropriate format and sends it to the server. The server receives this data and begins the process of analyzing the user's emotional state using an emotion engine. The emotion engine identifies the user's emotional state by analyzing the tone, pitch, and speed of the voice from the audio data, and by analyzing changes in facial expressions from the video data.
[0529] Subsequently, the server constructs an electronic surrogate that reflects the user's personality based on the analysis results. This electronic surrogate lists appropriate responses and attitudes to specific emotions and uses them in various communication situations. This process incorporates dynamic adjustment mechanisms to quickly respond to potential emotional shifts.
[0530] For example, if a user uses an emotion engine to recognize that they are "happy" in a particular conversation, the electronic proxy can respond with appropriate words and facial expressions to express that joy. Conversely, if "anger" or "sadness" is recognized, the electronic proxy will show a corresponding attitude of comfort and empathy.
[0531] Furthermore, the server records the conversation history of the electronic proxy, allowing the user to review the history later and provide feedback. This enables the electronic proxy's responses to become more sophisticated and adapt to changes in the user.
[0532] Through the embodiments described above, users can obtain a more personalized experience and embody communication that resonates with their emotions through AI. This invention enables communication based on the user's unique emotions and elevates interactions on social networking platforms to a new dimension.
[0533] The following describes the processing flow.
[0534] Step 1:
[0535] Users log in to their accounts and upload their own audio and video data to the platform. This data includes scenes that evoke a variety of emotions.
[0536] Step 2:
[0537] The device converts the data received from the user into an appropriate format, compresses and encrypts the data, and then sends it to the server.
[0538] Step 3:
[0539] To analyze the user's emotional state, the server extracts voice tone and pitch from received audio data and detects subtle changes in facial expressions from video data. The emotion engine then generates an emotion label based on this information.
[0540] Step 4:
[0541] The server uses the analyzed emotion data to build an electronic proxy profile that corresponds to the user's emotional response. This allows the electronic proxy to generate responses optimized for the user's emotions.
[0542] Step 5:
[0543] Users select someone they want to communicate with on the SNS platform and start a conversation session.
[0544] Step 6:
[0545] The server automatically sets an appropriate conversation style and response for the electronic proxy based on the user's emotions, and initiates a conversation with the other party's electronic proxy.
[0546] Step 7:
[0547] During the conversation, the server adjusts the electronic proxy's responses in real time and dynamically applies feedback based on the latest emotional state.
[0548] Step 8:
[0549] The server records the conversation history and changes in emotions, and stores the analysis results in a database. Users can review this log later.
[0550] Step 9:
[0551] Later, users will review the provided conversation logs, provide feedback, and improve the electronic proxy response tactics to build a more user-friendly communication experience.
[0552] (Example 2)
[0553] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0554] In today's communication environment, there is a lack of response that responds to users' emotions, making natural communication that reflects individual feelings and situations difficult. Furthermore, flexibility to respond immediately to changes in emotions is also required, but conventional technologies do not adequately address this. Therefore, there is a need to provide adaptive communication methods based on the individuality and emotions of users.
[0555] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0556] In this invention, the server includes means for analyzing the user's emotional state, means for generating an electronic surrogate that reflects the user's personality, and means for incorporating a dynamic adjustment function. This enables personalized responses that are in line with the user's emotions.
[0557] A "user" is an entity that registers with the system and provides audio and video information.
[0558] "Emotional state" refers to the user's mental condition as identified through the analysis of audio and video.
[0559] An "electronic proxy" is a virtual entity created based on personality and emotions, used for communication.
[0560] "Communication" refers to the process of exchanging information with other users through electronic agents.
[0561] "Dynamic adjustment function" refers to a function that automatically optimizes the electronic proxy's response in accordance with changes in emotions and situations.
[0562] A "generative AI model" is an artificial intelligence system that creates electronic surrogates based on the analysis of emotions and personality.
[0563] A "prompt" refers to a text-based input used to give instructions to a generative AI model.
[0564] To implement this invention, the user must first register with the system and provide audio and video data. The user uses a smartphone or computer to record audio using a microphone and transmits the audio data to the system. They also record video of their face using a camera and upload it as video data.
[0565] The terminal converts the audio and video data received from the user into the appropriate format. This process converts audio data to WAV or MP3 format and video data to MP4 format. This ensures the data is suitable for analysis on the server.
[0566] The server receives the converted data and performs analysis using an emotion engine. The emotion engine identifies the user's emotional state by analyzing the tone, pitch, speed, and facial expressions of their voice. For example, if the voice is high-pitched and fast-paced, it is determined that the user is likely excited.
[0567] Based on the analysis results, the server uses a generative AI model to generate an electronic proxy that reflects the user's personality. The electronic proxy lists response patterns to appropriately respond to the user's emotions and uses them during conversations. Dynamic adjustment functions are also incorporated, allowing for immediate response to real-time changes in emotions.
[0568] For example, if a user tells the system, "I have some great news today!", the electronic agent will understand the user's emotion and reply, "That's wonderful! Congratulations!" In this way, communication that is sensitive to the user's feelings is realized.
[0569] An example of a prompt might be, "List the responses from the user's most recent conversation history that elicited the most positive emotions." Based on such prompts, the generative AI model can continuously learn and become capable of performing increasingly accurate dialogues.
[0570] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0571] Step 1:
[0572] Users log in to the system and input audio and video data. They use the microphone and camera on their smartphone or computer to record audio and video, including everyday conversations and specific emotional states, and upload them to the system. The input data consists of raw audio and video files.
[0573] Step 2:
[0574] The terminal converts the raw data received from the user into the appropriate format. Audio data is converted to WAV or MP3 format, and video data is converted to MP4 format. This format conversion makes the data ready for processing on the server. The input is raw data, and the output is data converted to the appropriate format.
[0575] Step 3:
[0576] The server inputs formatted data received from the terminal into the emotion engine for analysis. Here, the user's emotional state is analyzed based on data such as voice tone, pitch, speed, and facial expression changes obtained from the video. For example, voice pitch and speed are considered indicators of emotions such as joy or excitement. The input is formatted audio and video data, and the output is the analyzed emotional state.
[0577] Step 4:
[0578] The server uses a generative AI model to construct an electronic proxy that reflects the user's personality based on the analyzed emotional state. The electronic proxy prepares responses and attitudes that match that emotion and lists them. The generative AI model outputs an appropriate dialogue format that corresponds to the user's emotions and personality based on a given prompt sentence. For example, the prompt might be, "Generate a dialogue format that expresses the user's joy." The input is the emotional state and the prompt sentence, and the output is a response pattern generated based on the emotion.
[0579] Step 5:
[0580] The server interacts with the user using a generated electronic proxy. Here, it utilizes real-time dynamic adjustment capabilities to respond immediately to changes in the user's emotions. For example, if the user is suddenly surprised, the electronic proxy will instantly provide an appropriate response. The input is the user's current emotional state, and the output is the dynamically adjusted response.
[0581] Step 6:
[0582] The server records a history of the interactions that take place. This allows users to review past responses and provide feedback to help improve the system. This interaction history includes the user's emotions and the responses of the electronic proxy. The input is the data of the interactions that took place, and the output is the saved interaction history.
[0583] (Application Example 2)
[0584] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0585] Traditionally, customer service in commercial facilities has been uniform, making it difficult to provide appropriate suggestions and services tailored to the emotional state and preferences of individual customers. Furthermore, there has been a lack of means to recommend the most suitable products and services to meet the diverse needs of customers.
[0586] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0587] In this invention, the server includes means for generating an electronic surrogate that reflects the individual's personality, means for collecting and analyzing user information, and means for analyzing emotions and suggesting products based on that analysis. This makes it possible to provide appropriate product suggestions and services tailored to each customer's emotional state.
[0588] An "electronic proxy that reflects individuality" is a digital entity that takes into account the user's personal characteristics and preferences and communicates on their behalf.
[0589] "User information" refers to audio data, video data, and other related data provided by the user, and serves as the basis for analyzing the user's emotions and characteristics.
[0590] "Means for collecting and analyzing data" refers to functions that receive and analyze data obtained from users in an appropriate format to identify the user's emotional state and characteristics.
[0591] "Means of managing interactions with other users using electronic proxies" refers to functions that perform communication and optimize interactions on behalf of the user.
[0592] "Means for retaining generated communication history" refers to a function for saving records of interactions with the user and keeping them in a state where they can be referenced later.
[0593] "A means of analyzing emotions and proposing products based on them" refers to a function that analyzes the user's emotional state and provides appropriate products or services based on the results.
[0594] This invention aims to build a system for analyzing customer emotions in physical stores and providing optimal product recommendations. Specific embodiments of the present invention are shown below.
[0595] Users first access the platform using smartphones or tablet devices installed within the store. These devices are equipped with voice input capabilities and cameras, making it possible to collect the user's voice and image data.
[0596] The device converts the collected data into an appropriate format and sends it to the server over the network. The server receives each user's data and analyzes it using speech recognition APIs (e.g., Google Cloud Speech-to-Text API) and facial recognition APIs (e.g., Microsoft Azure Face API). This analysis identifies the user's emotional state.
[0597] Based on the analysis results, the server generates an electronic proxy and recommends products and services that match the user's emotions. This process utilizes an emotion engine, which analyzes the user's facial expressions and vocal characteristics.
[0598] For example, if a user is observed looking at products in a store and appears to be deep in thought, their facial expression can be interpreted as "undecided." In this case, the server will display recommended product reviews and suggestions tailored to that user on their device, providing a more engaging shopping experience.
[0599] A concrete example of a prompt might be, "This user does not have a smartphone, but suggest a way to estimate his emotions from his facial expressions and generate a customer service profile." By inputting this prompt into the AI model, the data analysis and product suggestion process is automated.
[0600] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0601] Step 1:
[0602] The device collects audio and image data provided by the user. This collected data is processed in real time within the device and converted into a format that can be analyzed on the server. The input is raw audio and image data, and the output is formatted digital data.
[0603] Step 2:
[0604] The converted digital data is transmitted from the terminal to the server via the network. Specifically, this involves generating data packets and transferring them to the server using the appropriate communication protocol. The input is the formatted digital data, and the output is the transmission of data to the server.
[0605] Step 3:
[0606] The server inputs the received data into the speech recognition API and facial expression recognition API, and analyzes the user's voice tone and facial expressions. Data processing involves extracting features from the voice tone and detecting facial expression features from the image. The input is digital data sent to the server, and the output is the analyzed emotional state data.
[0607] Step 4:
[0608] Based on the analysis results, the server uses a generative AI model to generate product and service suggestions that are appropriate for the user's emotions. A prompt is input to the generative AI model, and the output is the suggested content. This process specifically involves sending a prompt to the generative AI model and receiving the result. The input is the analyzed emotional state data, and the output is the suggested content.
[0609] Step 5:
[0610] The server sends the generated suggestions to the terminal and displays them to the user. This allows the user to check the suggested products and services in the store. The specific operations include data transmission and screen display. The input is the generated suggestions, and the output is information that the user can visually confirm.
[0611] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0612] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0613] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0614] [Fourth Embodiment]
[0615] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0616] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0617] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0618] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0619] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0620] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0621] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0622] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0623] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0624] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0625] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0626] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0627] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0628] This invention relates to a system for generating a proxy that electronically mimics a user's personality and automatically communicates with other users on social networking services (SNS). Specific embodiments are described below.
[0629] First, users connect their social media accounts to the platform and provide data that reflects their personality. This includes text-based conversation logs, audio clips, and video clips. This data forms the basis for capturing user characteristics.
[0630] Next, the terminal formats the received data and sends it to the server. Upon receiving this data, the server uses natural language processing techniques to analyze the user's conversation patterns and unique writing style. Simultaneously, it uses speech synthesis technology to extract features of the user's voice and build a profile to mimic the user's voice.
[0631] Furthermore, video analysis technology is used to generate avatars that reflect non-verbal elements such as the user's gestures and facial expressions. These avatars visually represent the user's personality and are used when the electronic proxy engages in conversation.
[0632] The generated electronic proxy initiates a conversation with other users on the SNS platform via a server. During this process, the electronic proxy generates responses based on the user's previously learned conversational style. While the conversation content is generated in real time, the generation process incorporates individual contextual understanding and adaptive capabilities.
[0633] For example, if a user frequently uses the characteristic daily greeting "Good morning!", the electronic proxy will also use the same greeting at the same time. In addition, the electronic proxy records conversation logs and further refines the user's personality based on the content and frequency of conversations. This allows the proxy to more accurately reflect the user's evolving communication style over time.
[0634] In this way, users can save their time while maintaining their individuality and continuing to have rich interactions with other users on social media. This system makes it possible to further enhance the communication experience that the platform provides.
[0635] The following describes the processing flow.
[0636] Step 1:
[0637] Users upload necessary information to the platform, such as their profile information, past conversation history, audio data, and video data.
[0638] Step 2:
[0639] The terminal receives this user data, converts it to the required format, and then securely and efficiently transmits the data to the server.
[0640] Step 3:
[0641] The server analyzes the user data it receives and uses natural language processing techniques to extract characteristic expressions and frequently occurring phrases from the conversation logs.
[0642] Step 4:
[0643] The server uses speech synthesis technology based on the audio data to analyze the characteristics of the user's voice, generate a voice profile, and prepare to reproduce the user's voice.
[0644] Step 5:
[0645] The server analyzes the video data and generates an avatar that reflects the user's gestures and facial expressions. This avatar then acts as an electronic proxy.
[0646] Step 6:
[0647] Based on these analysis results, the server builds individually customized electronic proxies that reflect the user's personality and style.
[0648] Step 7:
[0649] Users choose which users they communicate with on the social networking platform.
[0650] Step 8:
[0651] The server sets up a conversation session with the selected electronic proxy and initiates a dialogue between the electronic proxies.
[0652] Step 9:
[0653] The server generates the conversation in real time, adjusting the response at each step to ensure that the meaning and flow of the conversation are natural.
[0654] Step 10:
[0655] The server automatically saves the conversation history and manages logs so that users can review the content later.
[0656] Step 11:
[0657] Throughout this process, users occasionally provide feedback, which is used to further customize the electronic proxy and improve its accuracy.
[0658] (Example 1)
[0659] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0660] In modern society, while communication on social media is active, a challenge remains in that users find it difficult to maintain their individuality and continue interacting while reducing the time burden. In particular, there is a need for proxy systems that can reflect a user's unique writing style, expression, and non-verbal characteristics such as voice and gestures.
[0661] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0662] In this invention, the server includes means for generating an electronic surrogate that mimics the user's characteristics, means for collecting and analyzing the user's conversation history, voice data, and video data, means for autonomously communicating with other users through the electronic surrogate, and means for recording the generated communication data and learning the user's personality. As a result, the user can communicate smoothly on social networking services while reducing the time burden by using a surrogate that effectively reflects their own personality.
[0663] A "user" is an individual or group that utilizes the system and provides data that reflects their own personality.
[0664] "Characteristics" refer to elements that represent a user's individuality, such as their unique conversation patterns, writing style, tone of voice, and gestures.
[0665] An "electronic proxy" is a virtual entity that mimics the characteristics of a user and engages in autonomous communication with other users on social networking services (SNS).
[0666] "Conversation history" refers to a record of the communication a user has had so far, including text-based logs.
[0667] "Audio data" refers to digital audio files that contain information about the user's voice, and are data that possesses characteristics such as timbre and pitch.
[0668] "Video data" refers to video files that include the user's gestures and facial expressions, and is digital data that represents non-verbal elements.
[0669] "Analyzing" refers to a method of extracting user characteristics and behavioral patterns by examining data in detail.
[0670] "Autonomously" refers to a system that performs functions independently without requiring explicit user intervention.
[0671] This invention is a system that generates an electronic proxy that accurately mimics the user's characteristics and autonomously communicates with other users on social networking services (SNS).
[0672] Data collection and analysis
[0673] Users connect their social media accounts to the system and provide text-based conversation history, audio data, and video data. This forms a foundation for reflecting the user's unique conversation patterns and nonverbal characteristics.
[0674] The terminal converts the data received from the user into a standard format and sends it to the server. At this stage, technologies such as BERT and GPT are used as natural language processing engines to analyze the conversation data. A speech synthesis library is used for feature analysis of the audio data. For video data, for example, OpenCV is used to perform facial recognition and gesture extraction to capture the user's nonverbal characteristics.
[0675] Generation of electronic proxy
[0676] The server generates an electronic avatar that mimics the user's voice and video characteristics based on the analyzed data. It visually constructs an avatar, including the user's facial expressions and gestures, using 3D modeling. It then forms a speech synthesis profile and generates a voice similar to the user's.
[0677] Autonomous communication
[0678] The server enables autonomous communication on social networking services (SNS) through generated electronic proxies. In this process, a generative AI model is used to provide real-time responses. A concrete example of a prompt is, "Generate a response based on a user's frequent conversation about the weather." This prompt allows the electronic proxies to form natural responses such as, "It's sunny today."
[0679] This system allows users to maintain their individuality while continuing online interactions, reducing the time burden involved.
[0680] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0681] Step 1:
[0682] Users connect their social media accounts to the system and provide past conversation history, audio data, and video data as input data. This data is accumulated as a foundation to reflect the user's personality.
[0683] Step 2:
[0684] The terminal processes the input data received from the user into a standard format (e.g., JSON) and sends it to the server. Audio data is converted to an appropriate sample rate, and video data is converted to an appropriate resolution and frame rate. This creates formatted data.
[0685] Step 3:
[0686] The server analyzes the formatted data it receives. First, it uses a natural language processing engine (such as BERT or GPT) to identify the user's conversation patterns and writing style from the conversation history. Next, it performs acoustic analysis to extract voice features (tone and pitch) from the audio data. Furthermore, it uses technologies such as OpenCV to recognize gestures and facial expressions from the video data. This generates a user personality profile.
[0687] Step 4:
[0688] The server generates an electronic surrogate based on the analysis results. Specifically, it uses 3D modeling and speech synthesis to construct an electronic avatar that mimics the user's voice and video characteristics. This is how the electronic surrogate is generated.
[0689] Step 5:
[0690] The server utilizes a generative AI model to communicate with other users on the social networking service (SNS) through an electronic proxy. Specifically, it generates real-time conversations using prompts such as "Generate an appropriate response when the user talks about the weather." Based on the input conversation content, the electronic proxy returns a natural response. In this way, communication on the SNS is managed autonomously.
[0691] Step 6:
[0692] The server records the communications made by the electronic proxy and their outcomes. This record is analyzed and used to further refine the user's personality. This improves the accuracy of the electronic proxy and its reflection of the user's characteristics over time.
[0693] (Application Example 1)
[0694] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0695] In today's digital content world, there is a lack of systems that allow for broad interaction while preserving individual personality. In particular, there is a need for methods that enable content creators to interact with many fans simultaneously, efficiently managing large volumes of communication while maintaining their own unique style.
[0696] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0697] In this invention, the server includes information processing means for generating an electronic surrogate that mimics personality, computation means for collecting and analyzing user records, and means for displaying a virtual character via a smart device. This enables content creators to maintain their individuality while efficiently engaging in real-time, interactive communication with a wide range of fans.
[0698] An "electronic surrogate that mimics personality" is a virtual character created using information processing technology to reproduce the user's personality and expression.
[0699] "Information processing means" refers to a set of hardware and software for collecting, analyzing, and processing digital data.
[0700] "User records" refers to all data that represents the user's behavior and characteristics, including audio, video, and text.
[0701] A "computational means" is a system that has the function of performing calculations and analyses according to a specific purpose based on user records.
[0702] "Communication methods" refer to technologies that enable the transmission and reception of data, and facilitate the exchange of information between devices.
[0703] "Storage means" refers to devices and media used to retain digital data for extended periods.
[0704] A "smart device" is a general term for electronic devices that have internet connectivity and application execution capabilities.
[0705] A "virtual character" is a visually represented surrogate being created using digital technology.
[0706] A system for carrying out this invention has a configuration that generates an electronic surrogate that mimics the personality of a content creator and displays it as a virtual character via a smart device.
[0707] The server first collects and analyzes user records, i.e., audio, video, and text data, using information processing tools. For analysis, it uses natural language processing (NLP) libraries such as NLTK and spaCy, and for voice feature extraction, it uses speech synthesis engines such as Google Cloud Text-to-Speech. This allows the server to understand the user's characteristics and conversational style and generate a virtual character appropriate to their personality.
[0708] The generated virtual character is visualized using avatar generation tools such as Unity3D. This virtual character is displayed in real time via smart devices, specifically smart glasses, and interacts with fans and other users.
[0709] As a concrete example, consider a scenario where a content creator holds a virtual event with fans through smart glasses. In this event, the virtual character would mimic phrases and greetings frequently used by the creator in the past, and generate responses to fans' questions that reflect the creator's style.
[0710] An example of a prompt to input into the generation AI model would be, "Generate a response that greets the audience warmly and in a friendly manner, while maintaining the humorous tone that is characteristic of the creator."
[0711] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0712] Step 1:
[0713] The user inputs their own voice, video, and text data into the device. This data serves as foundational information for mimicking the user's personality. The device collects this data, formats it, and sends it to the server. The input is raw audio and video, and the output is formalized data.
[0714] Step 2:
[0715] The server analyzes the received user data using natural language processing tools. Specifically, it extracts voice characteristics and trains an AI model to generate conversational styles that reflect the creator's individuality. In this process, the user's past conversation logs are used as input, and a profile modeling the user's conversational patterns is obtained as output.
[0716] Step 3:
[0717] The server uses a speech synthesis engine to generate a voice profile to mimic the user's voice. The input is the user's voice data, and the output is the mimicked voice profile. This allows the virtual character to communicate using the user's voice.
[0718] Step 4:
[0719] The server uses an avatar generation tool based on the user's video data to create a virtual character that reflects the user's appearance and gestures. The input is the analyzed video data, and the output is a 3D modeled virtual character.
[0720] Step 5:
[0721] The generated virtual character is displayed through smart glasses on the device. This allows the user to use the virtual character as their proxy and interact with other users in real time. The input is the data of the virtual character, and the output is the visual interface displayed on the smart device.
[0722] Step 6:
[0723] The user's smart glasses control the actions of a virtual character and perform prompt-based conversations on social media and communication platforms on the user's behalf. Here, responses generated by a generative AI model are input and displayed as output in a format that allows for real-time interaction with viewers. An example of a prompt to input to the generative AI model is, "Generate a response that greets the viewers warmly and in a friendly manner, while maintaining the humorous tone that is characteristic of the creator."
[0724] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0725] This invention relates to a system for recognizing a user's emotions and for an electronic surrogate to appropriately adjust communication based on those emotions. The invention generates an electronic surrogate that reflects the user's personality and incorporates an emotion engine to achieve more natural and adaptive communication.
[0726] First, users register on the platform and provide profile information along with audio and video data. This data includes typical conversational scenes and situations involving specific emotions. This data serves as foundational information for accurately understanding the user's emotional state.
[0727] Next, the device converts the provided data into the appropriate format and sends it to the server. The server receives this data and begins the process of analyzing the user's emotional state using an emotion engine. The emotion engine identifies the user's emotional state by analyzing the tone, pitch, and speed of the voice from the audio data, and by analyzing changes in facial expressions from the video data.
[0728] Subsequently, the server constructs an electronic surrogate that reflects the user's personality based on the analysis results. This electronic surrogate lists appropriate responses and attitudes to specific emotions and uses them in various communication situations. This process incorporates dynamic adjustment mechanisms to quickly respond to potential emotional shifts.
[0729] For example, if a user uses an emotion engine to recognize that they are "happy" in a particular conversation, the electronic proxy can respond with appropriate words and facial expressions to express that joy. Conversely, if "anger" or "sadness" is recognized, the electronic proxy will show a corresponding attitude of comfort and empathy.
[0730] Furthermore, the server records the conversation history of the electronic proxy, allowing the user to review the history later and provide feedback. This enables the electronic proxy's responses to become more sophisticated and adapt to changes in the user.
[0731] Through the embodiments described above, users can obtain a more personalized experience and embody communication that resonates with their emotions through AI. This invention enables communication based on the user's unique emotions and elevates interactions on social networking platforms to a new dimension.
[0732] The following describes the processing flow.
[0733] Step 1:
[0734] Users log in to their accounts and upload their own audio and video data to the platform. This data includes scenes that evoke a variety of emotions.
[0735] Step 2:
[0736] The device converts the data received from the user into an appropriate format, compresses and encrypts the data, and then sends it to the server.
[0737] Step 3:
[0738] To analyze the user's emotional state, the server extracts voice tone and pitch from received audio data and detects subtle changes in facial expressions from video data. The emotion engine then generates an emotion label based on this information.
[0739] Step 4:
[0740] The server uses the analyzed emotion data to build an electronic proxy profile that corresponds to the user's emotional response. This allows the electronic proxy to generate responses optimized for the user's emotions.
[0741] Step 5:
[0742] Users select someone they want to communicate with on the SNS platform and start a conversation session.
[0743] Step 6:
[0744] The server automatically sets an appropriate conversation style and response for the electronic proxy based on the user's emotions, and initiates a conversation with the other party's electronic proxy.
[0745] Step 7:
[0746] During the conversation, the server adjusts the electronic proxy's responses in real time and dynamically applies feedback based on the latest emotional state.
[0747] Step 8:
[0748] The server records the conversation history and changes in emotions, and stores the analysis results in a database. Users can review this log later.
[0749] Step 9:
[0750] Later, users will review the provided conversation logs, provide feedback, and improve the electronic proxy response tactics to build a more user-friendly communication experience.
[0751] (Example 2)
[0752] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0753] In today's communication environment, there is a lack of response that responds to users' emotions, making natural communication that reflects individual feelings and situations difficult. Furthermore, flexibility to respond immediately to changes in emotions is also required, but conventional technologies do not adequately address this. Therefore, there is a need to provide adaptive communication methods based on the individuality and emotions of users.
[0754] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0755] In this invention, the server includes means for analyzing the user's emotional state, means for generating an electronic surrogate that reflects the user's personality, and means for incorporating a dynamic adjustment function. This enables personalized responses that are in line with the user's emotions.
[0756] A "user" is an entity that registers with the system and provides audio and video information.
[0757] "Emotional state" refers to the user's mental condition as identified through the analysis of audio and video.
[0758] An "electronic proxy" is a virtual entity created based on personality and emotions, used for communication.
[0759] "Communication" refers to the process of exchanging information with other users through electronic agents.
[0760] "Dynamic adjustment function" refers to a function that automatically optimizes the electronic proxy's response in accordance with changes in emotions and situations.
[0761] A "generative AI model" is an artificial intelligence system that creates electronic surrogates based on the analysis of emotions and personality.
[0762] A "prompt" refers to a text-based input used to give instructions to a generative AI model.
[0763] To implement this invention, the user must first register with the system and provide audio and video data. The user uses a smartphone or computer to record audio using a microphone and transmits the audio data to the system. They also record video of their face using a camera and upload it as video data.
[0764] The terminal converts the audio and video data received from the user into the appropriate format. This process converts audio data to WAV or MP3 format and video data to MP4 format. This ensures the data is suitable for analysis on the server.
[0765] The server receives the converted data and performs analysis using an emotion engine. The emotion engine identifies the user's emotional state by analyzing the tone, pitch, speed, and facial expressions of their voice. For example, if the voice is high-pitched and fast-paced, it is determined that the user is likely excited.
[0766] Based on the analysis results, the server uses a generative AI model to generate an electronic proxy that reflects the user's personality. The electronic proxy lists response patterns to appropriately respond to the user's emotions and uses them during conversations. Dynamic adjustment functions are also incorporated, allowing for immediate response to real-time changes in emotions.
[0767] For example, if a user tells the system, "I have some great news today!", the electronic agent will understand the user's emotion and reply, "That's wonderful! Congratulations!" In this way, communication that is sensitive to the user's feelings is realized.
[0768] An example of a prompt might be, "List the responses from the user's most recent conversation history that elicited the most positive emotions." Based on such prompts, the generative AI model can continuously learn and become capable of performing increasingly accurate dialogues.
[0769] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0770] Step 1:
[0771] Users log in to the system and input audio and video data. They use the microphone and camera on their smartphone or computer to record audio and video, including everyday conversations and specific emotional states, and upload them to the system. The input data consists of raw audio and video files.
[0772] Step 2:
[0773] The terminal converts the raw data received from the user into the appropriate format. Audio data is converted to WAV or MP3 format, and video data is converted to MP4 format. This format conversion makes the data ready for processing on the server. The input is raw data, and the output is data converted to the appropriate format.
[0774] Step 3:
[0775] The server inputs formatted data received from the terminal into the emotion engine for analysis. Here, the user's emotional state is analyzed based on data such as voice tone, pitch, speed, and facial expression changes obtained from the video. For example, voice pitch and speed are considered indicators of emotions such as joy or excitement. The input is formatted audio and video data, and the output is the analyzed emotional state.
[0776] Step 4:
[0777] The server uses a generative AI model to construct an electronic proxy that reflects the user's personality based on the analyzed emotional state. The electronic proxy prepares responses and attitudes that match that emotion and lists them. The generative AI model outputs an appropriate dialogue format that corresponds to the user's emotions and personality based on a given prompt sentence. For example, the prompt might be, "Generate a dialogue format that expresses the user's joy." The input is the emotional state and the prompt sentence, and the output is a response pattern generated based on the emotion.
[0778] Step 5:
[0779] The server interacts with the user using a generated electronic proxy. Here, it utilizes real-time dynamic adjustment capabilities to respond immediately to changes in the user's emotions. For example, if the user is suddenly surprised, the electronic proxy will instantly provide an appropriate response. The input is the user's current emotional state, and the output is the dynamically adjusted response.
[0780] Step 6:
[0781] The server records a history of the interactions that take place. This allows users to review past responses and provide feedback to help improve the system. This interaction history includes the user's emotions and the responses of the electronic proxy. The input is the data of the interactions that took place, and the output is the saved interaction history.
[0782] (Application Example 2)
[0783] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0784] Traditionally, customer service in commercial facilities has been uniform, making it difficult to provide appropriate suggestions and services tailored to the emotional state and preferences of individual customers. Furthermore, there has been a lack of means to recommend the most suitable products and services to meet the diverse needs of customers.
[0785] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0786] In this invention, the server includes means for generating an electronic surrogate that reflects the individual's personality, means for collecting and analyzing user information, and means for analyzing emotions and suggesting products based on that analysis. This makes it possible to provide appropriate product suggestions and services tailored to each customer's emotional state.
[0787] An "electronic proxy that reflects individuality" is a digital entity that takes into account the user's personal characteristics and preferences and communicates on their behalf.
[0788] "User information" refers to audio data, video data, and other related data provided by the user, and serves as the basis for analyzing the user's emotions and characteristics.
[0789] "Means for collecting and analyzing data" refers to functions that receive and analyze data obtained from users in an appropriate format to identify the user's emotional state and characteristics.
[0790] "Means of managing interactions with other users using electronic proxies" refers to functions that perform communication and optimize interactions on behalf of the user.
[0791] "Means for retaining generated communication history" refers to a function for saving records of interactions with the user and keeping them in a state where they can be referenced later.
[0792] "A means of analyzing emotions and proposing products based on them" refers to a function that analyzes the user's emotional state and provides appropriate products or services based on the results.
[0793] This invention aims to build a system for analyzing customer emotions in physical stores and providing optimal product recommendations. Specific embodiments of the present invention are shown below.
[0794] Users first access the platform using smartphones or tablet devices installed within the store. These devices are equipped with voice input capabilities and cameras, making it possible to collect the user's voice and image data.
[0795] The device converts the collected data into an appropriate format and sends it to the server over the network. The server receives each user's data and analyzes it using speech recognition APIs (e.g., Google Cloud Speech-to-Text API) and facial recognition APIs (e.g., Microsoft Azure Face API). This analysis identifies the user's emotional state.
[0796] Based on the analysis results, the server generates an electronic proxy and recommends products and services that match the user's emotions. This process utilizes an emotion engine, which analyzes the user's facial expressions and vocal characteristics.
[0797] For example, if a user is observed looking at products in a store and appears to be deep in thought, their facial expression can be interpreted as "undecided." In this case, the server will display recommended product reviews and suggestions tailored to that user on their device, providing a more engaging shopping experience.
[0798] A concrete example of a prompt might be, "This user does not have a smartphone, but suggest a way to estimate his emotions from his facial expressions and generate a customer service profile." By inputting this prompt into the AI model, the data analysis and product suggestion process is automated.
[0799] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0800] Step 1:
[0801] The device collects audio and image data provided by the user. This collected data is processed in real time within the device and converted into a format that can be analyzed on the server. The input is raw audio and image data, and the output is formatted digital data.
[0802] Step 2:
[0803] The converted digital data is transmitted from the terminal to the server via the network. Specifically, this involves generating data packets and transferring them to the server using the appropriate communication protocol. The input is the formatted digital data, and the output is the transmission of data to the server.
[0804] Step 3:
[0805] The server inputs the received data into the speech recognition API and facial expression recognition API, and analyzes the user's voice tone and facial expressions. Data processing involves extracting features from the voice tone and detecting facial expression features from the image. The input is digital data sent to the server, and the output is the analyzed emotional state data.
[0806] Step 4:
[0807] Based on the analysis results, the server uses a generative AI model to generate product and service suggestions that are appropriate for the user's emotions. A prompt is input to the generative AI model, and the output is the suggested content. This process specifically involves sending a prompt to the generative AI model and receiving the result. The input is the analyzed emotional state data, and the output is the suggested content.
[0808] Step 5:
[0809] The server sends the generated suggestions to the terminal and displays them to the user. This allows the user to check the suggested products and services in the store. The specific operations include data transmission and screen display. The input is the generated suggestions, and the output is information that the user can visually confirm.
[0810] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0811] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0812] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0813] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0814] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0815] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0816] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0817] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0818] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0819] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0820] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0821] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0822] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0823] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0824] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0825] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0826] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0827] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0828] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0829] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0830] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0831] The following is further disclosed regarding the embodiments described above.
[0832] (Claim 1)
[0833] A means of generating electronic surrogates that reflect individuality,
[0834] Means for collecting and analyzing user data,
[0835] A means of managing communication with other users using electronic proxies,
[0836] A means of saving the generated conversation history,
[0837] A system that includes this.
[0838] (Claim 2)
[0839] The system according to claim 1, comprising means for constructing an electronic proxy based on the user's voice data and video data.
[0840] (Claim 3)
[0841] The system according to claim 1, comprising means for analyzing the user's personality using natural language processing.
[0842] "Example 1"
[0843] (Claim 1)
[0844] Means for generating an electronic surrogate that mimics the user's characteristics,
[0845] A means for collecting and analyzing user conversation history, audio data, and video data,
[0846] A means of autonomously communicating with other users through electronic proxies,
[0847] A means of recording generated communication data and learning the user's personality,
[0848] A system that includes this.
[0849] (Claim 2)
[0850] The system according to claim 1, comprising means for analyzing the user's voice and video characteristics and forming an electronic avatar.
[0851] (Claim 3)
[0852] The system according to claim 1, comprising means for analyzing the user's writing style and conversational characteristics using natural language processing technology.
[0853] "Application Example 1"
[0854] (Claim 1)
[0855] Information processing means for generating electronic proxies that mimic individuality,
[0856] A computational means for collecting and analyzing user records,
[0857] A means of communication for coordinating interactions with other users using electronic proxies,
[0858] A storage means for storing a record of the generated interactions,
[0859] A means of displaying a virtual character via a smart device,
[0860] A system that includes this.
[0861] (Claim 2)
[0862] The system according to claim 1, comprising a configuration for generating an electronic proxy based on the user's voice recordings and video recordings.
[0863] (Claim 3)
[0864] The system according to claim 1, comprising an algorithm for analyzing the user's personality using natural language processing.
[0865] "Example 2 of combining an emotion engine"
[0866] (Claim 1)
[0867] A means of analyzing the emotional state of users,
[0868] A means of generating an electronic proxy that reflects individuality,
[0869] A means of managing communication with other users using an electronic proxy,
[0870] A means of saving the generated conversation history,
[0871] Means for incorporating dynamic adjustment functions,
[0872] A system that includes this.
[0873] (Claim 2)
[0874] The system according to claim 1, comprising means for constructing an electronic agent based on the user's voice information and video information.
[0875] (Claim 3)
[0876] The system according to claim 1, comprising means for analyzing a user's personality using natural language interpretation.
[0877] "Application example 2 when combining with an emotional engine"
[0878] (Claim 1)
[0879] A means of generating electronic surrogates that reflect individuality,
[0880] Means for collecting and analyzing user information,
[0881] A means of managing interactions with other users using electronic proxies,
[0882] A means for storing the generated communication history,
[0883] A means of analyzing emotions and proposing products based on them,
[0884] A system that includes this.
[0885] (Claim 2)
[0886] The system according to claim 1, comprising means for constructing an electronic proxy based on the user's voice and video information.
[0887] (Claim 3)
[0888] The system according to claim 1, comprising means for analyzing user characteristics using natural language processing. [Explanation of symbols]
[0889] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of generating electronic surrogates that reflect individuality, Means for collecting and analyzing user data, A means of managing communication with other users using electronic proxies, A means of saving the generated conversation history, A system that includes this.
2. The system according to claim 1, comprising means for constructing an electronic proxy based on the user's voice data and video data.
3. The system according to claim 1, comprising means for analyzing the user's personality using natural language processing.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A