system

A system using voice analysis and interactive games addresses the challenge of distant family communication, enabling effective information sharing and bond strengthening through real-time voice processing and personalized interaction.

JP2026070864APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The physical distance between family members hinders natural and smooth communication, leading to weakened bonds and inadequate information sharing, necessitating a technological solution that utilizes voice recognition and conversation analysis to automate information organization and sharing.

Method used

A system incorporating voice analysis, conversation pattern learning, interactive games, and information management to facilitate rich communication and information sharing among family members, utilizing voice data processing, machine learning, and real-time interaction.

Benefits of technology

Enables efficient communication and information exchange across distances, enhancing family bonds through accurate voice recognition, pattern learning, and tailored interactive experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070864000001_ABST
    Figure 2026070864000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A voice analysis method for acquiring voice data in real time and identifying the characteristics of the speaker, A learning method for analyzing past conversation history and learning conversation patterns, A means of providing and displaying interactive games to participants, A means of managing information to collect and organize family events and health information, A means of distributing information to notify user terminals of organized information. Includes system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] It is difficult to maintain natural and smooth communication among family members living at a distance, and there is a problem that the physical distance becomes an obstacle to bonds and information sharing. For this reason, there is a need for an epoch-making means that utilizes voice recognition and conversation analysis to smooth family conversations beyond geographical constraints, and further automates information organization and sharing among family members.

Means for Solving the Problems

[0005] To solve this problem, the present invention proposes a system that includes a voice analysis means for acquiring voice data in real time and identifying speaker characteristics, a learning means for analyzing past conversation history and learning conversation patterns, a game provision means for providing interactive games to participants, an information management means for collecting and organizing family events and health information, and an information distribution means for notifying user terminals of the organized information. This enables rich communication and information sharing among family members that transcends physical distance.

[0006] "Audio data" refers to sound information acquired during a conversation, and is particularly used in digital format for speech recognition and speaker identification.

[0007] "Voice analysis means" refers to technical means that process collected voice data and have the function of identifying the speaker's characteristics and voiceprint.

[0008] A "learning tool" is a technological tool that uses machine learning based on past data to model conversation patterns and characteristics.

[0009] An "interactive game" is an activity or game in which multiple participants can enjoy themselves by providing two-way feedback.

[0010] "Game delivery means" refers to the technical means of selecting an interactive game and presenting it to the user in an interactive format.

[0011] "Information management tools" are technological tools that have the function of collecting and organizing family-related event and health information and prioritizing it based on its importance.

[0012] "Information distribution means" refers to technical means for appropriately notifying and sharing organized information with user terminals. [Brief explanation of the drawing]

[0013] [Figure 1]This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0019] In the following embodiments, a labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention provides a system that facilitates smooth communication between distant family members and includes functions for voice analysis, conversation pattern learning, interactive games, and information organization and distribution. The system operates through the interaction of a server, a terminal, and a user.

[0035] Forms of speech analysis

[0036] When a user initiates a video call, the device uses its microphone to acquire audio data in real time. This allows the device to efficiently capture the user's speech and transmit highly accurate audio data to the server. The server analyzes the received audio data using a speech recognition algorithm to identify the speaker and learn the unique voice characteristics of each user, enabling it to accurately understand the conversation.

[0037] Forms of learning conversation patterns

[0038] The server accumulates past conversation history and learns conversation patterns using frequency analysis and topic models based on that data. This allows it to understand the content of what was said and the patterns of the conversation flow, and to utilize the user's unique communication style in future conversations.

[0039] Interactive game delivery methods

[0040] The server selects an appropriate interactive game based on the conversation context and participants' interests. The game is displayed to the user via their device and is designed for intuitive operation. Specifically, it supports the user's participation in a quiz game, guiding them through the game's progression to elicit smart responses.

[0041] Forms of information organization and distribution

[0042] The server aggregates family event information and health data, organizing important information through an information management system. This information is then categorized based on urgency and importance, and notified to terminals using an information distribution system. This notification allows users to quickly take necessary actions.

[0043] In this way, the system of the present invention can efficiently carry out a series of processes from voice recognition to learning and information provision, thereby enriching communication among family members that transcends physical distance.

[0044] The following describes the processing flow.

[0045] Step 1:

[0046] When a user initiates a video call, the device immediately acquires audio data using its built-in microphone. This audio data is then processed to ensure clarity using noise-canceling technology.

[0047] Step 2:

[0048] The terminal compresses the acquired audio data in real time and sends it to the server via a secure protocol. The data is transferred using a highly efficient encoding method, minimizing latency.

[0049] Step 3:

[0050] The server inputs the received audio data into a speech recognition algorithm for analysis. Here, speech features are extracted to identify each speaker, and speaker identification is performed.

[0051] Step 4:

[0052] The server saves the identified speaker's voice characteristics to a cloud database and updates the dataset used for subsequent conversation analysis. This data is also protected by security protocols.

[0053] Step 5:

[0054] The server performs frequency analysis based on past conversation history to extract characteristic conversation patterns. It then uses topic models to analyze the flow of topics and the relationships between important keywords.

[0055] Step 6:

[0056] Based on learned conversation patterns, the server prepares to naturally suggest relevant information and topics in the next conversation, thereby helping to keep the conversation flowing smoothly.

[0057] Step 7:

[0058] The server selects an appropriate interactive game based on the conversation content and the user's interests, and prepares it for play. This selection process also utilizes the user's past game history stored in the database.

[0059] Step 8:

[0060] The device presents the user with an interactive game provided by the server and encourages them to participate. If the user consents, the game starts and leaderboard and score display functions are enabled.

[0061] Step 9:

[0062] The server aggregates and organizes family calendar events and health information. This information is categorized according to priority, and only the information relevant to each user is selected.

[0063] Step 10:

[0064] The server notifies the terminal of organized information at the necessary time, helping users to respond quickly. Each notification is sent as a push notification, so it reaches the user immediately.

[0065] (Example 1)

[0066] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0067] A key challenge is achieving smooth and intimate communication among family members living far apart. In particular, technological solutions are needed to prevent the weakening of communication caused by physical distance and to deepen family ties. Furthermore, improvements in voice recognition accuracy and the provision of information tailored to individual communication styles are essential.

[0068] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0069] In this invention, the server includes voice acquisition means, data transmission means, data analysis means, conversation analysis means, activity provision means, and information organization and distribution means. This enables real-time and effective communication that transcends physical distance and allows for the provision of interactive support to strengthen family bonds.

[0070] "Voice acquisition means" refers to technology for accurately acquiring the user's voice and collecting clear audio data by reducing external noise.

[0071] "Data transmission means" refers to technology for encoding acquired audio data and sending it to a server using a secure communication protocol.

[0072] "Data analysis means" refers to a technology that analyzes audio data received on a server using speech recognition technology, converts the spoken content into text, and identifies individual speakers.

[0073] A "conversation analysis method" is a technique that uses natural language processing technology to learn conversation patterns based on accumulated conversation data and identify frequently occurring topics and expressions.

[0074] "Activity delivery means" refers to technology that selects and provides the most suitable interactive activity to the user according to the user's conversation situation and profile.

[0075] "Information organization and distribution means" refers to technology that organizes family event information and health data and accurately notifies users of their devices according to their importance.

[0076] This invention is a system for facilitating smooth communication between family members living far apart. The specific implementation of this system is described below.

[0077] When a user initiates a video call, the device uses its built-in microphone to acquire audio data in real time. During this process, noise cancellation technology is used to reduce external noise and capture clearer audio.

[0078] The terminal encodes the collected audio data and sends it to the server using a secure communication protocol. Commonly used codecs can be used as the encoding technology. Data security is ensured by using encryption protocols such as TLS during this communication.

[0079] The server analyzes the received audio data using a speech recognition algorithm and converts the spoken content into text data. During this process, speaker identification technology analyzes the characteristics of each voice to identify who is speaking. Generally, commercial speech recognition APIs are available for this purpose.

[0080] Furthermore, the server uses the analyzed text data to analyze conversation patterns. The conversation data is stored in a database, and frequent expressions and topics are extracted using natural language processing techniques. By utilizing topic models such as LDA (Latent Dirichlet Allocation), it is possible to understand the trends in conversation.

[0081] Furthermore, the server selects and provides appropriate interactive activities based on each user's conversation status and interests. These selected activities, such as quiz games and puzzles, are designed to be enjoyable for families to participate in together. These activities are designed to be intuitive and easy to use.

[0082] The server organizes family events and health information in a cloud database and notifies devices based on their importance. A general cloud platform can be used to organize this information.

[0083] Through the process described above, it is possible to enrich communication among family members, transcending physical distance. An example of a prompt would be, "Suggest ways to deepen family connections using a remote communication system." This prompt can be used to obtain further insights and suggestions from the generative AI model.

[0084] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0085] Step 1:

[0086] When a user starts a video call, the device automatically activates the microphone and acquires audio data in real time. The input is the user's voice, and the acquired audio data is output as clear audio data after noise reduction using noise cancellation technology. Specifically, the device samples the audio signal at regular intervals and converts it into digital data.

[0087] Step 2:

[0088] The terminal encodes the acquired audio data and sends it to the server using a secure protocol. The input is the unencoded audio data, and the output is the encoded audio data in binary format. Specifically, the terminal applies data encryption technology and establishes a communication channel for transmitting data over the network.

[0089] Step 3:

[0090] The server analyzes the received audio data using a speech recognition algorithm and converts it into text data. The input is encoded audio data, and the output is the analyzed text data. Specifically, the server analyzes the audio waveform using an acoustic model and a language model and maps it to the corresponding text.

[0091] Step 4:

[0092] The server stores text data in a database and analyzes conversation patterns using natural language processing techniques. The input is text data, and the output is metadata indicating conversation patterns. Specifically, the server uses topic modeling techniques to identify topics and frequently occurring expressions in the conversation.

[0093] Step 5:

[0094] The server selects and provides appropriate interactive activities to the user based on the results of conversation analysis. The input is metadata about the conversation pattern, and the output is the game or activity presented to the user. Specifically, the server uses an AI recommendation algorithm to select the optimal game and displays it via the terminal.

[0095] Step 6:

[0096] The server integrates family events and health information, and delivers organized information to devices. The input is a collection of unorganized data, and the output is notification information delivered based on priority. Specifically, the server classifies the data in a cloud database and delivers the information to devices using a push notification system.

[0097] In this way, this system enables efficient and interactive long-distance family communication.

[0098] (Application Example 1)

[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0100] Communication between family members living in different locations can be difficult due to the constraints of physical distance. Furthermore, real-time situation monitoring to ensure family safety and providing education to raise security awareness are crucial. This invention aims to solve these problems and realize safer and more effective communication.

[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0102] In this invention, the server includes voice analysis means for acquiring voice data in real time and identifying speaker characteristics, situation monitoring means for analyzing voice data to monitor security status and issuing alerts when an anomaly is detected, and educational means for learning safety measures through an interactive game. This enables smooth communication among family members in remote locations and the realization of a highly secure living environment.

[0103] "Voice data" refers to digital information used for communication, which records a user's speech in real time.

[0104] "Speaker characteristics" refer to individually identifiable phonetic attributes such as voice quality and vocalization patterns.

[0105] "Voice analysis means" refers to technology used to analyze voice data and identify the speaker.

[0106] "Conversation patterns" are data that shows the flow and content trends of past conversations.

[0107] "Learning methods" refer to the process of acquiring new knowledge by analyzing past conversation history.

[0108] An "interactive game" is dynamic entertainment content that users can directly interact with and participate in.

[0109] "Game delivery methods" refer to technologies that deliver interactive games to users.

[0110] "Family events" refer to plans and occurrences that are shared among family members.

[0111] "Health information" refers to data related to an individual's health status.

[0112] "Information management means" refers to technologies and methods for organizing data and keeping it in a usable state.

[0113] "Information distribution means" refers to technologies for transmitting organized information to users.

[0114] "Situation monitoring means" refers to technologies for analyzing environmental data and detecting anomalies.

[0115] "Educational tools" are methods and devices used to enable users to learn information and skills.

[0116] The system for implementing the present invention primarily functions as real-time analysis of voice data, monitoring of security status, and provision of education through interactive games.

[0117] The terminal acquires voice data in real time when the user engages in voice communication and transmits this data to the server in a highly accurate format. The server analyzes the data using a speech recognition engine to identify the speaker's characteristics and monitors the surrounding environmental sounds. This provides a mechanism to immediately send an alert to the user's terminal if an anomaly or danger is detected.

[0118] The hardware used includes smart glasses and smartphones, and the software employs speech recognition technologies such as Google® Cloud Speech-to-Text. The server stores past conversation history and environmental data in a database and uses a learning algorithm to analyze conversation patterns and sound changes. Based on this analysis, it improves its ability to detect anomalies from everyday patterns that pose a low security risk.

[0119] Interactive games allow users to experience various security scenarios and are developed using game engines such as Unity. Within the game, users can learn security measures in a virtual environment and apply what they learn to real-world actions.

[0120] For example, if a user is walking in a park and suddenly hears a loud noise nearby, the system will analyze the sound, and if it detects an unusual pattern, it will immediately display an alert on the user's smart glasses. Furthermore, it is possible to provide the generating AI model with a prompt such as, "What technologies should be incorporated into daily life to improve family security?" to obtain additional advice and information regarding the system's security features.

[0121] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0122] Step 1:

[0123] When a user initiates voice communication, the device uses its microphone to capture ambient sounds and speech in real time. This results in voice data being input, which is then transmitted to a server via wireless communication.

[0124] Step 2:

[0125] The server analyzes the received audio data through a speech recognition engine to identify the speaker's voiceprint. This process converts the audio data into text information and compares it with voiceprint information in a database to improve the accuracy of speaker identification. The output obtained here is the audio data converted into text information.

[0126] Step 3:

[0127] The system monitors the ambient sounds in the user's environment in real time and detects any deviations from a defined safety pattern as an anomaly. The server uses this input data to perform algorithmic calculations based on the surrounding sound patterns, generating an anomaly detection output. An anomaly alert is then generated based on this output.

[0128] Step 4:

[0129] The server generates an interactive game and sends it to the user's device. Through this, the user practices skills to deal with a virtual security situation. Using a game engine (Unity), simulation data is generated, and the resulting game content is displayed on the user's device.

[0130] Step 5:

[0131] The user inputs security-related prompts into the generated AI model to obtain necessary additional information and advice. The input prompts are sent to the server as text data, the model analyzes their content, generates output as useful recommendations, and presents them to the user.

[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0133] This invention is a family communication support system incorporating an emotion recognition engine. In addition to voice analysis, conversation pattern learning, interactive game provision, and information management and distribution, the emotion recognition function works in conjunction with these other functions. The specific form of the system is described below.

[0134] Forms of speech analysis and emotion recognition

[0135] When a user initiates a video call, the device acquires audio data in real time. The audio data is de-noised, compressed, and securely transmitted to the server. The server uses voice analysis to identify the speaker and an emotion recognition engine to evaluate the voice tone and speech intensity. The emotion recognition engine identifies the user's emotional state based on the audio data and uses this information to support the conversation.

[0136] Conversation pattern learning and forms of emotional response

[0137] The server analyzes past conversation history and sentiment data, learning family-specific conversation patterns using frequency analysis and topic models. This process enables flexible topic suggestions and comment support tailored to the user's emotions. For example, if a user is feeling stressed, it will offer relaxing topics.

[0138] Interactive game delivery methods

[0139] Based on the user's emotional state, as determined by emotion recognition, the server dynamically selects the content of the interactive game. For depressed users, games that encourage or promote a sense of accomplishment are chosen. On the other hand, if a cheerful emotion is recognized, a competitive game is provided, and the difficulty level is adjusted accordingly.

[0140] Information management and distribution methods

[0141] The server aggregates family event information and health data, classifying the organized information based on its importance and the user's current emotional state. Information is delivered to the device at the appropriate time, allowing users to intuitively review and utilize the information. For example, if a user is in a positive mood, notifications encouraging participation in new events are prioritized.

[0142] In this way, the present invention utilizes an emotion engine to build a system that supports rich dialogue and information sharing among family members, enabling communication that senses emotions even remotely.

[0143] The following describes the processing flow.

[0144] Step 1:

[0145] When a user initiates a video call, the device acquires audio data in real time through the microphone. This audio data is then noise-filtered to ensure clear data is obtained.

[0146] Step 2:

[0147] The terminal compresses the acquired audio data to efficiently utilize network bandwidth and then securely transmits it to the server.

[0148] Step 3:

[0149] The server uses a speech recognition algorithm to identify the speaker in order to analyze the received audio data. This process identifies speaker-specific vocal features and voiceprints.

[0150] Step 4:

[0151] The server uses an emotion recognition engine to evaluate the user's emotional state based on their tone of voice and word choice. This allows it to identify emotions such as happiness, sadness, and anger in real time.

[0152] Step 5:

[0153] The server integrates past conversation history and sentiment data, and uses frequency analysis and topic models to learn conversation patterns between users. Based on the user's sentiment history, it can suggest topics that take into account predicted responses.

[0154] Step 6:

[0155] The server selects an appropriate interactive game based on emotional data. For example, if a user is feeling down, it will choose a game that is likely to elicit positive feedback. The server then notifies the user of the selected game.

[0156] Step 7:

[0157] The device displays game information sent from the server to the user and offers the option to participate in the game. If the user consents, the game starts and the progress is displayed in real time.

[0158] Step 8:

[0159] The server collects important family events and health information, organizing and prioritizing it. In particular, it takes the user's emotional state into consideration when deciding which information to notify them of first.

[0160] Step 9:

[0161] The server sends organized information as a push notification to the user's device, enabling a quick response. Users can then easily check this information on their device and take the necessary actions.

[0162] (Example 2)

[0163] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0164] In today's digital society, communication between family members living far apart is becoming increasingly difficult. In particular, accurately conveying and sharing emotions is challenging, necessitating richer and more effective communication methods. Furthermore, the inability to offer appropriate information tailored to the context of the conversation and individual emotional states can lead to a decline in the quality of communication.

[0165] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0166] In this invention, the server includes voice analysis means for instantly acquiring voice information and identifying the characteristics of the speaker, analysis means for analyzing past dialogue history and learning dialogue patterns, and emotion evaluation means for evaluating the emotional state of the speaker using an emotion recognition engine. This enables emotionally responsive communication between family members in remote locations and realizes flexible and appropriate dialogue suggestions tailored to the user's emotions.

[0167] "Voice information" refers to speech data collected from users, and voice analysis and emotion recognition are performed based on this data.

[0168] "Instant acquisition" means that voice information is collected in real time, with the aim of responding to user speech without delay.

[0169] "Speaker characteristics" refers to the individual voice and speaking style characteristics of each user, identified based on audio data.

[0170] "Speech analysis means" refers to technical methods for analyzing speech information to identify the characteristics of the speaker and the speaker themselves.

[0171] "Dialogue history" refers to a record of past user-to-user conversation data, and dialogue patterns are learned based on this.

[0172] "Analysis methods" refer to technical techniques used to analyze specific patterns and characteristics using collected data.

[0173] An "emotion recognition engine" refers to a program or system that has the function of identifying and evaluating a user's emotional state based on voice information.

[0174] "Emotional evaluation method" refers to a technical method for evaluating a user's emotional state based on collected audio information.

[0175] An "interactive game" refers to an entertainment activity that allows for active interaction with the user, and whose content is adjusted according to the user's emotional state.

[0176] "Entertainment provision means" refers to methods and technologies for providing interactive games to users.

[0177] "Information organization methods" refer to technical techniques for organizing collected information about household events and health-related matters.

[0178] "Information transmission means" refers to methods and technologies for notifying users of organized information to their terminals in a timely manner.

[0179] "Suggestion generation means" refers to programs and technologies for dynamically generating personalized dialogue and entertainment suggestions for users.

[0180] This invention is implemented as a family communication support system incorporating an emotion recognition engine. The system operates in conjunction with functions for voice analysis, conversation pattern learning, interactive entertainment provision, and information management and distribution. Specific embodiments are described below.

[0181] Voice analysis and emotion recognition

[0182] When a user initiates a voice call, the device uses its microphone to acquire voice information in real time. This voice data is then denoised, compressed, encrypted, and transmitted to a server via the internet. The server uses voice analysis tools to identify the speaker's characteristics and an emotion recognition engine to analyze the tone and intensity of speech to determine the user's emotional state. This process enables emotionally expressive communication between users.

[0183] Conversation pattern learning and suggestion generation

[0184] The server analyzes past conversation history and current emotional state. Based on the dialogue history stored in the database, it performs frequency analysis and topic modeling to learn unique dialogue patterns between users. Using this information, the server leverages a generative AI model to dynamically generate flexible and accurate dialogue and entertainment suggestions tailored to the user.

[0185] Interactive entertainment provider

[0186] Based on the emotional state identified by the emotion recognition engine, the server selects the most appropriate interactive entertainment. If the user is in a cheerful state, a competitive game is provided to the device, offering a higher level of challenge. Conversely, if factors indicating a need for relaxation are detected, a game with encouraging content is selected.

[0187] Information management and distribution

[0188] The server aggregates and organizes family event information and health data. The information is categorized according to the user's emotional state and notified to the user's device at the appropriate time through various communication channels. For example, users with positive emotions receive priority notifications of invitations to new events.

[0189] One concrete example is video chat between family members. For instance, in a scene where a grandchild and grandfather are talking, voice analysis can recognize the grandchild's enjoyment from their voice and suggest competitive games or other activities.

[0190] Examples of prompts to input into the generative AI model include, "Think about how to help a user relax when they are feeling stressed." This allows for natural and meaningful interactions that respond to the user's emotions.

[0191] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0192] Step 1:

[0193] The device instantly acquires audio data using the microphone as soon as the user initiates a video call. The acquired audio data is then filtered to remove background noise. Subsequently, data compression technology is used to reduce the size of the audio data, making it suitable for efficient transmission. The input to this process is the user's raw voice, and the output is compressed and encrypted audio data. The encrypted audio data is sent to the server through a secure communication channel.

[0194] Step 2:

[0195] The server receives audio data transmitted from the terminal and identifies the speaker's characteristics using speech analysis means. In this step, a speech recognition algorithm is used to characterize the speaker's voiceprint from the audio data. The input is encrypted audio data, and the output is the analyzed speaker characteristics. Simultaneously, an emotion recognition engine is interfaced to identify the user's emotional state by analyzing the voice tone and speech intensity.

[0196] Step 3:

[0197] The server performs analysis using previously collected dialogue history and current sentiment data. This process includes frequency analysis and topic modeling to learn unique dialogue patterns between users. The input is past dialogue history and sentiment data, and the output is the learned conversation patterns. Based on this data, the server utilizes a generative AI model to construct dialogues and entertainment suggestions that match the user's current situation.

[0198] Step 4:

[0199] The server selects the most appropriate interactive entertainment or conversation based on information obtained from the emotion recognition engine. It prompts a generative AI model, which automatically generates content tailored to the user's emotions. For example, if the user is feeling down, it will select an uplifting game or a relaxing topic. The input for this step is the user's emotion data, and the output is the content of the entertainment or conversation provided to the user.

[0200] Step 5:

[0201] The server collects organized family event information and health-related data, and uses information management and distribution tools to notify the user of the identified information. The timing of this notification is determined by the user's emotional state and the importance of the information. For example, when the user is in a positive emotional state, an invitation to participate in a new event will be sent to the device. The inputs to this process are family data and emotional state, and the output is the notification information sent to the device.

[0202] (Application Example 2)

[0203] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0204] In modern society, while the demand for online shopping and virtual experiences is increasing, there is a lack of customized product recommendations tailored to each user's individual emotional state and preferences. As a result, users may miss out on products and services they truly need, leading to decreased satisfaction with their purchasing experience. Therefore, there is a need to develop a system that can recognize a user's emotional state in real time and provide appropriate product recommendations based on that understanding.

[0205] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0206] In this invention, the server includes voice analysis means for acquiring voice data in real time and identifying the user's characteristics, means for making product suggestions based on the user's emotional state using an emotion recognition engine, and information distribution means for notifying the user terminal of the organized information. This makes it possible to make product suggestions based on the user's emotional state.

[0207] "Voice data" refers to information obtained by digitally capturing the user's voice in real time and processing it using a speech recognition algorithm.

[0208] "User characteristics" refers to information that describes the voice characteristics and patterns used to identify individual users.

[0209] "Voice analysis means" refers to a technical method for acquiring voice data in real time and identifying the characteristics of the user.

[0210] "Past conversation history" refers to information that records previous conversations between the user and others.

[0211] "Learning methods" refer to technical techniques for analyzing past conversation history and learning conversation patterns.

[0212] "Interactive entertainment" refers to entertainment content that actively engages the user and dynamically changes based on emotional recognition.

[0213] "Means of providing entertainment" refers to technical methods that provide interactive entertainment and allow users to experience it interactively.

[0214] "Domestic events" is a concept that refers to events and activities that occur among family members and close friends.

[0215] "Health information" refers to information about the health status of individual users or people within their household.

[0216] "Information management methods" refer to technical techniques for collecting and organizing information about events and health within the household.

[0217] "User terminal" refers to an electronic device used by a user to receive and operate information.

[0218] "Information distribution means" refers to technical methods for appropriately notifying users of organized information at their terminals.

[0219] An "emotion recognition engine" is an algorithm that identifies the user's emotional state based on factors such as voice tone and speech intensity.

[0220] "Product recommendation" is the act of recommending appropriate products or services based on the user's emotional state.

[0221] This invention is realized by acquiring voice data in real time using a terminal held by the user and processing that data using a voice analysis means. The terminal first has the function of converting the user's voice into digital data and securely transmitting it to a server. The server receives this voice data, removes noise, and then performs voice analysis. As a result, the user's characteristics are identified, and the user's emotional state is determined by an emotion recognition engine.

[0222] The server also analyzes past conversation history and learns user-specific conversation patterns using frequency analysis and topic models. Based on this, it can provide product recommendations that are best suited to the user's emotions. Product data is managed in a cloud-based database, and recommendations are made in real time, taking into account the user's emotional state.

[0223] For example, if a positive emotion is identified by emotion recognition while a user is shopping, products and services with a high affinity for that emotion can be recommended. The user can then use their smartphone to review and purchase the suggested products. An example of a prompt message would be, "Based on the user's current emotional state, please suggest the most suitable products for them," which could be given to the generative AI model.

[0224] In this way, a system that integrates emotion recognition and voice analysis technologies can enhance the user's purchasing experience. The system utilizes a communication network, a cloud-based database, a voice analysis algorithm, and an emotion recognition engine. This enables the provision of a personalized purchasing experience to users, thereby increasing satisfaction.

[0225] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0226] Step 1:

[0227] The device acquires the user's voice via a microphone, converts that voice into digital data, and sends it to the server. The input is the user's spoken voice, and the output is digital audio data. The device uses internet communication for data acquisition and transmission.

[0228] Step 2:

[0229] The server removes noise from the received audio data and analyzes the user's characteristics using a speech recognition algorithm. The input is audio data transmitted from the terminal, and the output is the identified user's voice profile. The speech recognition algorithm utilizes a cloud-based service to identify each user's voice characteristics.

[0230] Step 3:

[0231] The server operates an emotion recognition engine to identify the user's emotional state by evaluating voice tone and speech intensity. The input is a voice profile based on voice data, and the output is the user's emotional state. In this process, the emotion recognition engine analyzes the tone and intensity to determine characteristic emotions.

[0232] Step 4:

[0233] The server learns conversation patterns using recorded past conversation history and sentiment data, and generates frequency analyses and topic models. Inputs include past conversation history and current sentiment states, and output is the learned conversation patterns. Machine learning techniques are used to analyze the data.

[0234] Step 5:

[0235] The server retrieves product data from a cloud-based database and suggests the most suitable product based on the user's emotional state. The input is the user's emotional state and conversation patterns, and the output is a list of suggested products. The prompt is passed to a generating AI model, and the product recommendation algorithm selects the most suitable product for the user.

[0236] Step 6:

[0237] The server notifies the terminal of the suggested product information, allowing the user to view the suggestions. The input is the list of suggested products, and the output is the product information displayed on the user's terminal. The notification function uses push notifications, allowing the user to instantly check the information on their smartphone.

[0238] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0239] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0240] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0241] [Second Embodiment]

[0242] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0243] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0244] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0245] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0246] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0247] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0248] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0249] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0250] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0251] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0252] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0253] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0254] This invention provides a system that facilitates smooth communication between distant family members and includes functions for voice analysis, conversation pattern learning, interactive games, and information organization and distribution. The system operates through the interaction of a server, a terminal, and a user.

[0255] Forms of speech analysis

[0256] When a user initiates a video call, the device uses its microphone to acquire audio data in real time. This allows the device to efficiently capture the user's speech and transmit highly accurate audio data to the server. The server analyzes the received audio data using a speech recognition algorithm to identify the speaker and learn the unique voice characteristics of each user, enabling it to accurately understand the conversation.

[0257] Forms of learning conversation patterns

[0258] The server accumulates past conversation history and learns conversation patterns using frequency analysis and topic models based on that data. This allows it to understand the content of what was said and the patterns of the conversation flow, and to utilize the user's unique communication style in future conversations.

[0259] Interactive game delivery methods

[0260] The server selects an appropriate interactive game based on the conversation context and participants' interests. The game is displayed to the user via their device and is designed for intuitive operation. Specifically, it supports the user's participation in a quiz game, guiding them through the game's progression to elicit smart responses.

[0261] Forms of information organization and distribution

[0262] The server aggregates family event information and health data, organizing important information through an information management system. This information is then categorized based on urgency and importance, and notified to terminals using an information distribution system. This notification allows users to quickly take necessary actions.

[0263] In this way, the system of the present invention can efficiently carry out a series of processes from voice recognition to learning and information provision, thereby enriching communication among family members that transcends physical distance.

[0264] The following describes the processing flow.

[0265] Step 1:

[0266] When a user initiates a video call, the device immediately acquires audio data using its built-in microphone. This audio data is then processed to ensure clarity using noise-canceling technology.

[0267] Step 2:

[0268] The terminal compresses the acquired audio data in real time and sends it to the server via a secure protocol. The data is transferred using a highly efficient encoding method, minimizing latency.

[0269] Step 3:

[0270] The server inputs the received audio data into a speech recognition algorithm for analysis. Here, speech features are extracted to identify each speaker, and speaker identification is performed.

[0271] Step 4:

[0272] The server saves the identified speaker's voice characteristics to a cloud database and updates the dataset used for subsequent conversation analysis. This data is also protected by security protocols.

[0273] Step 5:

[0274] The server performs frequency analysis based on past conversation history to extract characteristic conversation patterns. It then uses topic models to analyze the flow of topics and the relationships between important keywords.

[0275] Step 6:

[0276] Based on learned conversation patterns, the server prepares to naturally suggest relevant information and topics in the next conversation, thereby helping to keep the conversation flowing smoothly.

[0277] Step 7:

[0278] The server selects an appropriate interactive game based on the conversation content and the user's interests, and prepares for the play. The past game history of the user in the database is also utilized for this selection.

[0279] Step 8:

[0280] The terminal presents the interactive game provided by the server to the user and encourages the user to participate in the game. When the consent of the user is obtained, the game is started, and the leaderboard and score display functions are enabled.

[0281] Step 9:

[0282] The server aggregates and organizes the family's calendar events and health information. These information are classified according to the priority, and those relevant to each user are selected.

[0283] Step 10:

[0284] The server notifies the terminal of the organized information at the necessary timing to assist the user to respond promptly. Since each notification is sent as a push notification, it can be delivered to the user immediately.

[0285] (Example 1)

[0286] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0287] It is an issue to realize smooth and intimate communication among family members who are at a distance. In particular, there is a need for a technical solution to prevent the thinning of communication due to the physical distance and deepen the bond of the family. Furthermore, it is also essential to improve the recognition accuracy of voice and provide information corresponding to individual communication styles.

[0288] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0289] In this invention, the server includes voice acquisition means, data transmission means, data analysis means, conversation analysis means, activity provision means, and information organization and distribution means. This enables real-time and effective communication that transcends physical distance and allows for the provision of interactive support to strengthen family bonds.

[0290] "Voice acquisition means" refers to technology for accurately acquiring the user's voice and collecting clear audio data by reducing external noise.

[0291] "Data transmission means" refers to technology for encoding acquired audio data and sending it to a server using a secure communication protocol.

[0292] "Data analysis means" refers to a technology that analyzes audio data received on a server using speech recognition technology, converts the spoken content into text, and identifies individual speakers.

[0293] A "conversation analysis method" is a technique that uses natural language processing technology to learn conversation patterns based on accumulated conversation data and identify frequently occurring topics and expressions.

[0294] "Activity delivery means" refers to technology that selects and provides the most suitable interactive activity to the user according to the user's conversation situation and profile.

[0295] "Information organization and distribution means" refers to technology that organizes family event information and health data and accurately notifies users of their devices according to their importance.

[0296] This invention is a system for facilitating smooth communication between family members living far apart. The specific implementation of this system is described below.

[0297] When a user initiates a video call, the device uses its built-in microphone to acquire audio data in real time. During this process, noise cancellation technology is used to reduce external noise and capture clearer audio.

[0298] The terminal encodes the collected audio data and sends it to the server using a secure communication protocol. Commonly used codecs can be used as the encoding technology. Data security is ensured by using encryption protocols such as TLS during this communication.

[0299] The server analyzes the received audio data using a speech recognition algorithm and converts the spoken content into text data. During this process, speaker identification technology analyzes the characteristics of each voice to identify who is speaking. Generally, commercial speech recognition APIs are available for this purpose.

[0300] Furthermore, the server uses the analyzed text data to analyze conversation patterns. The conversation data is stored in a database, and frequent expressions and topics are extracted using natural language processing techniques. By utilizing topic models such as LDA (Latent Dirichlet Allocation), it is possible to understand the trends in conversation.

[0301] Furthermore, the server selects and provides appropriate interactive activities based on each user's conversation status and interests. These selected activities, such as quiz games and puzzles, are designed to be enjoyable for families to participate in together. These activities are designed to be intuitive and easy to use.

[0302] The server organizes family events and health information in a cloud database and notifies devices based on their importance. A general cloud platform can be used to organize this information.

[0303] Through the above process, it is possible to enrich family communication beyond physical distance. As an example of a prompt sentence, "Propose a method for deepening family connection in a remote communication system" can be considered. Using this prompt, further insights and proposals can be obtained from the generative AI model.

[0304] The flow of the specific process in Example 1 will be described using FIG. 11.

[0305] Step 1:

[0306] When the user starts a video call, the terminal automatically activates the microphone and acquires voice data in real time. The input is the user's voice, and the acquired voice data is output as clear voice data after reducing noise using noise cancellation technology. Specifically, the terminal samples the voice signal at regular intervals and converts it into digital data.

[0307] Step 2:

[0308] The terminal encodes the acquired voice data and transmits it to the server using a secure protocol. The input is the voice data before encoding, and the output is the encoded voice data in binary format. As a specific operation, the terminal applies data encryption technology to establish a communication channel for transmitting data through the network.

[0309] Step 3:

[0310] The server analyzes the received voice data using a voice recognition algorithm and converts it into text data. The input is the encoded voice data, and the output is the analyzed text data. Specifically, the server uses an acoustic model and a language model to analyze the waveform of the voice and perform a process of mapping it to the corresponding text.

[0311] Step 4:

[0312] The server stores text data in a database and analyzes conversation patterns using natural language processing techniques. The input is text data, and the output is metadata indicating conversation patterns. Specifically, the server uses topic modeling techniques to identify topics and frequently occurring expressions in the conversation.

[0313] Step 5:

[0314] The server selects and provides appropriate interactive activities to the user based on the results of conversation analysis. The input is metadata about the conversation pattern, and the output is the game or activity presented to the user. Specifically, the server uses an AI recommendation algorithm to select the optimal game and displays it via the terminal.

[0315] Step 6:

[0316] The server integrates family events and health information, and delivers organized information to devices. The input is a collection of unorganized data, and the output is notification information delivered based on priority. Specifically, the server classifies the data in a cloud database and delivers the information to devices using a push notification system.

[0317] In this way, this system enables efficient and interactive long-distance family communication.

[0318] (Application Example 1)

[0319] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0320] Communication between family members living in different locations can be difficult due to the constraints of physical distance. Furthermore, real-time situation monitoring to ensure family safety and providing education to raise security awareness are crucial. This invention aims to solve these problems and realize safer and more effective communication.

[0321] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0322] In this invention, the server includes voice analysis means for acquiring voice data in real time and identifying speaker characteristics, situation monitoring means for analyzing voice data to monitor security status and issuing alerts when an anomaly is detected, and educational means for learning safety measures through an interactive game. This enables smooth communication among family members in remote locations and the realization of a highly secure living environment.

[0323] "Voice data" refers to digital information used for communication, which records a user's speech in real time.

[0324] "Speaker characteristics" refer to individually identifiable phonetic attributes such as voice quality and vocalization patterns.

[0325] "Voice analysis means" refers to technology used to analyze voice data and identify the speaker.

[0326] "Conversation patterns" are data that shows the flow and content trends of past conversations.

[0327] "Learning methods" refer to the process of acquiring new knowledge by analyzing past conversation history.

[0328] An "interactive game" is dynamic entertainment content that users can directly interact with and participate in.

[0329] "Game delivery methods" refer to technologies that deliver interactive games to users.

[0330] "Family events" refer to plans and occurrences that are shared among family members.

[0331] "Health information" refers to data related to an individual's health status.

[0332] "Information management means" refers to technologies and methods for organizing data and keeping it in a usable state.

[0333] "Information distribution means" refers to technologies for transmitting organized information to users.

[0334] "Situation monitoring means" refers to technologies for analyzing environmental data and detecting anomalies.

[0335] "Educational tools" are methods and devices used to enable users to learn information and skills.

[0336] The system for implementing the present invention primarily functions as real-time analysis of voice data, monitoring of security status, and provision of education through interactive games.

[0337] The terminal acquires voice data in real time when the user engages in voice communication and transmits this data to the server in a highly accurate format. The server analyzes the data using a speech recognition engine to identify the speaker's characteristics and monitors the surrounding environmental sounds. This provides a mechanism to immediately send an alert to the user's terminal if an anomaly or danger is detected.

[0338] The hardware used includes smart glasses and smartphones, while the software employs speech recognition technologies such as Google Cloud Speech-to-Text. The server stores past conversation history and environmental data in a database and uses a learning algorithm to analyze conversation patterns and sound changes. Based on this analysis, the system improves its ability to detect anomalies from everyday patterns that pose a low security risk.

[0339] Interactive games allow users to experience various security scenarios and are developed using game engines such as Unity. Within the game, users can learn security measures in a virtual environment and apply what they learn to real-world actions.

[0340] For example, if a user is walking in a park and suddenly hears a loud noise nearby, the system will analyze the sound, and if it detects an unusual pattern, it will immediately display an alert on the user's smart glasses. Furthermore, it is possible to provide the generating AI model with a prompt such as, "What technologies should be incorporated into daily life to improve family security?" to obtain additional advice and information regarding the system's security features.

[0341] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0342] Step 1:

[0343] When a user initiates voice communication, the device uses its microphone to capture ambient sounds and speech in real time. This results in voice data being input, which is then transmitted to a server via wireless communication.

[0344] Step 2:

[0345] The server analyzes the received audio data through a speech recognition engine to identify the speaker's voiceprint. This process converts the audio data into text information and compares it with voiceprint information in a database to improve the accuracy of speaker identification. The output obtained here is the audio data converted into text information.

[0346] Step 3:

[0347] The system monitors the ambient sounds in the user's environment in real time and detects any deviations from a defined safety pattern as an anomaly. The server uses this input data to perform algorithmic calculations based on the surrounding sound patterns, generating an anomaly detection output. An anomaly alert is then generated based on this output.

[0348] Step 4:

[0349] The server generates an interactive game and sends it to the user's device. Through this, the user practices skills to deal with a virtual security situation. Using a game engine (Unity), simulation data is generated, and the resulting game content is displayed on the user's device.

[0350] Step 5:

[0351] The user inputs security-related prompts into the generated AI model to obtain necessary additional information and advice. The input prompts are sent to the server as text data, the model analyzes their content, generates output as useful recommendations, and presents them to the user.

[0352] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0353] This invention is a family communication support system incorporating an emotion recognition engine. In addition to voice analysis, conversation pattern learning, interactive game provision, and information management and distribution, the emotion recognition function works in conjunction with these other functions. The specific form of the system is described below.

[0354] Forms of speech analysis and emotion recognition

[0355] When a user initiates a video call, the device acquires audio data in real time. The audio data is de-noised, compressed, and securely transmitted to the server. The server uses voice analysis to identify the speaker and an emotion recognition engine to evaluate the voice tone and speech intensity. The emotion recognition engine identifies the user's emotional state based on the audio data and uses this information to support the conversation.

[0356] Conversation pattern learning and forms of emotional response

[0357] The server analyzes past conversation history and sentiment data, learning family-specific conversation patterns using frequency analysis and topic models. This process enables flexible topic suggestions and comment support tailored to the user's emotions. For example, if a user is feeling stressed, it will offer relaxing topics.

[0358] Interactive game delivery methods

[0359] Based on the user's emotional state, as determined by emotion recognition, the server dynamically selects the content of the interactive game. For depressed users, games that encourage or promote a sense of accomplishment are chosen. On the other hand, if a cheerful emotion is recognized, a competitive game is provided, and the difficulty level is adjusted accordingly.

[0360] Information management and distribution methods

[0361] The server aggregates family event information and health data, classifying the organized information based on its importance and the user's current emotional state. Information is delivered to the device at the appropriate time, allowing users to intuitively review and utilize the information. For example, if a user is in a positive mood, notifications encouraging participation in new events are prioritized.

[0362] In this way, the present invention utilizes an emotion engine to build a system that supports rich dialogue and information sharing among family members, enabling communication that senses emotions even remotely.

[0363] The following describes the processing flow.

[0364] Step 1:

[0365] When a user initiates a video call, the device acquires audio data in real time through the microphone. This audio data is then noise-filtered to ensure clear data is obtained.

[0366] Step 2:

[0367] The terminal compresses the acquired audio data to efficiently utilize network bandwidth and then securely transmits it to the server.

[0368] Step 3:

[0369] The server uses a speech recognition algorithm to identify the speaker in order to analyze the received audio data. This process identifies speaker-specific vocal features and voiceprints.

[0370] Step 4:

[0371] The server uses an emotion recognition engine to evaluate the user's emotional state based on their tone of voice and word choice. This allows it to identify emotions such as happiness, sadness, and anger in real time.

[0372] Step 5:

[0373] The server integrates past conversation history and sentiment data, and uses frequency analysis and topic models to learn conversation patterns between users. Based on the user's sentiment history, it can suggest topics that take into account predicted responses.

[0374] Step 6:

[0375] The server selects an appropriate interactive game based on emotional data. For example, if a user is feeling down, it will choose a game that is likely to elicit positive feedback. The server then notifies the user of the selected game.

[0376] Step 7:

[0377] The device displays game information sent from the server to the user and offers the option to participate in the game. If the user consents, the game starts and the progress is displayed in real time.

[0378] Step 8:

[0379] The server collects important family events and health information, organizing and prioritizing it. In particular, it takes the user's emotional state into consideration when deciding which information to notify them of first.

[0380] Step 9:

[0381] The server sends organized information as a push notification to the user's device, enabling a quick response. Users can then easily check this information on their device and take the necessary actions.

[0382] (Example 2)

[0383] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0384] In today's digital society, communication between family members living far apart is becoming increasingly difficult. In particular, accurately conveying and sharing emotions is challenging, necessitating richer and more effective communication methods. Furthermore, the inability to offer appropriate information tailored to the context of the conversation and individual emotional states can lead to a decline in the quality of communication.

[0385] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0386] In this invention, the server includes voice analysis means for instantly acquiring voice information and identifying the characteristics of the speaker, analysis means for analyzing past dialogue history and learning dialogue patterns, and emotion evaluation means for evaluating the emotional state of the speaker using an emotion recognition engine. This enables emotionally responsive communication between family members in remote locations and realizes flexible and appropriate dialogue suggestions tailored to the user's emotions.

[0387] "Voice information" refers to speech data collected from users, and voice analysis and emotion recognition are performed based on this data.

[0388] "Instant acquisition" means that voice information is collected in real time, with the aim of responding to user speech without delay.

[0389] "Speaker characteristics" refers to the individual voice and speaking style characteristics of each user, identified based on audio data.

[0390] "Speech analysis means" refers to technical methods for analyzing speech information to identify the characteristics of the speaker and the speaker themselves.

[0391] "Dialogue history" refers to a record of past user-to-user conversation data, and dialogue patterns are learned based on this.

[0392] "Analysis methods" refer to technical techniques used to analyze specific patterns and characteristics using collected data.

[0393] An "emotion recognition engine" refers to a program or system that has the function of identifying and evaluating a user's emotional state based on voice information.

[0394] "Emotional evaluation method" refers to a technical method for evaluating a user's emotional state based on collected audio information.

[0395] An "interactive game" refers to an entertainment activity that allows for active interaction with the user, and whose content is adjusted according to the user's emotional state.

[0396] "Entertainment provision means" refers to methods and technologies for providing interactive games to users.

[0397] "Information organization methods" refer to technical techniques for organizing collected information about household events and health-related matters.

[0398] "Information transmission means" refers to methods and technologies for notifying users of organized information to their terminals in a timely manner.

[0399] "Suggestion generation means" refers to programs and technologies for dynamically generating personalized dialogue and entertainment suggestions for users.

[0400] This invention is implemented as a family communication support system incorporating an emotion recognition engine. The system operates in conjunction with functions for voice analysis, conversation pattern learning, interactive entertainment provision, and information management and distribution. Specific embodiments are described below.

[0401] Voice analysis and emotion recognition

[0402] When a user initiates a voice call, the device uses its microphone to acquire voice information in real time. This voice data is then denoised, compressed, encrypted, and transmitted to a server via the internet. The server uses voice analysis tools to identify the speaker's characteristics and an emotion recognition engine to analyze the tone and intensity of speech to determine the user's emotional state. This process enables emotionally expressive communication between users.

[0403] Conversation pattern learning and suggestion generation

[0404] The server analyzes past conversation history and current emotional state. Based on the dialogue history stored in the database, it performs frequency analysis and topic modeling to learn unique dialogue patterns between users. Using this information, the server leverages a generative AI model to dynamically generate flexible and accurate dialogue and entertainment suggestions tailored to the user.

[0405] Interactive entertainment provider

[0406] Based on the emotional state identified by the emotion recognition engine, the server selects the most appropriate interactive entertainment. If the user is in a cheerful state, a competitive game is provided to the device, offering a higher level of challenge. Conversely, if factors indicating a need for relaxation are detected, a game with encouraging content is selected.

[0407] Information management and distribution

[0408] The server aggregates and organizes family event information and health data. The information is categorized according to the user's emotional state and notified to the user's device at the appropriate time through various communication channels. For example, users with positive emotions receive priority notifications of invitations to new events.

[0409] One concrete example is video chat between family members. For instance, in a scene where a grandchild and grandfather are talking, voice analysis can recognize the grandchild's enjoyment from their voice and suggest competitive games or other activities.

[0410] Examples of prompts to input into the generative AI model include, "Think about how to help a user relax when they are feeling stressed." This allows for natural and meaningful interactions that respond to the user's emotions.

[0411] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0412] Step 1:

[0413] The device instantly acquires audio data using the microphone as soon as the user initiates a video call. The acquired audio data is then filtered to remove background noise. Subsequently, data compression technology is used to reduce the size of the audio data, making it suitable for efficient transmission. The input to this process is the user's raw voice, and the output is compressed and encrypted audio data. The encrypted audio data is sent to the server through a secure communication channel.

[0414] Step 2:

[0415] The server receives audio data transmitted from the terminal and identifies the speaker's characteristics using speech analysis means. In this step, a speech recognition algorithm is used to characterize the speaker's voiceprint from the audio data. The input is encrypted audio data, and the output is the analyzed speaker characteristics. Simultaneously, an emotion recognition engine is interfaced to identify the user's emotional state by analyzing the voice tone and speech intensity.

[0416] Step 3:

[0417] The server performs analysis using previously collected dialogue history and current sentiment data. This process includes frequency analysis and topic modeling to learn unique dialogue patterns between users. The input is past dialogue history and sentiment data, and the output is the learned conversation patterns. Based on this data, the server utilizes a generative AI model to construct dialogues and entertainment suggestions that match the user's current situation.

[0418] Step 4:

[0419] The server selects the most appropriate interactive entertainment or conversation based on information obtained from the emotion recognition engine. It prompts a generative AI model, which automatically generates content tailored to the user's emotions. For example, if the user is feeling down, it will select an uplifting game or a relaxing topic. The input for this step is the user's emotion data, and the output is the content of the entertainment or conversation provided to the user.

[0420] Step 5:

[0421] The server collects organized family event information and health-related data, and uses information management and distribution tools to notify the user of the identified information. The timing of this notification is determined by the user's emotional state and the importance of the information. For example, when the user is in a positive emotional state, an invitation to participate in a new event will be sent to the device. The inputs to this process are family data and emotional state, and the output is the notification information sent to the device.

[0422] (Application Example 2)

[0423] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0424] In modern society, while the demand for online shopping and virtual experiences is increasing, there is a lack of customized product recommendations tailored to each user's individual emotional state and preferences. As a result, users may miss out on products and services they truly need, leading to decreased satisfaction with their purchasing experience. Therefore, there is a need to develop a system that can recognize a user's emotional state in real time and provide appropriate product recommendations based on that understanding.

[0425] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0426] In this invention, the server includes voice analysis means for acquiring voice data in real time and identifying the user's characteristics, means for making product suggestions based on the user's emotional state using an emotion recognition engine, and information distribution means for notifying the user terminal of the organized information. This makes it possible to make product suggestions based on the user's emotional state.

[0427] "Voice data" refers to information obtained by digitally capturing the user's voice in real time and processing it using a speech recognition algorithm.

[0428] "User characteristics" refers to information that describes the voice characteristics and patterns used to identify individual users.

[0429] "Voice analysis means" refers to a technical method for acquiring voice data in real time and identifying the characteristics of the user.

[0430] "Past conversation history" refers to information that records previous conversations between the user and others.

[0431] "Learning methods" refer to technical techniques for analyzing past conversation history and learning conversation patterns.

[0432] "Interactive entertainment" refers to entertainment content that actively engages the user and dynamically changes based on emotional recognition.

[0433] "Means of providing entertainment" refers to technical methods that provide interactive entertainment and allow users to experience it interactively.

[0434] "Domestic events" is a concept that refers to events and activities that occur among family members and close friends.

[0435] "Health information" refers to information about the health status of individual users or people within their household.

[0436] "Information management methods" refer to technical techniques for collecting and organizing information about events and health within the household.

[0437] "User terminal" refers to an electronic device used by a user to receive and operate information.

[0438] "Information distribution means" refers to technical methods for appropriately notifying users of organized information at their terminals.

[0439] An "emotion recognition engine" is an algorithm that identifies the user's emotional state based on factors such as voice tone and speech intensity.

[0440] "Product recommendation" is the act of recommending appropriate products or services based on the user's emotional state.

[0441] This invention is realized by acquiring voice data in real time using a terminal held by the user and processing that data using a voice analysis means. The terminal first has the function of converting the user's voice into digital data and securely transmitting it to a server. The server receives this voice data, removes noise, and then performs voice analysis. As a result, the user's characteristics are identified, and the user's emotional state is determined by an emotion recognition engine.

[0442] The server also analyzes past conversation history and learns user-specific conversation patterns using frequency analysis and topic models. Based on this, it can provide product recommendations that are best suited to the user's emotions. Product data is managed in a cloud-based database, and recommendations are made in real time, taking into account the user's emotional state.

[0443] For example, if a positive emotion is identified by emotion recognition while a user is shopping, products and services with a high affinity for that emotion can be recommended. The user can then use their smartphone to review and purchase the suggested products. An example of a prompt message would be, "Based on the user's current emotional state, please suggest the most suitable products for them," which could be given to the generative AI model.

[0444] In this way, a system that integrates emotion recognition and voice analysis technologies can enhance the user's purchasing experience. The system utilizes a communication network, a cloud-based database, a voice analysis algorithm, and an emotion recognition engine. This enables the provision of a personalized purchasing experience to users, thereby increasing satisfaction.

[0445] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0446] Step 1:

[0447] The device acquires the user's voice via a microphone, converts that voice into digital data, and sends it to the server. The input is the user's spoken voice, and the output is digital audio data. The device uses internet communication for data acquisition and transmission.

[0448] Step 2:

[0449] The server removes noise from the received audio data and analyzes the user's characteristics using a speech recognition algorithm. The input is audio data transmitted from the terminal, and the output is the identified user's voice profile. The speech recognition algorithm utilizes a cloud-based service to identify each user's voice characteristics.

[0450] Step 3:

[0451] The server operates an emotion recognition engine to identify the user's emotional state by evaluating voice tone and speech intensity. The input is a voice profile based on voice data, and the output is the user's emotional state. In this process, the emotion recognition engine analyzes the tone and intensity to determine characteristic emotions.

[0452] Step 4:

[0453] The server learns conversation patterns using recorded past conversation history and sentiment data, and generates frequency analyses and topic models. Inputs include past conversation history and current sentiment states, and output is the learned conversation patterns. Machine learning techniques are used to analyze the data.

[0454] Step 5:

[0455] The server retrieves product data from a cloud-based database and suggests the most suitable product based on the user's emotional state. The input is the user's emotional state and conversation patterns, and the output is a list of suggested products. The prompt is passed to a generating AI model, and the product recommendation algorithm selects the most suitable product for the user.

[0456] Step 6:

[0457] The server notifies the terminal of the suggested product information, allowing the user to view the suggestions. The input is the list of suggested products, and the output is the product information displayed on the user's terminal. The notification function uses push notifications, allowing the user to instantly check the information on their smartphone.

[0458] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0459] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0460] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0461] [Third Embodiment]

[0462] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0463] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0464] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0465] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0466] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0467] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0468] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0469] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0470] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0471] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0472] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0473] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0474] This invention provides a system that facilitates smooth communication between distant family members and includes functions for voice analysis, conversation pattern learning, interactive games, and information organization and distribution. The system operates through the interaction of a server, a terminal, and a user.

[0475] Forms of speech analysis

[0476] When a user initiates a video call, the device uses its microphone to acquire audio data in real time. This allows the device to efficiently capture the user's speech and transmit highly accurate audio data to the server. The server analyzes the received audio data using a speech recognition algorithm to identify the speaker and learn the unique voice characteristics of each user, enabling it to accurately understand the conversation.

[0477] Forms of learning conversation patterns

[0478] The server accumulates past conversation history and learns conversation patterns using frequency analysis and topic models based on that data. This allows it to understand the content of what was said and the patterns of the conversation flow, and to utilize the user's unique communication style in future conversations.

[0479] Interactive game delivery methods

[0480] The server selects an appropriate interactive game based on the conversation context and participants' interests. The game is displayed to the user via their device and is designed for intuitive operation. Specifically, it supports the user's participation in a quiz game, guiding them through the game's progression to elicit smart responses.

[0481] Forms of information organization and distribution

[0482] The server aggregates family event information and health data, organizing important information through an information management system. This information is then categorized based on urgency and importance, and notified to terminals using an information distribution system. This notification allows users to quickly take necessary actions.

[0483] In this way, the system of the present invention can efficiently carry out a series of processes from voice recognition to learning and information provision, thereby enriching communication among family members that transcends physical distance.

[0484] The following describes the processing flow.

[0485] Step 1:

[0486] When a user initiates a video call, the device immediately acquires audio data using its built-in microphone. This audio data is then processed to ensure clarity using noise-canceling technology.

[0487] Step 2:

[0488] The terminal compresses the acquired audio data in real time and sends it to the server via a secure protocol. The data is transferred using a highly efficient encoding method, minimizing latency.

[0489] Step 3:

[0490] The server inputs the received audio data into a speech recognition algorithm for analysis. Here, speech features are extracted to identify each speaker, and speaker identification is performed.

[0491] Step 4:

[0492] The server saves the identified speaker's voice characteristics to a cloud database and updates the dataset used for subsequent conversation analysis. This data is also protected by security protocols.

[0493] Step 5:

[0494] The server performs frequency analysis based on past conversation history to extract characteristic conversation patterns. It then uses topic models to analyze the flow of topics and the relationships between important keywords.

[0495] Step 6:

[0496] Based on learned conversation patterns, the server prepares to naturally suggest relevant information and topics in the next conversation, thereby helping to keep the conversation flowing smoothly.

[0497] Step 7:

[0498] The server selects an appropriate interactive game based on the conversation content and the user's interests, and prepares it for play. This selection process also utilizes the user's past game history stored in the database.

[0499] Step 8:

[0500] The device presents the user with an interactive game provided by the server and encourages them to participate. If the user consents, the game starts and leaderboard and score display functions are enabled.

[0501] Step 9:

[0502] The server aggregates and organizes family calendar events and health information. This information is categorized according to priority, and only the information relevant to each user is selected.

[0503] Step 10:

[0504] The server notifies the terminal of organized information at the necessary time, helping users to respond quickly. Each notification is sent as a push notification, so it reaches the user immediately.

[0505] (Example 1)

[0506] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0507] A key challenge is achieving smooth and intimate communication among family members living far apart. In particular, technological solutions are needed to prevent the weakening of communication caused by physical distance and to deepen family ties. Furthermore, improvements in voice recognition accuracy and the provision of information tailored to individual communication styles are essential.

[0508] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0509] In this invention, the server includes voice acquisition means, data transmission means, data analysis means, conversation analysis means, activity provision means, and information organization and distribution means. This enables real-time and effective communication that transcends physical distance and allows for the provision of interactive support to strengthen family bonds.

[0510] "Voice acquisition means" refers to technology for accurately acquiring the user's voice and collecting clear audio data by reducing external noise.

[0511] "Data transmission means" refers to technology for encoding acquired audio data and sending it to a server using a secure communication protocol.

[0512] "Data analysis means" refers to a technology that analyzes audio data received on a server using speech recognition technology, converts the spoken content into text, and identifies individual speakers.

[0513] A "conversation analysis method" is a technique that uses natural language processing technology to learn conversation patterns based on accumulated conversation data and identify frequently occurring topics and expressions.

[0514] "Activity delivery means" refers to technology that selects and provides the most suitable interactive activity to the user according to the user's conversation situation and profile.

[0515] "Information organization and distribution means" refers to technology that organizes family event information and health data and accurately notifies users of their devices according to their importance.

[0516] This invention is a system for facilitating smooth communication between family members living far apart. The specific implementation of this system is described below.

[0517] When a user initiates a video call, the device uses its built-in microphone to acquire audio data in real time. During this process, noise cancellation technology is used to reduce external noise and capture clearer audio.

[0518] The terminal encodes the collected audio data and sends it to the server using a secure communication protocol. Commonly used codecs can be used as the encoding technology. Data security is ensured by using encryption protocols such as TLS during this communication.

[0519] The server analyzes the received audio data using a speech recognition algorithm and converts the spoken content into text data. During this process, speaker identification technology analyzes the characteristics of each voice to identify who is speaking. Generally, commercial speech recognition APIs are available for this purpose.

[0520] Furthermore, the server uses the analyzed text data to analyze conversation patterns. The conversation data is stored in a database, and frequent expressions and topics are extracted using natural language processing techniques. By utilizing topic models such as LDA (Latent Dirichlet Allocation), it is possible to understand the trends in conversation.

[0521] Furthermore, the server selects and provides appropriate interactive activities based on each user's conversation status and interests. These selected activities, such as quiz games and puzzles, are designed to be enjoyable for families to participate in together. These activities are designed to be intuitive and easy to use.

[0522] The server organizes family events and health information in a cloud database and notifies devices based on their importance. A general cloud platform can be used to organize this information.

[0523] Through the process described above, it is possible to enrich communication among family members, transcending physical distance. An example of a prompt would be, "Suggest ways to deepen family connections using a remote communication system." This prompt can be used to obtain further insights and suggestions from the generative AI model.

[0524] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0525] Step 1:

[0526] When a user starts a video call, the device automatically activates the microphone and acquires audio data in real time. The input is the user's voice, and the acquired audio data is output as clear audio data after noise reduction using noise cancellation technology. Specifically, the device samples the audio signal at regular intervals and converts it into digital data.

[0527] Step 2:

[0528] The terminal encodes the acquired audio data and sends it to the server using a secure protocol. The input is the unencoded audio data, and the output is the encoded audio data in binary format. Specifically, the terminal applies data encryption technology and establishes a communication channel for transmitting data over the network.

[0529] Step 3:

[0530] The server analyzes the received audio data using a speech recognition algorithm and converts it into text data. The input is encoded audio data, and the output is the analyzed text data. Specifically, the server analyzes the audio waveform using an acoustic model and a language model and maps it to the corresponding text.

[0531] Step 4:

[0532] The server stores text data in a database and analyzes conversation patterns using natural language processing techniques. The input is text data, and the output is metadata indicating conversation patterns. Specifically, the server uses topic modeling techniques to identify topics and frequently occurring expressions in the conversation.

[0533] Step 5:

[0534] The server selects and provides appropriate interactive activities to the user based on the results of conversation analysis. The input is metadata about the conversation pattern, and the output is the game or activity presented to the user. Specifically, the server uses an AI recommendation algorithm to select the optimal game and displays it via the terminal.

[0535] Step 6:

[0536] The server integrates family events and health information, and delivers organized information to devices. The input is a collection of unorganized data, and the output is notification information delivered based on priority. Specifically, the server classifies the data in a cloud database and delivers the information to devices using a push notification system.

[0537] In this way, this system enables efficient and interactive long-distance family communication.

[0538] (Application Example 1)

[0539] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0540] Communication between family members living in different locations can be difficult due to the constraints of physical distance. Furthermore, real-time situation monitoring to ensure family safety and providing education to raise security awareness are crucial. This invention aims to solve these problems and realize safer and more effective communication.

[0541] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0542] In this invention, the server includes voice analysis means for acquiring voice data in real time and identifying speaker characteristics, situation monitoring means for analyzing voice data to monitor security status and issuing alerts when an anomaly is detected, and educational means for learning safety measures through an interactive game. This enables smooth communication among family members in remote locations and the realization of a highly secure living environment.

[0543] "Voice data" refers to digital information used for communication, which records a user's speech in real time.

[0544] "Speaker characteristics" refer to individually identifiable phonetic attributes such as voice quality and vocalization patterns.

[0545] "Voice analysis means" refers to technology used to analyze voice data and identify the speaker.

[0546] "Conversation patterns" are data that shows the flow and content trends of past conversations.

[0547] "Learning methods" refer to the process of acquiring new knowledge by analyzing past conversation history.

[0548] An "interactive game" is dynamic entertainment content that users can directly interact with and participate in.

[0549] "Game delivery methods" refer to technologies that deliver interactive games to users.

[0550] "Family events" refer to plans and occurrences that are shared among family members.

[0551] "Health information" refers to data related to an individual's health status.

[0552] "Information management means" refers to technologies and methods for organizing data and keeping it in a usable state.

[0553] "Information distribution means" refers to technologies for transmitting organized information to users.

[0554] "Situation monitoring means" refers to technologies for analyzing environmental data and detecting anomalies.

[0555] "Educational tools" are methods and devices used to enable users to learn information and skills.

[0556] The system for implementing the present invention primarily functions as real-time analysis of voice data, monitoring of security status, and provision of education through interactive games.

[0557] The terminal acquires voice data in real time when the user engages in voice communication and transmits this data to the server in a highly accurate format. The server analyzes the data using a speech recognition engine to identify the speaker's characteristics and monitors the surrounding environmental sounds. This provides a mechanism to immediately send an alert to the user's terminal if an anomaly or danger is detected.

[0558] The hardware used includes smart glasses and smartphones, while the software employs speech recognition technologies such as Google Cloud Speech-to-Text. The server stores past conversation history and environmental data in a database and uses a learning algorithm to analyze conversation patterns and sound changes. Based on this analysis, the system improves its ability to detect anomalies from everyday patterns that pose a low security risk.

[0559] Interactive games allow users to experience various security scenarios and are developed using game engines such as Unity. Within the game, users can learn security measures in a virtual environment and apply what they learn to real-world actions.

[0560] For example, if a user is walking in a park and suddenly hears a loud noise nearby, the system will analyze the sound, and if it detects an unusual pattern, it will immediately display an alert on the user's smart glasses. Furthermore, it is possible to provide the generating AI model with a prompt such as, "What technologies should be incorporated into daily life to improve family security?" to obtain additional advice and information regarding the system's security features.

[0561] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0562] Step 1:

[0563] When a user initiates voice communication, the device uses its microphone to capture ambient sounds and speech in real time. This results in voice data being input, which is then transmitted to a server via wireless communication.

[0564] Step 2:

[0565] The server analyzes the received audio data through a speech recognition engine to identify the speaker's voiceprint. This process converts the audio data into text information and compares it with voiceprint information in a database to improve the accuracy of speaker identification. The output obtained here is the audio data converted into text information.

[0566] Step 3:

[0567] The system monitors the ambient sounds in the user's environment in real time and detects any deviations from a defined safety pattern as an anomaly. The server uses this input data to perform algorithmic calculations based on the surrounding sound patterns, generating an anomaly detection output. An anomaly alert is then generated based on this output.

[0568] Step 4:

[0569] The server generates an interactive game and sends it to the user's device. Through this, the user practices skills to deal with a virtual security situation. Using a game engine (Unity), simulation data is generated, and the resulting game content is displayed on the user's device.

[0570] Step 5:

[0571] The user inputs security-related prompts into the generated AI model to obtain necessary additional information and advice. The input prompts are sent to the server as text data, the model analyzes their content, generates output as useful recommendations, and presents them to the user.

[0572] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0573] This invention is a family communication support system incorporating an emotion recognition engine. In addition to voice analysis, conversation pattern learning, interactive game provision, and information management and distribution, the emotion recognition function works in conjunction with these other functions. The specific form of the system is described below.

[0574] Forms of speech analysis and emotion recognition

[0575] When a user initiates a video call, the device acquires audio data in real time. The audio data is de-noised, compressed, and securely transmitted to the server. The server uses voice analysis to identify the speaker and an emotion recognition engine to evaluate the voice tone and speech intensity. The emotion recognition engine identifies the user's emotional state based on the audio data and uses this information to support the conversation.

[0576] Conversation pattern learning and forms of emotional response

[0577] The server analyzes past conversation history and sentiment data, learning family-specific conversation patterns using frequency analysis and topic models. This process enables flexible topic suggestions and comment support tailored to the user's emotions. For example, if a user is feeling stressed, it will offer relaxing topics.

[0578] Interactive game delivery methods

[0579] Based on the user's emotional state, as determined by emotion recognition, the server dynamically selects the content of the interactive game. For depressed users, games that encourage or promote a sense of accomplishment are chosen. On the other hand, if a cheerful emotion is recognized, a competitive game is provided, and the difficulty level is adjusted accordingly.

[0580] Information management and distribution methods

[0581] The server aggregates family event information and health data, classifying the organized information based on its importance and the user's current emotional state. Information is delivered to the device at the appropriate time, allowing users to intuitively review and utilize the information. For example, if a user is in a positive mood, notifications encouraging participation in new events are prioritized.

[0582] In this way, the present invention utilizes an emotion engine to build a system that supports rich dialogue and information sharing among family members, enabling communication that senses emotions even remotely.

[0583] The following describes the processing flow.

[0584] Step 1:

[0585] When a user initiates a video call, the device acquires audio data in real time through the microphone. This audio data is then noise-filtered to ensure clear data is obtained.

[0586] Step 2:

[0587] The terminal compresses the acquired audio data to efficiently utilize network bandwidth and then securely transmits it to the server.

[0588] Step 3:

[0589] The server uses a speech recognition algorithm to identify the speaker in order to analyze the received audio data. This process identifies speaker-specific vocal features and voiceprints.

[0590] Step 4:

[0591] The server uses an emotion recognition engine to evaluate the user's emotional state based on their tone of voice and word choice. This allows it to identify emotions such as happiness, sadness, and anger in real time.

[0592] Step 5:

[0593] The server integrates past conversation history and sentiment data, and uses frequency analysis and topic models to learn conversation patterns between users. Based on the user's sentiment history, it can suggest topics that take into account predicted responses.

[0594] Step 6:

[0595] The server selects an appropriate interactive game based on emotional data. For example, if a user is feeling down, it will choose a game that is likely to elicit positive feedback. The server then notifies the user of the selected game.

[0596] Step 7:

[0597] The device displays game information sent from the server to the user and offers the option to participate in the game. If the user consents, the game starts and the progress is displayed in real time.

[0598] Step 8:

[0599] The server collects important family events and health information, organizing and prioritizing it. In particular, it takes the user's emotional state into consideration when deciding which information to notify them of first.

[0600] Step 9:

[0601] The server sends organized information as a push notification to the user's device, enabling a quick response. Users can then easily check this information on their device and take the necessary actions.

[0602] (Example 2)

[0603] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0604] In today's digital society, communication between family members living far apart is becoming increasingly difficult. In particular, accurately conveying and sharing emotions is challenging, necessitating richer and more effective communication methods. Furthermore, the inability to offer appropriate information tailored to the context of the conversation and individual emotional states can lead to a decline in the quality of communication.

[0605] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0606] In this invention, the server includes voice analysis means for instantly acquiring voice information and identifying the characteristics of the speaker, analysis means for analyzing past dialogue history and learning dialogue patterns, and emotion evaluation means for evaluating the emotional state of the speaker using an emotion recognition engine. This enables emotionally responsive communication between family members in remote locations and realizes flexible and appropriate dialogue suggestions tailored to the user's emotions.

[0607] "Voice information" refers to speech data collected from users, and voice analysis and emotion recognition are performed based on this data.

[0608] "Instant acquisition" means that voice information is collected in real time, with the aim of responding to user speech without delay.

[0609] "Speaker characteristics" refers to the individual voice and speaking style characteristics of each user, identified based on audio data.

[0610] "Speech analysis means" refers to technical methods for analyzing speech information to identify the characteristics of the speaker and the speaker themselves.

[0611] "Dialogue history" refers to a record of past user-to-user conversation data, and dialogue patterns are learned based on this.

[0612] "Analysis methods" refer to technical techniques used to analyze specific patterns and characteristics using collected data.

[0613] An "emotion recognition engine" refers to a program or system that has the function of identifying and evaluating a user's emotional state based on voice information.

[0614] "Emotional evaluation method" refers to a technical method for evaluating a user's emotional state based on collected audio information.

[0615] An "interactive game" refers to an entertainment activity that allows for active interaction with the user, and whose content is adjusted according to the user's emotional state.

[0616] "Entertainment provision means" refers to methods and technologies for providing interactive games to users.

[0617] "Information organization methods" refer to technical techniques for organizing collected information about household events and health-related matters.

[0618] "Information transmission means" refers to methods and technologies for notifying users of organized information to their terminals in a timely manner.

[0619] "Suggestion generation means" refers to programs and technologies for dynamically generating personalized dialogue and entertainment suggestions for users.

[0620] This invention is implemented as a family communication support system incorporating an emotion recognition engine. The system operates in conjunction with functions for voice analysis, conversation pattern learning, interactive entertainment provision, and information management and distribution. Specific embodiments are described below.

[0621] Voice analysis and emotion recognition

[0622] When a user initiates a voice call, the device uses its microphone to acquire voice information in real time. This voice data is then denoised, compressed, encrypted, and transmitted to a server via the internet. The server uses voice analysis tools to identify the speaker's characteristics and an emotion recognition engine to analyze the tone and intensity of speech to determine the user's emotional state. This process enables emotionally expressive communication between users.

[0623] Conversation pattern learning and suggestion generation

[0624] The server analyzes past conversation history and current emotional state. Based on the dialogue history stored in the database, it performs frequency analysis and topic modeling to learn unique dialogue patterns between users. Using this information, the server leverages a generative AI model to dynamically generate flexible and accurate dialogue and entertainment suggestions tailored to the user.

[0625] Interactive entertainment provider

[0626] Based on the emotional state identified by the emotion recognition engine, the server selects the most appropriate interactive entertainment. If the user is in a cheerful state, a competitive game is provided to the device, offering a higher level of challenge. Conversely, if factors indicating a need for relaxation are detected, a game with encouraging content is selected.

[0627] Information management and distribution

[0628] The server aggregates and organizes family event information and health data. The information is categorized according to the user's emotional state and notified to the user's device at the appropriate time through various communication channels. For example, users with positive emotions receive priority notifications of invitations to new events.

[0629] One concrete example is video chat between family members. For instance, in a scene where a grandchild and grandfather are talking, voice analysis can recognize the grandchild's enjoyment from their voice and suggest competitive games or other activities.

[0630] Examples of prompts to input into the generative AI model include, "Think about how to help a user relax when they are feeling stressed." This allows for natural and meaningful interactions that respond to the user's emotions.

[0631] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0632] Step 1:

[0633] The device instantly acquires audio data using the microphone as soon as the user initiates a video call. The acquired audio data is then filtered to remove background noise. Subsequently, data compression technology is used to reduce the size of the audio data, making it suitable for efficient transmission. The input to this process is the user's raw voice, and the output is compressed and encrypted audio data. The encrypted audio data is sent to the server through a secure communication channel.

[0634] Step 2:

[0635] The server receives audio data transmitted from the terminal and identifies the speaker's characteristics using speech analysis means. In this step, a speech recognition algorithm is used to characterize the speaker's voiceprint from the audio data. The input is encrypted audio data, and the output is the analyzed speaker characteristics. Simultaneously, an emotion recognition engine is interfaced to identify the user's emotional state by analyzing the voice tone and speech intensity.

[0636] Step 3:

[0637] The server performs analysis using previously collected dialogue history and current sentiment data. This process includes frequency analysis and topic modeling to learn unique dialogue patterns between users. The input is past dialogue history and sentiment data, and the output is the learned conversation patterns. Based on this data, the server utilizes a generative AI model to construct dialogues and entertainment suggestions that match the user's current situation.

[0638] Step 4:

[0639] The server selects the most appropriate interactive entertainment or conversation based on information obtained from the emotion recognition engine. It prompts a generative AI model, which automatically generates content tailored to the user's emotions. For example, if the user is feeling down, it will select an uplifting game or a relaxing topic. The input for this step is the user's emotion data, and the output is the content of the entertainment or conversation provided to the user.

[0640] Step 5:

[0641] The server collects organized family event information and health-related data, and uses information management and distribution tools to notify the user of the identified information. The timing of this notification is determined by the user's emotional state and the importance of the information. For example, when the user is in a positive emotional state, an invitation to participate in a new event will be sent to the device. The inputs to this process are family data and emotional state, and the output is the notification information sent to the device.

[0642] (Application Example 2)

[0643] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0644] In modern society, while the demand for online shopping and virtual experiences is increasing, there is a lack of customized product recommendations tailored to each user's individual emotional state and preferences. As a result, users may miss out on products and services they truly need, leading to decreased satisfaction with their purchasing experience. Therefore, there is a need to develop a system that can recognize a user's emotional state in real time and provide appropriate product recommendations based on that understanding.

[0645] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0646] In this invention, the server includes voice analysis means for acquiring voice data in real time and identifying the user's characteristics, means for making product suggestions based on the user's emotional state using an emotion recognition engine, and information distribution means for notifying the user terminal of the organized information. This makes it possible to make product suggestions based on the user's emotional state.

[0647] "Voice data" refers to information obtained by digitally capturing the user's voice in real time and processing it using a speech recognition algorithm.

[0648] "User characteristics" refers to information that describes the voice characteristics and patterns used to identify individual users.

[0649] "Voice analysis means" refers to a technical method for acquiring voice data in real time and identifying the characteristics of the user.

[0650] "Past conversation history" refers to information that records previous conversations between the user and others.

[0651] "Learning methods" refer to technical techniques for analyzing past conversation history and learning conversation patterns.

[0652] "Interactive entertainment" refers to entertainment content that actively engages the user and dynamically changes based on emotional recognition.

[0653] "Means of providing entertainment" refers to technical methods that provide interactive entertainment and allow users to experience it interactively.

[0654] "Domestic events" is a concept that refers to events and activities that occur among family members and close friends.

[0655] "Health information" refers to information about the health status of individual users or people within their household.

[0656] "Information management methods" refer to technical techniques for collecting and organizing information about events and health within the household.

[0657] "User terminal" refers to an electronic device used by a user to receive and operate information.

[0658] "Information distribution means" refers to technical methods for appropriately notifying users of organized information at their terminals.

[0659] An "emotion recognition engine" is an algorithm that identifies the user's emotional state based on factors such as voice tone and speech intensity.

[0660] "Product recommendation" is the act of recommending appropriate products or services based on the user's emotional state.

[0661] This invention is realized by acquiring voice data in real time using a terminal held by the user and processing that data using a voice analysis means. The terminal first has the function of converting the user's voice into digital data and securely transmitting it to a server. The server receives this voice data, removes noise, and then performs voice analysis. As a result, the user's characteristics are identified, and the user's emotional state is determined by an emotion recognition engine.

[0662] The server also analyzes past conversation history and learns user-specific conversation patterns using frequency analysis and topic models. Based on this, it can provide product recommendations that are best suited to the user's emotions. Product data is managed in a cloud-based database, and recommendations are made in real time, taking into account the user's emotional state.

[0663] For example, if a positive emotion is identified by emotion recognition while a user is shopping, products and services with a high affinity for that emotion can be recommended. The user can then use their smartphone to review and purchase the suggested products. An example of a prompt message would be, "Based on the user's current emotional state, please suggest the most suitable products for them," which could be given to the generative AI model.

[0664] In this way, a system that integrates emotion recognition and voice analysis technologies can enhance the user's purchasing experience. The system utilizes a communication network, a cloud-based database, a voice analysis algorithm, and an emotion recognition engine. This enables the provision of a personalized purchasing experience to users, thereby increasing satisfaction.

[0665] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0666] Step 1:

[0667] The device acquires the user's voice via a microphone, converts that voice into digital data, and sends it to the server. The input is the user's spoken voice, and the output is digital audio data. The device uses internet communication for data acquisition and transmission.

[0668] Step 2:

[0669] The server removes noise from the received audio data and analyzes the user's characteristics using a speech recognition algorithm. The input is audio data transmitted from the terminal, and the output is the identified user's voice profile. The speech recognition algorithm utilizes a cloud-based service to identify each user's voice characteristics.

[0670] Step 3:

[0671] The server operates an emotion recognition engine to identify the user's emotional state by evaluating voice tone and speech intensity. The input is a voice profile based on voice data, and the output is the user's emotional state. In this process, the emotion recognition engine analyzes the tone and intensity to determine characteristic emotions.

[0672] Step 4:

[0673] The server learns conversation patterns using recorded past conversation history and sentiment data, and generates frequency analyses and topic models. Inputs include past conversation history and current sentiment states, and output is the learned conversation patterns. Machine learning techniques are used to analyze the data.

[0674] Step 5:

[0675] The server retrieves product data from a cloud-based database and suggests the most suitable product based on the user's emotional state. The input is the user's emotional state and conversation patterns, and the output is a list of suggested products. The prompt is passed to a generating AI model, and the product recommendation algorithm selects the most suitable product for the user.

[0676] Step 6:

[0677] The server notifies the terminal of the suggested product information, allowing the user to view the suggestions. The input is the list of suggested products, and the output is the product information displayed on the user's terminal. The notification function uses push notifications, allowing the user to instantly check the information on their smartphone.

[0678] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0679] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0680] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0681] [Fourth Embodiment]

[0682] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0683] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0684] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0685] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0686] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0687] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0688] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0689] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0690] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0691] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0692] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0693] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0694] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0695] This invention provides a system that facilitates smooth communication between distant family members and includes functions for voice analysis, conversation pattern learning, interactive games, and information organization and distribution. The system operates through the interaction of a server, a terminal, and a user.

[0696] Forms of speech analysis

[0697] When a user initiates a video call, the device uses its microphone to acquire audio data in real time. This allows the device to efficiently capture the user's speech and transmit highly accurate audio data to the server. The server analyzes the received audio data using a speech recognition algorithm to identify the speaker and learn the unique voice characteristics of each user, enabling it to accurately understand the conversation.

[0698] Forms of learning conversation patterns

[0699] The server accumulates past conversation history and learns conversation patterns using frequency analysis and topic models based on that data. This allows it to understand the content of what was said and the patterns of the conversation flow, and to utilize the user's unique communication style in future conversations.

[0700] Interactive game delivery methods

[0701] The server selects an appropriate interactive game based on the conversation context and participants' interests. The game is displayed to the user via their device and is designed for intuitive operation. Specifically, it supports the user's participation in a quiz game, guiding them through the game's progression to elicit smart responses.

[0702] Forms of information organization and distribution

[0703] The server aggregates family event information and health data, organizing important information through an information management system. This information is then categorized based on urgency and importance, and notified to terminals using an information distribution system. This notification allows users to quickly take necessary actions.

[0704] In this way, the system of the present invention can efficiently carry out a series of processes from voice recognition to learning and information provision, thereby enriching communication among family members that transcends physical distance.

[0705] The following describes the processing flow.

[0706] Step 1:

[0707] When a user initiates a video call, the device immediately acquires audio data using its built-in microphone. This audio data is then processed to ensure clarity using noise-canceling technology.

[0708] Step 2:

[0709] The terminal compresses the acquired audio data in real time and sends it to the server via a secure protocol. The data is transferred using a highly efficient encoding method, minimizing latency.

[0710] Step 3:

[0711] The server inputs the received audio data into a speech recognition algorithm for analysis. Here, speech features are extracted to identify each speaker, and speaker identification is performed.

[0712] Step 4:

[0713] The server saves the identified speaker's voice characteristics to a cloud database and updates the dataset used for subsequent conversation analysis. This data is also protected by security protocols.

[0714] Step 5:

[0715] The server performs frequency analysis based on past conversation history to extract characteristic conversation patterns. It then uses topic models to analyze the flow of topics and the relationships between important keywords.

[0716] Step 6:

[0717] Based on learned conversation patterns, the server prepares to naturally suggest relevant information and topics in the next conversation, thereby helping to keep the conversation flowing smoothly.

[0718] Step 7:

[0719] The server selects an appropriate interactive game based on the conversation content and the user's interests, and prepares it for play. This selection process also utilizes the user's past game history stored in the database.

[0720] Step 8:

[0721] The device presents the user with an interactive game provided by the server and encourages them to participate. If the user consents, the game starts and leaderboard and score display functions are enabled.

[0722] Step 9:

[0723] The server aggregates and organizes family calendar events and health information. This information is categorized according to priority, and only the information relevant to each user is selected.

[0724] Step 10:

[0725] The server notifies the terminal of organized information at the necessary time, helping users to respond quickly. Each notification is sent as a push notification, so it reaches the user immediately.

[0726] (Example 1)

[0727] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0728] A key challenge is achieving smooth and intimate communication among family members living far apart. In particular, technological solutions are needed to prevent the weakening of communication caused by physical distance and to deepen family ties. Furthermore, improvements in voice recognition accuracy and the provision of information tailored to individual communication styles are essential.

[0729] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0730] In this invention, the server includes voice acquisition means, data transmission means, data analysis means, conversation analysis means, activity provision means, and information organization and distribution means. This enables real-time and effective communication that transcends physical distance and allows for the provision of interactive support to strengthen family bonds.

[0731] "Voice acquisition means" refers to technology for accurately acquiring the user's voice and collecting clear audio data by reducing external noise.

[0732] "Data transmission means" refers to technology for encoding acquired audio data and sending it to a server using a secure communication protocol.

[0733] "Data analysis means" refers to a technology that analyzes audio data received on a server using speech recognition technology, converts the spoken content into text, and identifies individual speakers.

[0734] A "conversation analysis method" is a technique that uses natural language processing technology to learn conversation patterns based on accumulated conversation data and identify frequently occurring topics and expressions.

[0735] "Activity delivery means" refers to technology that selects and provides the most suitable interactive activity to the user according to the user's conversation situation and profile.

[0736] "Information organization and distribution means" refers to technology that organizes family event information and health data and accurately notifies users of their devices according to their importance.

[0737] This invention is a system for facilitating smooth communication between family members living far apart. The specific implementation of this system is described below.

[0738] When a user initiates a video call, the device uses its built-in microphone to acquire audio data in real time. During this process, noise cancellation technology is used to reduce external noise and capture clearer audio.

[0739] The terminal encodes the collected audio data and sends it to the server using a secure communication protocol. Commonly used codecs can be used as the encoding technology. Data security is ensured by using encryption protocols such as TLS during this communication.

[0740] The server analyzes the received audio data using a speech recognition algorithm and converts the spoken content into text data. During this process, speaker identification technology analyzes the characteristics of each voice to identify who is speaking. Generally, commercial speech recognition APIs are available for this purpose.

[0741] Furthermore, the server uses the analyzed text data to analyze conversation patterns. The conversation data is stored in a database, and frequent expressions and topics are extracted using natural language processing techniques. By utilizing topic models such as LDA (Latent Dirichlet Allocation), it is possible to understand the trends in conversation.

[0742] Furthermore, the server selects and provides appropriate interactive activities based on each user's conversation status and interests. These selected activities, such as quiz games and puzzles, are designed to be enjoyable for families to participate in together. These activities are designed to be intuitive and easy to use.

[0743] The server organizes family events and health information in a cloud database and notifies devices based on their importance. A general cloud platform can be used to organize this information.

[0744] Through the process described above, it is possible to enrich communication among family members, transcending physical distance. An example of a prompt would be, "Suggest ways to deepen family connections using a remote communication system." This prompt can be used to obtain further insights and suggestions from the generative AI model.

[0745] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0746] Step 1:

[0747] When a user starts a video call, the device automatically activates the microphone and acquires audio data in real time. The input is the user's voice, and the acquired audio data is output as clear audio data after noise reduction using noise cancellation technology. Specifically, the device samples the audio signal at regular intervals and converts it into digital data.

[0748] Step 2:

[0749] The terminal encodes the acquired audio data and sends it to the server using a secure protocol. The input is the unencoded audio data, and the output is the encoded audio data in binary format. Specifically, the terminal applies data encryption technology and establishes a communication channel for transmitting data over the network.

[0750] Step 3:

[0751] The server analyzes the received audio data using a speech recognition algorithm and converts it into text data. The input is encoded audio data, and the output is the analyzed text data. Specifically, the server analyzes the audio waveform using an acoustic model and a language model and maps it to the corresponding text.

[0752] Step 4:

[0753] The server stores text data in a database and analyzes conversation patterns using natural language processing techniques. The input is text data, and the output is metadata indicating conversation patterns. Specifically, the server uses topic modeling techniques to identify topics and frequently occurring expressions in the conversation.

[0754] Step 5:

[0755] The server selects and provides appropriate interactive activities to the user based on the results of conversation analysis. The input is metadata about the conversation pattern, and the output is the game or activity presented to the user. Specifically, the server uses an AI recommendation algorithm to select the optimal game and displays it via the terminal.

[0756] Step 6:

[0757] The server integrates family events and health information, and delivers organized information to devices. The input is a collection of unorganized data, and the output is notification information delivered based on priority. Specifically, the server classifies the data in a cloud database and delivers the information to devices using a push notification system.

[0758] In this way, this system enables efficient and interactive long-distance family communication.

[0759] (Application Example 1)

[0760] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0761] Communication between family members living in different locations can be difficult due to the constraints of physical distance. Furthermore, real-time situation monitoring to ensure family safety and providing education to raise security awareness are crucial. This invention aims to solve these problems and realize safer and more effective communication.

[0762] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0763] In this invention, the server includes voice analysis means for acquiring voice data in real time and identifying speaker characteristics, situation monitoring means for analyzing voice data to monitor security status and issuing alerts when an anomaly is detected, and educational means for learning safety measures through an interactive game. This enables smooth communication among family members in remote locations and the realization of a highly secure living environment.

[0764] "Voice data" refers to digital information used for communication, which records a user's speech in real time.

[0765] "Speaker characteristics" refer to individually identifiable phonetic attributes such as voice quality and vocalization patterns.

[0766] "Voice analysis means" refers to technology used to analyze voice data and identify the speaker.

[0767] "Conversation patterns" are data that shows the flow and content trends of past conversations.

[0768] "Learning methods" refer to the process of acquiring new knowledge by analyzing past conversation history.

[0769] An "interactive game" is dynamic entertainment content that users can directly interact with and participate in.

[0770] "Game delivery methods" refer to technologies that deliver interactive games to users.

[0771] "Family events" refer to plans and occurrences that are shared among family members.

[0772] "Health information" refers to data related to an individual's health status.

[0773] "Information management means" refers to technologies and methods for organizing data and keeping it in a usable state.

[0774] "Information distribution means" refers to technologies for transmitting organized information to users.

[0775] "Situation monitoring means" refers to technologies for analyzing environmental data and detecting anomalies.

[0776] "Educational tools" are methods and devices used to enable users to learn information and skills.

[0777] The system for implementing the present invention primarily functions as real-time analysis of voice data, monitoring of security status, and provision of education through interactive games.

[0778] The terminal acquires voice data in real time when the user engages in voice communication and transmits this data to the server in a highly accurate format. The server analyzes the data using a speech recognition engine to identify the speaker's characteristics and monitors the surrounding environmental sounds. This provides a mechanism to immediately send an alert to the user's terminal if an anomaly or danger is detected.

[0779] The hardware used includes smart glasses and smartphones, while the software employs speech recognition technologies such as Google Cloud Speech-to-Text. The server stores past conversation history and environmental data in a database and uses a learning algorithm to analyze conversation patterns and sound changes. Based on this analysis, the system improves its ability to detect anomalies from everyday patterns that pose a low security risk.

[0780] Interactive games allow users to experience various security scenarios and are developed using game engines such as Unity. Within the game, users can learn security measures in a virtual environment and apply what they learn to real-world actions.

[0781] For example, if a user is walking in a park and suddenly hears a loud noise nearby, the system will analyze the sound, and if it detects an unusual pattern, it will immediately display an alert on the user's smart glasses. Furthermore, it is possible to provide the generating AI model with a prompt such as, "What technologies should be incorporated into daily life to improve family security?" to obtain additional advice and information regarding the system's security features.

[0782] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0783] Step 1:

[0784] When a user initiates voice communication, the device uses its microphone to capture ambient sounds and speech in real time. This results in voice data being input, which is then transmitted to a server via wireless communication.

[0785] Step 2:

[0786] The server analyzes the received audio data through a speech recognition engine to identify the speaker's voiceprint. This process converts the audio data into text information and compares it with voiceprint information in a database to improve the accuracy of speaker identification. The output obtained here is the audio data converted into text information.

[0787] Step 3:

[0788] The system monitors the ambient sounds in the user's environment in real time and detects any deviations from a defined safety pattern as an anomaly. The server uses this input data to perform algorithmic calculations based on the surrounding sound patterns, generating an anomaly detection output. An anomaly alert is then generated based on this output.

[0789] Step 4:

[0790] The server generates an interactive game and sends it to the user's device. Through this, the user practices skills to deal with a virtual security situation. Using a game engine (Unity), simulation data is generated, and the resulting game content is displayed on the user's device.

[0791] Step 5:

[0792] The user inputs security-related prompts into the generated AI model to obtain necessary additional information and advice. The input prompts are sent to the server as text data, the model analyzes their content, generates output as useful recommendations, and presents them to the user.

[0793] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0794] This invention is a family communication support system incorporating an emotion recognition engine. In addition to voice analysis, conversation pattern learning, interactive game provision, and information management and distribution, the emotion recognition function works in conjunction with these other functions. The specific form of the system is described below.

[0795] Forms of speech analysis and emotion recognition

[0796] When a user initiates a video call, the device acquires audio data in real time. The audio data is de-noised, compressed, and securely transmitted to the server. The server uses voice analysis to identify the speaker and an emotion recognition engine to evaluate the voice tone and speech intensity. The emotion recognition engine identifies the user's emotional state based on the audio data and uses this information to support the conversation.

[0797] Conversation pattern learning and forms of emotional response

[0798] The server analyzes past conversation history and sentiment data, learning family-specific conversation patterns using frequency analysis and topic models. This process enables flexible topic suggestions and comment support tailored to the user's emotions. For example, if a user is feeling stressed, it will offer relaxing topics.

[0799] Interactive game delivery methods

[0800] Based on the user's emotional state, as determined by emotion recognition, the server dynamically selects the content of the interactive game. For depressed users, games that encourage or promote a sense of accomplishment are chosen. On the other hand, if a cheerful emotion is recognized, a competitive game is provided, and the difficulty level is adjusted accordingly.

[0801] Information management and distribution methods

[0802] The server aggregates family event information and health data, classifying the organized information based on its importance and the user's current emotional state. Information is delivered to the device at the appropriate time, allowing users to intuitively review and utilize the information. For example, if a user is in a positive mood, notifications encouraging participation in new events are prioritized.

[0803] In this way, the present invention utilizes an emotion engine to build a system that supports rich dialogue and information sharing among family members, enabling communication that senses emotions even remotely.

[0804] The following describes the processing flow.

[0805] Step 1:

[0806] When a user initiates a video call, the device acquires audio data in real time through the microphone. This audio data is then noise-filtered to ensure clear data is obtained.

[0807] Step 2:

[0808] The terminal compresses the acquired audio data to efficiently utilize network bandwidth and then securely transmits it to the server.

[0809] Step 3:

[0810] The server uses a speech recognition algorithm to identify the speaker in order to analyze the received audio data. This process identifies speaker-specific vocal features and voiceprints.

[0811] Step 4:

[0812] The server uses an emotion recognition engine to evaluate the user's emotional state based on their tone of voice and word choice. This allows it to identify emotions such as happiness, sadness, and anger in real time.

[0813] Step 5:

[0814] The server integrates past conversation history and sentiment data, and uses frequency analysis and topic models to learn conversation patterns between users. Based on the user's sentiment history, it can suggest topics that take into account predicted responses.

[0815] Step 6:

[0816] The server selects an appropriate interactive game based on emotional data. For example, if a user is feeling down, it will choose a game that is likely to elicit positive feedback. The server then notifies the user of the selected game.

[0817] Step 7:

[0818] The device displays game information sent from the server to the user and offers the option to participate in the game. If the user consents, the game starts and the progress is displayed in real time.

[0819] Step 8:

[0820] The server collects important family events and health information, organizing and prioritizing it. In particular, it takes the user's emotional state into consideration when deciding which information to notify them of first.

[0821] Step 9:

[0822] The server sends organized information as a push notification to the user's device, enabling a quick response. Users can then easily check this information on their device and take the necessary actions.

[0823] (Example 2)

[0824] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0825] In today's digital society, communication between family members living far apart is becoming increasingly difficult. In particular, accurately conveying and sharing emotions is challenging, necessitating richer and more effective communication methods. Furthermore, the inability to offer appropriate information tailored to the context of the conversation and individual emotional states can lead to a decline in the quality of communication.

[0826] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0827] In this invention, the server includes voice analysis means for instantly acquiring voice information and identifying the characteristics of the speaker, analysis means for analyzing past dialogue history and learning dialogue patterns, and emotion evaluation means for evaluating the emotional state of the speaker using an emotion recognition engine. This enables emotionally responsive communication between family members in remote locations and realizes flexible and appropriate dialogue suggestions tailored to the user's emotions.

[0828] "Voice information" refers to speech data collected from users, and voice analysis and emotion recognition are performed based on this data.

[0829] "Instant acquisition" means that voice information is collected in real time, with the aim of responding to user speech without delay.

[0830] "Speaker characteristics" refers to the individual voice and speaking style characteristics of each user, identified based on audio data.

[0831] "Speech analysis means" refers to technical methods for analyzing speech information to identify the characteristics of the speaker and the speaker themselves.

[0832] "Dialogue history" refers to a record of past user-to-user conversation data, and dialogue patterns are learned based on this.

[0833] "Analysis methods" refer to technical techniques used to analyze specific patterns and characteristics using collected data.

[0834] An "emotion recognition engine" refers to a program or system that has the function of identifying and evaluating a user's emotional state based on voice information.

[0835] "Emotional evaluation method" refers to a technical method for evaluating a user's emotional state based on collected audio information.

[0836] An "interactive game" refers to an entertainment activity that allows for active interaction with the user, and whose content is adjusted according to the user's emotional state.

[0837] "Entertainment provision means" refers to methods and technologies for providing interactive games to users.

[0838] "Information organization methods" refer to technical techniques for organizing collected information about household events and health-related matters.

[0839] "Information transmission means" refers to methods and technologies for notifying users of organized information to their terminals in a timely manner.

[0840] "Suggestion generation means" refers to programs and technologies for dynamically generating personalized dialogue and entertainment suggestions for users.

[0841] This invention is implemented as a family communication support system incorporating an emotion recognition engine. The system operates in conjunction with functions for voice analysis, conversation pattern learning, interactive entertainment provision, and information management and distribution. Specific embodiments are described below.

[0842] Voice analysis and emotion recognition

[0843] When a user initiates a voice call, the device uses its microphone to acquire voice information in real time. This voice data is then denoised, compressed, encrypted, and transmitted to a server via the internet. The server uses voice analysis tools to identify the speaker's characteristics and an emotion recognition engine to analyze the tone and intensity of speech to determine the user's emotional state. This process enables emotionally expressive communication between users.

[0844] Conversation pattern learning and suggestion generation

[0845] The server analyzes past conversation history and current emotional state. Based on the dialogue history stored in the database, it performs frequency analysis and topic modeling to learn unique dialogue patterns between users. Using this information, the server leverages a generative AI model to dynamically generate flexible and accurate dialogue and entertainment suggestions tailored to the user.

[0846] Interactive entertainment provider

[0847] Based on the emotional state identified by the emotion recognition engine, the server selects the most appropriate interactive entertainment. If the user is in a cheerful state, a competitive game is provided to the device, offering a higher level of challenge. Conversely, if factors indicating a need for relaxation are detected, a game with encouraging content is selected.

[0848] Information management and distribution

[0849] The server aggregates and organizes family event information and health data. The information is categorized according to the user's emotional state and notified to the user's device at the appropriate time through various communication channels. For example, users with positive emotions receive priority notifications of invitations to new events.

[0850] One concrete example is video chat between family members. For instance, in a scene where a grandchild and grandfather are talking, voice analysis can recognize the grandchild's enjoyment from their voice and suggest competitive games or other activities.

[0851] Examples of prompts to input into the generative AI model include, "Think about how to help a user relax when they are feeling stressed." This allows for natural and meaningful interactions that respond to the user's emotions.

[0852] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0853] Step 1:

[0854] The device instantly acquires audio data using the microphone as soon as the user initiates a video call. The acquired audio data is then filtered to remove background noise. Subsequently, data compression technology is used to reduce the size of the audio data, making it suitable for efficient transmission. The input to this process is the user's raw voice, and the output is compressed and encrypted audio data. The encrypted audio data is sent to the server through a secure communication channel.

[0855] Step 2:

[0856] The server receives audio data transmitted from the terminal and identifies the speaker's characteristics using speech analysis means. In this step, a speech recognition algorithm is used to characterize the speaker's voiceprint from the audio data. The input is encrypted audio data, and the output is the analyzed speaker characteristics. Simultaneously, an emotion recognition engine is interfaced to identify the user's emotional state by analyzing the voice tone and speech intensity.

[0857] Step 3:

[0858] The server performs analysis using previously collected dialogue history and current sentiment data. This process includes frequency analysis and topic modeling to learn unique dialogue patterns between users. The input is past dialogue history and sentiment data, and the output is the learned conversation patterns. Based on this data, the server utilizes a generative AI model to construct dialogues and entertainment suggestions that match the user's current situation.

[0859] Step 4:

[0860] The server selects the most appropriate interactive entertainment or conversation based on information obtained from the emotion recognition engine. It prompts a generative AI model, which automatically generates content tailored to the user's emotions. For example, if the user is feeling down, it will select an uplifting game or a relaxing topic. The input for this step is the user's emotion data, and the output is the content of the entertainment or conversation provided to the user.

[0861] Step 5:

[0862] The server collects organized family event information and health-related data, and uses information management and distribution tools to notify the user of the identified information. The timing of this notification is determined by the user's emotional state and the importance of the information. For example, when the user is in a positive emotional state, an invitation to participate in a new event will be sent to the device. The inputs to this process are family data and emotional state, and the output is the notification information sent to the device.

[0863] (Application Example 2)

[0864] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0865] In modern society, while the demand for online shopping and virtual experiences is increasing, there is a lack of customized product recommendations tailored to each user's individual emotional state and preferences. As a result, users may miss out on products and services they truly need, leading to decreased satisfaction with their purchasing experience. Therefore, there is a need to develop a system that can recognize a user's emotional state in real time and provide appropriate product recommendations based on that understanding.

[0866] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0867] In this invention, the server includes voice analysis means for acquiring voice data in real time and identifying the user's characteristics, means for making product suggestions based on the user's emotional state using an emotion recognition engine, and information distribution means for notifying the user terminal of the organized information. This makes it possible to make product suggestions based on the user's emotional state.

[0868] "Voice data" refers to information obtained by digitally capturing the user's voice in real time and processing it using a speech recognition algorithm.

[0869] "User characteristics" refers to information that describes the voice characteristics and patterns used to identify individual users.

[0870] "Voice analysis means" refers to a technical method for acquiring voice data in real time and identifying the characteristics of the user.

[0871] "Past conversation history" refers to information that records previous conversations between the user and others.

[0872] "Learning methods" refer to technical techniques for analyzing past conversation history and learning conversation patterns.

[0873] "Interactive entertainment" refers to entertainment content that actively engages the user and dynamically changes based on emotional recognition.

[0874] "Means of providing entertainment" refers to technical methods that provide interactive entertainment and allow users to experience it interactively.

[0875] "Domestic events" is a concept that refers to events and activities that occur among family members and close friends.

[0876] "Health information" refers to information about the health status of individual users or people within their household.

[0877] "Information management methods" refer to technical techniques for collecting and organizing information about events and health within the household.

[0878] "User terminal" refers to an electronic device used by a user to receive and operate information.

[0879] "Information distribution means" refers to technical methods for appropriately notifying users of organized information at their terminals.

[0880] An "emotion recognition engine" is an algorithm that identifies the user's emotional state based on factors such as voice tone and speech intensity.

[0881] "Product recommendation" is the act of recommending appropriate products or services based on the user's emotional state.

[0882] This invention is realized by acquiring voice data in real time using a terminal held by the user and processing that data using a voice analysis means. The terminal first has the function of converting the user's voice into digital data and securely transmitting it to a server. The server receives this voice data, removes noise, and then performs voice analysis. As a result, the user's characteristics are identified, and the user's emotional state is determined by an emotion recognition engine.

[0883] The server also analyzes past conversation history and learns user-specific conversation patterns using frequency analysis and topic models. Based on this, it can provide product recommendations that are best suited to the user's emotions. Product data is managed in a cloud-based database, and recommendations are made in real time, taking into account the user's emotional state.

[0884] For example, if a positive emotion is identified by emotion recognition while a user is shopping, products and services with a high affinity for that emotion can be recommended. The user can then use their smartphone to review and purchase the suggested products. An example of a prompt message would be, "Based on the user's current emotional state, please suggest the most suitable products for them," which could be given to the generative AI model.

[0885] In this way, a system that integrates emotion recognition and voice analysis technologies can enhance the user's purchasing experience. The system utilizes a communication network, a cloud-based database, a voice analysis algorithm, and an emotion recognition engine. This enables the provision of a personalized purchasing experience to users, thereby increasing satisfaction.

[0886] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0887] Step 1:

[0888] The device acquires the user's voice via a microphone, converts that voice into digital data, and sends it to the server. The input is the user's spoken voice, and the output is digital audio data. The device uses internet communication for data acquisition and transmission.

[0889] Step 2:

[0890] The server removes noise from the received audio data and analyzes the user's characteristics using a speech recognition algorithm. The input is audio data transmitted from the terminal, and the output is the identified user's voice profile. The speech recognition algorithm utilizes a cloud-based service to identify each user's voice characteristics.

[0891] Step 3:

[0892] The server operates an emotion recognition engine to identify the user's emotional state by evaluating voice tone and speech intensity. The input is a voice profile based on voice data, and the output is the user's emotional state. In this process, the emotion recognition engine analyzes the tone and intensity to determine characteristic emotions.

[0893] Step 4:

[0894] The server learns conversation patterns using recorded past conversation history and sentiment data, and generates frequency analyses and topic models. Inputs include past conversation history and current sentiment states, and output is the learned conversation patterns. Machine learning techniques are used to analyze the data.

[0895] Step 5:

[0896] The server retrieves product data from a cloud-based database and suggests the most suitable product based on the user's emotional state. The input is the user's emotional state and conversation patterns, and the output is a list of suggested products. The prompt is passed to a generating AI model, and the product recommendation algorithm selects the most suitable product for the user.

[0897] Step 6:

[0898] The server notifies the terminal of the suggested product information, allowing the user to view the suggestions. The input is the list of suggested products, and the output is the product information displayed on the user's terminal. The notification function uses push notifications, allowing the user to instantly check the information on their smartphone.

[0899] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0900] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0901] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0902] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0903] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0904] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0905] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0906] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0907] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0908] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0909] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0910] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0911] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0912] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0913] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0914] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0915] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0916] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0917] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0918] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0919] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0920] The following is further disclosed regarding the embodiments described above.

[0921] (Claim 1)

[0922] A voice analysis method for acquiring voice data in real time and identifying the characteristics of the speaker,

[0923] A learning method for analyzing past conversation history and learning conversation patterns,

[0924] A means of providing and displaying interactive games to participants,

[0925] A means of managing information to collect and organize family events and health information,

[0926] A means of distributing information to notify user terminals of organized information.

[0927] Includes system.

[0928] (Claim 2)

[0929] The system according to claim 1, characterized in that the voice analysis means identifies the speaker's voiceprint using a voice recognition algorithm.

[0930] (Claim 3)

[0931] The system according to claim 1, characterized in that the learning method identifies conversation patterns using conversation frequency analysis or topic modeling.

[0932] "Example 1"

[0933] (Claim 1)

[0934] A means for acquiring audio data and reducing noise,

[0935] A data transmission means for encoding acquired audio data and transmitting it through a secure protocol,

[0936] A data analysis means for analyzing transmitted audio using speech recognition technology, converting speech into text, and identifying the speaker,

[0937] A conversation analysis method for accumulating transcribed conversations and learning conversation patterns using natural language processing technology,

[0938] A means of providing activities to select and deliver appropriate interactive activities based on the context of the conversation and the profiles of the participants,

[0939] A means of organizing and distributing information to organize family events and health information using cloud technology, and notifying users on their devices as needed.

[0940] Includes system.

[0941] (Claim 2)

[0942] The system according to claim 1, characterized in that the data analysis means identifies the speaker's voice characteristics using speech recognition technology.

[0943] (Claim 3)

[0944] The system according to claim 1, characterized in that the conversation analysis means identifies conversation patterns using conversation frequency analysis and topic modeling methods.

[0945] "Application Example 1"

[0946] (Claim 1)

[0947] A voice analysis method for acquiring voice data in real time and identifying the characteristics of the speaker,

[0948] A learning method for analyzing past conversation history and learning conversation patterns,

[0949] A means of providing and displaying interactive games to participants,

[0950] A means of managing information to collect and organize family events and health information,

[0951] An information distribution means for notifying user terminals of organized information,

[0952] A situation monitoring system that analyzes audio data to monitor security status and issues alerts when an anomaly is detected,

[0953] An educational tool for learning safety measures through interactive games.

[0954] A system that includes this.

[0955] (Claim 2)

[0956] The system according to claim 1, characterized in that the voice analysis means identifies the speaker's voiceprint using a voice recognition algorithm.

[0957] (Claim 3)

[0958] The system according to claim 1, characterized in that the learning method identifies conversation patterns using conversation frequency analysis or topic modeling.

[0959] "Example 2 of combining an emotion engine"

[0960] (Claim 1)

[0961] A voice analysis means for instantly acquiring voice information and identifying the characteristics of the speaker,

[0962] An analytical means for analyzing past dialogue history and learning dialogue patterns,

[0963] An entertainment provision method for providing and presenting interactive games to participants,

[0964] A means of organizing information for collecting and organizing information about household events and health-related information,

[0965] A means of transmitting information to notify the user's terminal of organized information,

[0966] An emotion evaluation method for evaluating the emotional state of a speaker, utilizing an emotion recognition engine,

[0967] A proposal generation means for dynamically generating dialogue and entertainment suggestions based on emotional state.

[0968] Includes system.

[0969] (Claim 2)

[0970] The system according to claim 1, characterized in that the voice analysis means identifies the voiceprint of the speaker using voice recognition technology.

[0971] (Claim 3)

[0972] The system according to claim 1, characterized in that the analysis means identifies dialogue patterns using dialogue frequency analysis and topic classification.

[0973] "Application example 2 when combining with an emotional engine"

[0974] (Claim 1)

[0975] A voice analysis method for acquiring voice data in real time and identifying the user's characteristics,

[0976] A learning method for analyzing past conversation history and learning conversation patterns,

[0977] An entertainment provision means for providing and displaying interactive entertainment to the user,

[0978] A means of information management for collecting and organizing household events and health information,

[0979] An information distribution means for notifying user terminals of organized information,

[0980] A means of providing product recommendations based on the user's emotional state using an emotion recognition engine.

[0981] Includes system.

[0982] (Claim 2)

[0983] The system according to claim 1, characterized in that the voice analysis means identifies the user's voice characteristics using a voice recognition algorithm.

[0984] (Claim 3)

[0985] The system according to claim 1, characterized in that the learning method identifies conversation patterns using conversation frequency analysis or topic modeling. [Explanation of Symbols]

[0986] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A voice analysis method for acquiring voice data in real time and identifying the characteristics of the speaker, A learning method for analyzing past conversation history and learning conversation patterns, A means of providing and displaying interactive games to participants, A means of managing information to collect and organize family events and health information, A means of distributing information to notify user terminals of organized information. Includes system.

2. The system according to claim 1, characterized in that the voice analysis means identifies the speaker's voiceprint using a voice recognition algorithm.

3. The system according to claim 1, characterized in that the learning method identifies conversation patterns using conversation frequency analysis or topic models.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A