system

The system addresses the issue of superficial matchmaking by analyzing voice data to recommend compatible partners, improving relationship quality through emotional and communication compatibility analysis.

JP2026070922APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Conventional matchmaking services rely on superficial user profile information, leading to mismatches and relationship disharmony due to the inability to accurately assess emotional compatibility and communication styles.

Method used

A system that analyzes users' voice data to determine emotions, preferences, and communication styles, calculates compatibility scores, and provides partner recommendations, while continuously monitoring interactions for relationship improvement.

Benefits of technology

Enhances relationship compatibility by suggesting suitable partners based on detailed emotional and communication analysis, supporting a happy married life through personalized and continuous relationship support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070922000001_ABST
    Figure 2026070922000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of obtaining voice data from the user, A means for analyzing the aforementioned audio data to generate user voiceprint information and transcribed dialogue content, A means of analyzing the user's emotions, preferences, and communication style from the generated dialogue content and voiceprint information, Based on the aforementioned analysis results, a means for selecting compatible partner candidates, A means of providing the user with information on the aforementioned potential partners, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of this disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern times, it is a very difficult problem to evaluate emotional compatibility and communication compatibility in the selection of a marriage partner. Since conventional matching services rely on profile information provided by users, they tend to depend on superficial data and cannot fully reflect actual emotions and communication styles. For this reason, the matching with partners having low compatibility increases, often leading to disharmony after marriage and the breakdown of the relationship. Therefore, there is a need for a technology that can more accurately determine compatibility using deep information based on users' voice data and conversation content.

Means for Solving the Problems

[0005] This invention provides a system that analyzes a user's emotions, preferences, and communication style by acquiring voice data from the user, analyzing this voice data to generate the user's voiceprint information and transcribed dialogue content. Based on the analysis results, it selects and provides compatible partner candidates to the user. Furthermore, it achieves more accurate matching by calculating a compatibility score using past user data. In addition, it includes means for continuously monitoring the user's dialogue and generating advice that helps improve the relationship, thereby supporting a happy married life.

[0006] "Voice data" refers to digital data that records the user's voice.

[0007] "Voiceprint information" refers to data that shows the characteristics of an individual's voice, analyzed from audio data.

[0008] "Transcripted dialogue content" refers to audio data converted into text.

[0009] "Emotions" refer to the psychological state inferred from the user's voice and dialogue.

[0010] "Preferences" refer to the user's tastes and interests, which are estimated from the content of conversations and voiceprint information.

[0011] "Communication style" refers to the patterns of how users send and receive information.

[0012] A "potential partner" refers to another user who has been determined to be a good match for the user and is eligible for matching.

[0013] A "compatibility score" is a numerical representation of the compatibility between users, calculated based on in-depth information and historical data.

[0014] "Advice" refers to suggestions that a system provides to a user that are helpful in improving or maintaining relationships.

Brief Description of the Drawings

[0015] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Modes for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is a next-generation matchmaking system that uses the user's voice data to recommend compatible partners. This system acquires the user's voice data and, through analysis, derives their emotions, preferences, and communication style, thereby suggesting a partner who will help them maintain a happy life after marriage.

[0037] Program processing

[0038] 1. Data Acquisition and Analysis

[0039] The user uses the app to perform voice input. Through a series of conversations, the app elicits the user's thoughts and intentions.

[0040] The device sends the user's voice data to the server. The voice is recorded in real time and transferred to the server via streaming technology.

[0041] The server uses speech recognition technology to convert the voice data into text. Based on this text, a generating AI analyzes the user's voiceprint information to determine their emotions, preferences, and communication style.

[0042] 2. Compatibility assessment and partner selection

[0043] The server uses these analysis results to enrich user profiles. Furthermore, it compares them with information on other users stored in the database to select suitable partner candidates.

[0044] The server calculates a compatibility score and lists the candidates with the highest scores.

[0045] 3. Providing suggestions and advice

[0046] The server proposes selected partner candidates to the user via the terminal.

[0047] Users can review the candidates and decide whether or not to interact with the suggested individuals.

[0048] The server continuously monitors user interaction patterns and feedback, and generates advice for improvement, thereby supporting the deepening of relationships.

[0049] Specific example

[0050] Example 1: Proposal for a potential partner

[0051] Input an audio recording of User A happily discussing their vacation plans.

[0052] The server analyzes user A's voiceprint and text to determine their characteristics, such as their love of travel and enjoyment of conversation.

[0053] The server identifies user B, who has similar hobbies and communication styles, as a candidate and suggests it to user A via the terminal.

[0054] User A is interested in the candidate, and a match is made.

[0055] Example 2: Communication support

[0056] The server analyzes the conversation between user A and user B after matching.

[0057] Based on the conversation, the server sends advice to user A's terminal recommending that they "try to get the other person to talk a little more."

[0058] User A puts the advice into practice, and communication proceeds smoothly.

[0059] In this way, this system analyzes the user's psychological characteristics and recommends an ideal partner. Furthermore, it supports the creation of happy relationships by improving continuous communication.

[0060] The following describes the processing flow.

[0061] Step 1:

[0062] Users register through the application. They enter basic information such as their name, age, gender, and hobbies, and then press the submit button.

[0063] Step 2:

[0064] The terminal formats the input data from the user and sends it to the server via secure communication.

[0065] Step 3:

[0066] The server stores the received user information in a database. This registers the user's basic profile within the system.

[0067] Step 4:

[0068] The user activates the device's voice input function to begin an initial dialogue session with the AI.

[0069] Step 5:

[0070] The device records the user's voice in real time and sends the audio data to the server in streaming format.

[0071] Step 6:

[0072] The server processes the received audio data through a speech recognition engine and converts it into text data.

[0073] Step 7:

[0074] The server uses a generation AI to analyze the transcribed dialogue content and voiceprint information to understand the user's emotions, preferences, and communication style.

[0075] Step 8:

[0076] Based on the analysis results, the server updates the user's profile and stores in-depth information in the database.

[0077] Step 9:

[0078] The server runs an algorithm that compares the database with other users' databases to select compatible partner candidates.

[0079] Step 10:

[0080] The server calculates a compatibility score, generates a list of potential partners with high scores, and sends it to the terminal.

[0081] Step 11:

[0082] The device displays proposed partner candidates to the user and provides an interface that allows for visual confirmation.

[0083] Step 12:

[0084] Users review the list of candidates and accept or reconsider matching with those they are interested in.

[0085] Step 13:

[0086] The server continuously monitors the conversations of matched pairs and generates advice to improve the quality of their communication.

[0087] Step 14:

[0088] The device periodically notifies users with advice to help them achieve better communication.

[0089] (Example 1)

[0090] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0091] In modern society, there is a growing need for more efficient relationship building between individuals. In particular, selecting the optimal partner while understanding each person's personality and preferences is not easy. Against this backdrop, there is a need to develop a system that utilizes voice information to provide more personalized suggestions.

[0092] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0093] In this invention, the server includes a device for acquiring voice information, a device for analyzing the voice information to generate voice characteristic information and written communication content, and a device for analyzing feelings, preferences, and communication styles from the generated communication content and voice characteristic information. This enables sophisticated selection of potential partners based on the individual characteristics of the user.

[0094] "Audio information" refers to sound data acquired for the purpose of recording and analyzing the user's speech.

[0095] "Voice feature information" refers to data that indicates the characteristics of a voice, obtained by analyzing voice information.

[0096] "Transcripted communication content" refers to the content of a conversation that has been generated from audio information and expressed in written form.

[0097] "Feelings" refers to the emotional state or sensations of the user.

[0098] "Preferences" is a concept that refers to the user's likes and interests.

[0099] "Communication style" refers to the methods and styles in which users communicate with others.

[0100] "Potential partners" refers to other users selected by the system who may be a good match for the user.

[0101] A "fitness index" is a numerical indicator that quantifies the degree of compatibility between a user's characteristics and the system.

[0102] "Advice" refers to suggestions and guidance provided to improve the relationship between the user and the potential partner.

[0103] This system is configured to suggest the most suitable partner based on the user's voice information and to support the continuous improvement of the relationship. Specifically, the invention will be implemented in the following form.

[0104] Users install a dedicated application on their device and use the voice input function to naturally express their thoughts and feelings. The voice information is recorded in real time by the device and transmitted to a server using streaming technology. This process incorporates advanced security and privacy protection features.

[0105] The server converts speech information into text using speech recognition software such as Google® Speech-to-Text API. Next, a generative AI model is used to analyze the transcribed conversation content and speech feature information. This analysis identifies the user's feelings, preferences, and communication style. The generative AI model extracts these characteristics by utilizing natural language processing techniques and sentiment analysis algorithms.

[0106] The analysis results are stored in a database on the server and used to enrich user profiles. These results are then compared with data from other users to calculate a suitability index. Based on this index, the server selects the most suitable candidates and sends the candidate list to the user's terminal.

[0107] Users can review candidate information displayed on their device and select those they are interested in. After selection, the server continuously monitors the user's interactions and provides advice using AI-generated content. This advice is designed to improve the quality of conversations and help deepen relationships.

[0108] For example, if a user inputs a voice message "talking happily about travel," the server will extract characteristics such as a love of travel and enjoyment of conversation, and suggest partners with similar interests. An example of a prompt message is, "Analyze the user's voice data and suggest the best partner based on emotions and preferences. Provide advice based on that and teach me how to improve communication." This system promotes interaction with others and supports the building of richer relationships.

[0109] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0110] Step 1:

[0111] The user launches a dedicated app and inputs voice information into the device. The voice information consists of words expressing the user's thoughts and feelings. The device records this in real time and sends the voice data to the server. In this process, the input is voice data, and the output is the transfer of voice data to the server.

[0112] Step 2:

[0113] The server analyzes the received audio data. First, it uses speech recognition software to convert the audio data into text data. This conversion yields audio feature information and the transcribed content of the conversation. Here, the input is audio data, and the output is text data.

[0114] Step 3:

[0115] The server uses a generative AI model to analyze text data. This analysis employs natural language processing techniques to extract user feelings, preferences, and communication styles from the text. The input is text data, and the output is the analyzed user profile information.

[0116] Step 4:

[0117] The server updates the user profile based on the analysis results. It then compares this profile with data from other users and calculates a compatibility index. This involves database searches and statistical analysis techniques. The input is the analyzed user profile information, and the output is the compatibility score.

[0118] Step 5:

[0119] The server selects the most suitable partner candidate based on the compatibility score. It then lists the information of the selected candidates and sends it to the user's terminal. The input is the compatibility score, and the output is the candidate list.

[0120] Step 6:

[0121] The user views a list of suggested candidates on their device and selects those they are interested in. The selections are fed back to the server. The input here is the candidate list, and the output is the feedback from the selected individuals.

[0122] Step 7:

[0123] After selection, the server continuously monitors user communication content. It utilizes generative AI to analyze the communication content and generate specific advice. This provides advice to improve the quality of user interactions. Inputs are feedback and communication content, while output is advice for communication improvement.

[0124] (Application Example 1)

[0125] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0126] In recent years, with the advancement of digital communication technology, there has been a growing demand from users for real-time, personalized experiences. However, existing systems struggle to analyze users' emotions and interests in real time and provide immediate, appropriate suggestions. As a result, users face the challenge of a reduced quality of experience due to unnatural pauses in communication and inappropriate suggestions.

[0127] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0128] In this invention, the server includes means for collecting voice data from the user in real time, means for analyzing the user's emotions and interests from the voice data in real time, and means for making appropriate suggestions according to the situation. As a result, the user can receive situation-appropriate suggestions in real time and enjoy a more personalized and high-quality communication experience.

[0129] A "user" is an individual or group that utilizes a particular service or system.

[0130] "Audio data" refers to information recorded in digital format from the voice spoken by a user.

[0131] "Means of real-time data collection" refers to technologies or methods that enable the immediate acquisition of information provided by users.

[0132] "Voiceprint information" refers to biometric authentication-related information obtained by analyzing the characteristics of a user's voice.

[0133] "Transcripted dialogue content" refers to information in text format extracted from audio data.

[0134] "Emotions" refer to information that indicates the user's psychological state.

[0135] "Preferences" refer to information about a user's likes and interests regarding a particular subject.

[0136] "Communication style" refers to the unique behaviors and patterns that users exhibit when exchanging information with others.

[0137] A "compatible partner" is another individual or digital agent that possesses characteristics that match the user's requirements and features.

[0138] "A means of presenting candidates to users in real time and enabling them to access dialogue" refers to a technology that provides users with immediately visible information on selected candidates, allowing them to interact with them directly.

[0139] "A means of analyzing a user's emotions and interests in real time from the user's voice data during a conversation" refers to a technology that instantly interprets voice information obtained during interaction and identifies the user's emotional and interest tendencies.

[0140] "Means of making appropriate suggestions according to the situation" refers to technologies or methods that present beneficial options or actions according to the environment and the user's state.

[0141] The system of the present invention acquires voice data from a smart device worn by the user or a terminal used by the user, analyzes it on a server, and provides the user with appropriate suggestions in real time.

[0142] The server uses streaming technology to acquire user voice data and converts the data to text using a speech recognition API. Specific hardware includes smart glasses and ear-mounted microphones, which collect voice with high accuracy. The software can utilize commonly used speech recognition APIs such as Google Cloud Speech-to-Text and Azure Cognitive Services for sentiment analysis. This allows for real-time analysis of user emotions, preferences, and communication styles from their voice, enabling appropriate suggestions tailored to the user's behavior in the digital environment.

[0143] When a user initiates a conversation in the digital space, the server collects the voice data and analyzes their emotions and interests. This allows it to make real-time suggestions to help them achieve their desired experience. These suggestions are continuously improved based on user feedback, and it's even possible to suggest compatible digital partners by comparing them with past data.

[0144] For example, if a user is talking about action movies during a virtual date, the server can analyze the audio and recommend new action movies that might interest them. The generative AI model used in this process includes the ability to analyze the flow of the conversation and generate appropriate options, and examples of prompts include the following:

[0145] "A user is talking about action movies during a virtual date. Analyze the user's comments and figure out how to recommend new action movies that might interest them."

[0146] Thus, the present invention enables high-quality, real-time interaction by proposing personalized digital experiences tailored to the user.

[0147] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0148] Step 1:

[0149] The user inputs voice data via a smart device. The device collects this voice data in real time and sends it to a server using streaming technology. In this process, the voice signal is converted into digital data and transmitted instantly.

[0150] Step 2:

[0151] The server inputs the received audio data into a speech recognition API, which converts the audio into text. This speech recognition API quickly converts the audio data into a string of characters. The output is the user's spoken content formatted as text.

[0152] Step 3:

[0153] The server uses a generative AI model to analyze the user's emotions, preferences, and communication style from the transcribed data. The input is the text data obtained in step 2, and the output is the result of analyzing its emotional and semantic nuances. This analysis clearly reveals the user's psychological state and interests.

[0154] Step 4:

[0155] Based on the analysis results, the server selects compatible digital content and partners from the available database and generates specific content. The input is the analysis results from step 3, and the output is content and action plans to be presented to the user. The generative AI model is also used here to derive options based on the prompt text.

[0156] Step 5:

[0157] The server instantly delivers selected content and information about the other party to the user's smart device and presents it as a user interface. Options for the next action or experience are displayed on the device, and the user uses these to proceed with the interaction.

[0158] Step 6:

[0159] When a user selects or reacts to presented content, the device sends this feedback back to the server. Using the returned feedback data, the server further learns the user's preferences and continuously improves its analysis algorithms.

[0160] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0161] This invention is a next-generation matchmaking system that analyzes users' emotions in detail and recommends compatible partners. This system analyzes the user's voiceprint information and transcribed dialogue content, and uses an emotion engine to achieve more sophisticated matching.

[0162] Program processing

[0163] 1. Data acquisition and voice analysis

[0164] The user launches the application and inputs voice commands into the device. They then engage in real-time dialogue, providing the system with their thoughts and wishes.

[0165] The terminal immediately transfers the acquired audio data to the server.

[0166] The server converts the audio to text and extracts voiceprint information. The generated text is then subjected to analysis by an emotion engine.

[0167] 2. Sentiment analysis and profile updates

[0168] The server uses an emotion engine to identify the user's emotions from the text. It recognizes the specific emotional state and reflects it in the profile.

[0169] We collect user sentiment data and compare it with past history to understand long-term trends.

[0170] 3. Partner Selection

[0171] Based on the analysis results, the server selects compatible partner candidates from the database, taking into account emotional characteristics and communication style compatibility.

[0172] We utilize past data to calculate compatibility scores and narrow down the candidates.

[0173] 4. Information provision and two-way advice

[0174] The device presents the user with selected partner candidates. Information is provided in a visually easy-to-understand format.

[0175] The server monitors user interactions and generates advice based on real-time sentiment analysis. It also sends timely notifications to the device to support effective communication.

[0176] Specific example

[0177] Example 1: Matching based on sentiment analysis

[0178] User A talks about stress through the app. The server analyzes User A's anxiety levels from their voiceprint and words. It also identifies that User A likes relaxation.

[0179] The server selects User B, who has resourceful hobbies, as a suitable partner for this profile, suggesting a high level of compatibility.

[0180] The device provides this information to user A, and user A becomes interested in user B.

[0181] Example 2: Providing real-time advice

[0182] User A and User B begin a conversation after being matched. The server monitors the conversation.

[0183] During a conversation, the emotion engine detects changes in user A's emotions (e.g., tension) and suggests topics to help them relax.

[0184] The device informs user A of this and supports the smooth progress of the conversation.

[0185] This system allows for a detailed analysis of the user's psychological characteristics, recommends an ideal partner, and continuously improves the relationship.

[0186] The following describes the processing flow.

[0187] Step 1:

[0188] The user opens the application and inputs voice commands into the device. The user can then introduce themselves or talk about topics that interest them.

[0189] Step 2:

[0190] The terminal receives voice input and captures it as audio data. At the same time, it prepares this audio data for transmission to the server.

[0191] Step 3:

[0192] The server receives the voice data sent from the terminal and uses a speech recognition engine to convert the data into text.

[0193] Step 4:

[0194] The server inputs text data into the emotion engine, which then analyzes the user's emotional state. This process identifies specific emotions such as joy, sadness, and anger.

[0195] Step 5:

[0196] The server analyzes voiceprint information to extract the unique characteristics of the user's voice. This allows for a more detailed understanding of the user's characteristics.

[0197] Step 6:

[0198] The server updates the user's profile using acquired sentiment data and voiceprint information. This includes emotional tendencies and preferences.

[0199] Step 7:

[0200] The server performs a database search based on profile data to select compatible partner candidates. The selection focuses on emotional compatibility and matching communication styles.

[0201] Step 8:

[0202] The server evaluates the selected partner candidates and calculates a compatibility score. This quantifies the compatibility between users.

[0203] Step 9:

[0204] The device receives information about potential partners along with compatibility scores from the server and provides it to the user. This information is displayed in a visually and intuitively easy-to-understand format.

[0205] Step 10:

[0206] Users review the information provided about potential partners and make a decision about matching. This is expressed as a contact request to the preferred candidate.

[0207] Step 11:

[0208] If a user match is successful, the server will continue to monitor the interaction between both parties in real time.

[0209] Step 12:

[0210] The server utilizes an emotion engine to track changes in emotions during conversations and generates communication advice as needed.

[0211] Step 13:

[0212] The device notifies the user of the generated advice, encouraging positive communication.

[0213] (Example 2)

[0214] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0215] In today's society, accurately understanding one's own emotional state and preferences, and finding a suitable partner based on that understanding, is crucial for building deeper relationships. However, conventional matching systems struggle to accurately analyze users' emotional states, limiting their ability to find compatible partner candidates. Furthermore, there is a lack of systems that provide appropriate advice to support communication after matching. Therefore, there is a need for a system that can more accurately and effectively select compatible partner candidates and support communication.

[0216] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0217] In this invention, the server includes means for acquiring voice data from a user, means for analyzing the voice data to generate the user's acoustic characteristic information and transcribed dialogue content, and means for analyzing the user's emotional state, preferences, and dialogue tendencies using the generated dialogue content and acoustic characteristic information. This makes it possible to accurately grasp the user's emotional state and preferences, select appropriate compatible relationship candidates based on them, and provide advice in real time that corresponds to emotional changes that occur during the conversation.

[0218] "Voice data" refers to information recorded in digital format from the voice signals spoken by the user.

[0219] "Acoustic feature information" refers to information that includes features such as spectrum and frequency extracted from audio data.

[0220] "Transcripted dialogue content" refers to linguistic information in text format converted from audio data.

[0221] A "generative AI model" refers to a technology that uses artificial intelligence algorithms learned from vast datasets to perform analysis and predictions that utilize the characteristics of emotions and preferences.

[0222] "Compatible social relationship candidates" refer to other users who are considered to have a high affinity with the user, selected based on the user's emotional state and preferences.

[0223] A "fit score" refers to a numerical value used to quantitatively evaluate the degree of compatibility between users.

[0224] "Monitoring" refers to the process of observing and recording user conversations in real time.

[0225] "Means of generating advice" refers to the ability to create suggestions to support user communication based on analysis results.

[0226] This invention relates to a system that precisely analyzes a user's emotions and preferences and recommends a compatible partner. This system primarily involves acquiring and analyzing the user's voice data, and performing a series of processes to support dialogue. Specific embodiments are described below.

[0227] Data acquisition and analysis

[0228] First, the user uses an application installed on their mobile device to input voice into the system. During this process, the user can freely express their thoughts and wishes.

[0229] The device instantly captures voice data from the user and transfers it to a cloud server using wireless communication. This transfer utilizes an internet connection.

[0230] Subsequently, the server converts the audio data into text data using automatic speech recognition software (e.g., an existing API for speech-to-text conversion). Based on the resulting text data and acoustic feature information, the server uses a generative AI model to analyze the user's emotional state, preferences, and dialogue tendencies.

[0231] Partner selection and provision

[0232] Based on the analysis results, the server selects candidates for highly compatible social relationships. Specifically, it matches emotional and preference data using prompt messages to find compatible candidates.

[0233] Information on selected candidates is presented visually to the user via their device. This allows the user to easily review the details of the candidates and choose those they are interested in.

[0234] Specific example

[0235] For example, if a user says, "I'm tired from work today, but I want to relax this weekend," the emotion engine will analyze the fatigue and desire for relaxation. Example prompt: "Recommend other users with relevant hobbies or activities when the user is seeking relaxation."

[0236] This embodiment enables sophisticated matching that takes into account the emotional characteristics of users, and also supports the building of good relationships by assisting with communication after the interaction.

[0237] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0238] Step 1:

[0239] The user launches the app on their mobile device and inputs voice data. What the user speaks is recorded in real time by the mobile device. Input: User's voice. Output: Recorded voice data.

[0240] Step 2:

[0241] The device acquires the recorded audio data and sends it to a cloud server via an internet connection. The audio data is converted to a compressed digital file format on the device. Input: Recorded audio data. Output: Audio data file sent to the server.

[0242] Step 3:

[0243] The server converts received audio data into text data using speech recognition software. In addition, it analyzes acoustic features to extract voiceprint information. Data processing includes analysis of audio waveform data and digital processing of the audio signal. Input: Audio data file sent to the server. Output: Transcripted dialogue content and voiceprint information.

[0244] Step 4:

[0245] The server utilizes a generative AI model to evaluate the user's emotional state, preferences, and communication style from transcribed dialogue content and voiceprint information. Natural language processing algorithms are used for data calculation to determine emotional tendencies. Input: Transcribed dialogue content, voiceprint information. Output: Analyzed emotional state, preferences, and communication style.

[0246] Step 5:

[0247] The server searches a database of past data based on the prompt message and selects suitable social relationship candidates. The server calculates a fit score and performs analysis to identify candidates. Input: Analyzed emotional state, preferences, and prompt message. Output: Suitable relationship candidates and fit score.

[0248] Step 6:

[0249] The terminal presents the user with information on potential relationships received from the server. It displays compatibility scores and candidate profiles in a visually easy-to-understand format. Input: Compatible relationship candidates, suitability scores. Output: Candidate information presented to the user.

[0250] Step 7:

[0251] The server monitors the content of the conversation between users after matching in real time and generates advice to facilitate communication using a generated AI model. It supports users by sending the advice to their devices in a timely manner. Input: Content of the conversation between users. Output: Generated communication advice.

[0252] (Application Example 2)

[0253] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0254] In today's world, a vast array of content exists, and there is a need to efficiently select and provide content that is suitable for each individual user. However, recommending appropriate content based on users' emotions and preferences is difficult, and users may not get the experience they desire.

[0255] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0256] In this invention, the server includes means for acquiring voice information from the user, means for analyzing the voice information to generate the user's voiceprint information and transcribed dialogue information, means for analyzing the user's emotions, preferences, and communication style from the generated dialogue information and voiceprint information, means for selecting appropriate information according to the emotional state, and means for providing the information to the user. This makes it possible to recommend optimal content based on the user's emotions.

[0257] "Audio information" refers to sound data obtained from the user, including information about speech and voice characteristics.

[0258] "Voiceprint information" refers to data that shows the unique vocal characteristics of an individual, extracted from audio information.

[0259] "Textualized dialogue information" refers to the content of the dialogue in string format, which is obtained by analyzing and converting the user's voice.

[0260] "Emotions" refer to the psychological state analyzed from the user's voice and dialogue content.

[0261] "Preferences" refer to information that indicates the subjects or things that a user is interested in or concerned with.

[0262] "Communication style" refers to information about the methods and styles of communication used by users.

[0263] "Emotional state" refers to a user's temporary psychological state, including the type and intensity of their emotions.

[0264] "Means of selecting information" refers to a method of selecting data appropriate for the user based on their analyzed emotional state.

[0265] "Means of providing information" refers to the methods and interfaces used to present selected data to users.

[0266] As an application example of this invention, a system is realized that recommends content based on the user's emotional state. The system uses a smartphone or smart glasses. First, the user inputs voice information into the smart device. This voice information is sent from the device to a server via the internet. The server uses speech recognition software (e.g., Google Cloud Speech-to-Text API) to convert the voice into textual dialogue information. Furthermore, the server uses sentiment analysis software (e.g., IBM Watson® Tone Analyzer) to identify the user's emotional state from the text.

[0267] Based on the identified emotional state, the server selects content appropriate for the user. This selection uses a relevance score based on past user data and data from similar users. Once content selection is complete, the server sends the selection results to the terminal, and the user receives the content. The received data is then displayed visually, allowing the user to experience it.

[0268] As a concrete example, suppose a user voice-inputs "I'm tired today." In this case, the system determines that the emotional state is "fatigue" and recommends relaxing music or stress-relieving videos to alleviate that state.

[0269] An example of a prompt is, "Please recommend music that can help me relax when I feel stressed." By inputting this prompt into a generating AI model, the system can quickly provide content that is suitable for the user.

[0270] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0271] Step 1:

[0272] The user inputs voice information using a smart device. This voice information is collected by the device and transmitted directly to the server. The input is voice data, and the output from the device to the server is also voice data.

[0273] Step 2:

[0274] The server uses speech recognition software to convert the received audio information into text. During this process, data processing is performed to convert the audio data into string data. The output is the textualized dialogue information.

[0275] Step 3:

[0276] The server processes the transcribed dialogue information using sentiment analysis software to extract the user's emotional state. The data calculation performed here involves assigning emotion labels through text analysis. The input is text information, and the output is emotional state data.

[0277] Step 4:

[0278] The server selects appropriate content from its content database based on the user's emotional state. The selection process uses accumulated data to calculate a relevance score. The input is emotional state data, and the output is information about the recommended content.

[0279] Step 5:

[0280] The server transmits selected content information to the terminal, which then visualizes and presents that information to the user. Here, visualization is used to communicate the data to the user. The input is content information, and the output is display data for the user.

[0281] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0282] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">)Generative AIs such as etc. can be mentioned. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input into the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is input. The data generation model 58 infers the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summary, etc.

[0283] In the above embodiment, a form example in which specific processing is performed by the data processing device 12 is given, but the technology of the present disclosure is not limited to this, and specific processing may be performed by the smart device 14.

[0284] [Second Embodiment]

[0285] FIG. 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0286] As shown in FIG. 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0287] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of the "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network), etc.

[0288] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0289] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0290] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0291] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0292] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0293] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0294] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0295] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0296] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0297] This invention is a next-generation matchmaking system that uses the user's voice data to recommend compatible partners. This system acquires the user's voice data and, through analysis, derives their emotions, preferences, and communication style, thereby suggesting a partner who will help them maintain a happy life after marriage.

[0298] Program processing

[0299] 1. Data Acquisition and Analysis

[0300] The user uses the app to perform voice input. Through a series of conversations, the app elicits the user's thoughts and intentions.

[0301] The terminal sends the user's voice data to the server. The voice is recorded in real time and transferred to the server through streaming technology.

[0302] The server uses voice recognition technology to convert the voice data into text. Based on this text, the generative AI analyzes the user's voiceprint information, emotions, preferences, and communication styles.

[0303] 2. Compatibility Judgment and Partner Selection

[0304] The server uses these analysis results to enrich the user's profile. Furthermore, it compares with the information of other users stored in the database to select suitable partner candidates.

[0305] The server calculates the compatibility score and lists the candidates with high scores.

[0306] 3. Proposal and Advice Provision

[0307] The server proposes the selected partner candidates to the user through the terminal.

[0308] The user can confirm the candidates and decide whether to communicate with the proposed partners.

[0309] The server continuously monitors the user's dialogue patterns and feedback, and generates advice for improvement to support the deepening of the relationship.

[0310] Specific Example

[0311] Example 1: Proposal of Partner Candidates

[0312] User A inputs voice talking cheerfully about vacation plans.

[0313] The server analyzes user A's voiceprint and text to determine their characteristics, such as their love of travel and enjoyment of conversation.

[0314] The server identifies user B, who has similar hobbies and communication styles, as a candidate and suggests it to user A via the terminal.

[0315] User A is interested in the candidate, and a match is made.

[0316] Example 2: Communication support

[0317] The server analyzes the conversation between user A and user B after matching.

[0318] Based on the conversation, the server sends advice to user A's terminal recommending that they "try to get the other person to talk a little more."

[0319] User A puts the advice into practice, and communication proceeds smoothly.

[0320] In this way, this system analyzes the user's psychological characteristics and recommends an ideal partner. Furthermore, it supports the creation of happy relationships by improving continuous communication.

[0321] The following describes the processing flow.

[0322] Step 1:

[0323] Users register through the application. They enter basic information such as their name, age, gender, and hobbies, and then press the submit button.

[0324] Step 2:

[0325] The terminal formats the input data from the user and sends it to the server via secure communication.

[0326] Step 3:

[0327] The server stores the received user information in a database. This registers the user's basic profile within the system.

[0328] Step 4:

[0329] The user activates the device's voice input function to begin an initial dialogue session with the AI.

[0330] Step 5:

[0331] The device records the user's voice in real time and sends the audio data to the server in streaming format.

[0332] Step 6:

[0333] The server processes the received audio data through a speech recognition engine and converts it into text data.

[0334] Step 7:

[0335] The server uses a generation AI to analyze the transcribed dialogue content and voiceprint information to understand the user's emotions, preferences, and communication style.

[0336] Step 8:

[0337] Based on the analysis results, the server updates the user's profile and stores in-depth information in the database.

[0338] Step 9:

[0339] The server runs an algorithm that compares the database with other users' databases to select compatible partner candidates.

[0340] Step 10:

[0341] The server calculates a compatibility score, generates a list of potential partners with high scores, and sends it to the terminal.

[0342] Step 11:

[0343] The device displays proposed partner candidates to the user and provides an interface that allows for visual confirmation.

[0344] Step 12:

[0345] Users review the list of candidates and accept or reconsider matching with those they are interested in.

[0346] Step 13:

[0347] The server continuously monitors the conversations of matched pairs and generates advice to improve the quality of their communication.

[0348] Step 14:

[0349] The device periodically notifies users with advice to help them achieve better communication.

[0350] (Example 1)

[0351] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0352] In modern society, there is a growing need for more efficient relationship building between individuals. In particular, selecting the optimal partner while understanding each person's personality and preferences is not easy. Against this backdrop, there is a need to develop a system that utilizes voice information to provide more personalized suggestions.

[0353] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0354] In this invention, the server includes a device for acquiring voice information, a device for analyzing the voice information to generate voice characteristic information and written communication content, and a device for analyzing feelings, preferences, and communication styles from the generated communication content and voice characteristic information. This enables sophisticated selection of potential partners based on the individual characteristics of the user.

[0355] "Audio information" refers to sound data acquired for the purpose of recording and analyzing the user's speech.

[0356] "Voice feature information" refers to data that indicates the characteristics of a voice, obtained by analyzing voice information.

[0357] "Transcripted communication content" refers to the content of a conversation that has been generated from audio information and expressed in written form.

[0358] "Feelings" refers to the emotional state or sensations of the user.

[0359] "Preferences" is a concept that refers to the user's likes and interests.

[0360] "Communication style" refers to the methods and styles in which users communicate with others.

[0361] "Potential partners" refers to other users selected by the system who may be a good match for the user.

[0362] A "fitness index" is a numerical indicator that quantifies the degree of compatibility between a user's characteristics and the system.

[0363] "Advice" refers to suggestions and guidance provided to improve the relationship between the user and the potential partner.

[0364] This system is configured to suggest the most suitable partner based on the user's voice information and to support the continuous improvement of the relationship. Specifically, the invention will be implemented in the following form.

[0365] Users install a dedicated application on their device and use the voice input function to naturally express their thoughts and feelings. The voice information is recorded in real time by the device and transmitted to a server using streaming technology. This process incorporates advanced security and privacy protection features.

[0366] The server uses speech recognition software, such as the Google Speech-to-Text API, to convert the spoken information into text. Next, a generative AI model is used to analyze the transcribed conversation content and speech feature information. This analysis identifies the user's feelings, preferences, and communication style. The generative AI model extracts these characteristics by utilizing natural language processing techniques and sentiment analysis algorithms.

[0367] The analysis results are stored in a database on the server and used to enrich user profiles. These results are then compared with data from other users to calculate a suitability index. Based on this index, the server selects the most suitable candidates and sends the candidate list to the user's terminal.

[0368] Users can review candidate information displayed on their device and select those they are interested in. After selection, the server continuously monitors the user's interactions and provides advice using AI-generated content. This advice is designed to improve the quality of conversations and help deepen relationships.

[0369] For example, if a user inputs a voice message "talking happily about travel," the server will extract characteristics such as a love of travel and enjoyment of conversation, and suggest partners with similar interests. An example of a prompt message is, "Analyze the user's voice data and suggest the best partner based on emotions and preferences. Provide advice based on that and teach me how to improve communication." This system promotes interaction with others and supports the building of richer relationships.

[0370] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0371] Step 1:

[0372] The user launches a dedicated app and inputs voice information into the device. The voice information consists of words expressing the user's thoughts and feelings. The device records this in real time and sends the voice data to the server. In this process, the input is voice data, and the output is the transfer of voice data to the server.

[0373] Step 2:

[0374] The server analyzes the received audio data. First, it uses speech recognition software to convert the audio data into text data. This conversion yields audio feature information and the transcribed content of the conversation. Here, the input is audio data, and the output is text data.

[0375] Step 3:

[0376] The server uses a generative AI model to analyze text data. This analysis employs natural language processing techniques to extract user feelings, preferences, and communication styles from the text. The input is text data, and the output is the analyzed user profile information.

[0377] Step 4:

[0378] The server updates the user profile based on the analysis results. It then compares this profile with data from other users and calculates a compatibility index. This involves database searches and statistical analysis techniques. The input is the analyzed user profile information, and the output is the compatibility score.

[0379] Step 5:

[0380] The server selects the most suitable partner candidate based on the compatibility score. It then lists the information of the selected candidates and sends it to the user's terminal. The input is the compatibility score, and the output is the candidate list.

[0381] Step 6:

[0382] The user views a list of suggested candidates on their device and selects those they are interested in. The selections are fed back to the server. The input here is the candidate list, and the output is the feedback from the selected individuals.

[0383] Step 7:

[0384] After selection, the server continuously monitors user communication content. It utilizes generative AI to analyze the communication content and generate specific advice. This provides advice to improve the quality of user interactions. Inputs are feedback and communication content, while output is advice for communication improvement.

[0385] (Application Example 1)

[0386] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0387] In recent years, with the advancement of digital communication technology, there has been a growing demand from users for real-time, personalized experiences. However, existing systems struggle to analyze users' emotions and interests in real time and provide immediate, appropriate suggestions. As a result, users face the challenge of a reduced quality of experience due to unnatural pauses in communication and inappropriate suggestions.

[0388] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0389] In this invention, the server includes means for collecting voice data from the user in real time, means for analyzing the user's emotions and interests from the voice data in real time, and means for making appropriate suggestions according to the situation. As a result, the user can receive situation-appropriate suggestions in real time and enjoy a more personalized and high-quality communication experience.

[0390] A "user" is an individual or group that utilizes a particular service or system.

[0391] "Audio data" refers to information recorded in digital format from the voice spoken by a user.

[0392] "Means of real-time data collection" refers to technologies or methods that enable the immediate acquisition of information provided by users.

[0393] "Voiceprint information" refers to biometric authentication-related information obtained by analyzing the characteristics of a user's voice.

[0394] "Transcripted dialogue content" refers to information in text format extracted from audio data.

[0395] "Emotions" refer to information that indicates the user's psychological state.

[0396] "Preferences" refer to information about a user's likes and interests regarding a particular subject.

[0397] "Communication style" refers to the unique behaviors and patterns that users exhibit when exchanging information with others.

[0398] A "compatible partner" is another individual or digital agent that possesses characteristics that match the user's requirements and features.

[0399] "A means of presenting candidates to users in real time and enabling them to access dialogue" refers to a technology that provides users with immediately visible information on selected candidates, allowing them to interact with them directly.

[0400] "A means of analyzing a user's emotions and interests in real time from the user's voice data during a conversation" refers to a technology that instantly interprets voice information obtained during interaction and identifies the user's emotional and interest tendencies.

[0401] "Means of making appropriate suggestions according to the situation" refers to technologies or methods that present beneficial options or actions according to the environment and the user's state.

[0402] The system of the present invention acquires voice data from a smart device worn by the user or a terminal used by the user, analyzes it on a server, and provides the user with appropriate suggestions in real time.

[0403] The server uses streaming technology to acquire user voice data and converts it to text using a speech recognition API. Specific hardware includes smart glasses and ear-mounted microphones, which collect voice with high accuracy. The software can utilize commonly used speech recognition APIs such as Google Cloud Speech-to-Text and Azure Cognitive Services for sentiment analysis. This allows for real-time analysis of user emotions, preferences, and communication styles from their voice, enabling appropriate suggestions tailored to the user's behavior in the digital environment.

[0404] When a user initiates a conversation in the digital space, the server collects the voice data and analyzes their emotions and interests. This allows it to make real-time suggestions to help them achieve their desired experience. These suggestions are continuously improved based on user feedback, and it's even possible to suggest compatible digital partners by comparing them with past data.

[0405] For example, if a user is talking about action movies during a virtual date, the server can analyze the audio and recommend new action movies that might interest them. The generative AI model used in this process includes the ability to analyze the flow of the conversation and generate appropriate options, and examples of prompts include the following:

[0406] "A user is talking about action movies during a virtual date. Analyze the user's comments and figure out how to recommend new action movies that might interest them."

[0407] Thus, the present invention enables high-quality, real-time interaction by proposing personalized digital experiences tailored to the user.

[0408] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0409] Step 1:

[0410] The user inputs voice data via a smart device. The device collects this voice data in real time and sends it to a server using streaming technology. In this process, the voice signal is converted into digital data and transmitted instantly.

[0411] Step 2:

[0412] The server inputs the received audio data into a speech recognition API, which converts the audio into text. This speech recognition API quickly converts the audio data into a string of characters. The output is the user's spoken content formatted as text.

[0413] Step 3:

[0414] The server uses a generative AI model to analyze the user's emotions, preferences, and communication style from the transcribed data. The input is the text data obtained in step 2, and the output is the result of analyzing its emotional and semantic nuances. This analysis clearly reveals the user's psychological state and interests.

[0415] Step 4:

[0416] Based on the analysis results, the server selects compatible digital content and partners from the available database and generates specific content. The input is the analysis results from step 3, and the output is content and action plans to be presented to the user. The generative AI model is also used here to derive options based on the prompt text.

[0417] Step 5:

[0418] The server instantly delivers selected content and information about the other party to the user's smart device and presents it as a user interface. Options for the next action or experience are displayed on the device, and the user uses these to proceed with the interaction.

[0419] Step 6:

[0420] When a user selects or reacts to presented content, the device sends this feedback back to the server. Using the returned feedback data, the server further learns the user's preferences and continuously improves its analysis algorithms.

[0421] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0422] This invention is a next-generation matchmaking system that analyzes users' emotions in detail and recommends compatible partners. This system analyzes the user's voiceprint information and transcribed dialogue content, and uses an emotion engine to achieve more sophisticated matching.

[0423] Program processing

[0424] 1. Data acquisition and voice analysis

[0425] The user launches the application and inputs voice commands into the device. They then engage in real-time dialogue, providing the system with their thoughts and wishes.

[0426] The terminal immediately transfers the acquired audio data to the server.

[0427] The server converts the audio to text and extracts voiceprint information. The generated text is then subjected to analysis by an emotion engine.

[0428] 2. Sentiment analysis and profile updates

[0429] The server uses an emotion engine to identify the user's emotions from the text. It recognizes the specific emotional state and reflects it in the profile.

[0430] We collect user sentiment data and compare it with past history to understand long-term trends.

[0431] 3. Partner Selection

[0432] Based on the analysis results, the server selects compatible partner candidates from the database, taking into account emotional characteristics and communication style compatibility.

[0433] We utilize past data to calculate compatibility scores and narrow down the candidates.

[0434] 4. Information provision and two-way advice

[0435] The device presents the user with selected partner candidates. Information is provided in a visually easy-to-understand format.

[0436] The server monitors user interactions and generates advice based on real-time sentiment analysis. It also sends timely notifications to the device to support effective communication.

[0437] Specific example

[0438] Example 1: Matching based on sentiment analysis

[0439] User A talks about stress through the app. The server analyzes User A's anxiety levels from their voiceprint and words. It also identifies that User A likes relaxation.

[0440] The server selects User B, who has resourceful hobbies, as a suitable partner for this profile, suggesting a high level of compatibility.

[0441] The device provides this information to user A, and user A becomes interested in user B.

[0442] Example 2: Providing real-time advice

[0443] User A and User B begin a conversation after being matched. The server monitors the conversation.

[0444] During a conversation, the emotion engine detects changes in user A's emotions (e.g., tension) and suggests topics to help them relax.

[0445] The device informs user A of this and supports the smooth progress of the conversation.

[0446] This system allows for a detailed analysis of the user's psychological characteristics, recommends an ideal partner, and continuously improves the relationship.

[0447] The following describes the processing flow.

[0448] Step 1:

[0449] The user opens the application and inputs voice commands into the device. The user can then introduce themselves or talk about topics that interest them.

[0450] Step 2:

[0451] The terminal receives voice input and captures it as audio data. At the same time, it prepares this audio data for transmission to the server.

[0452] Step 3:

[0453] The server receives the voice data sent from the terminal and uses a speech recognition engine to convert the data into text.

[0454] Step 4:

[0455] The server inputs text data into the emotion engine, which then analyzes the user's emotional state. This process identifies specific emotions such as joy, sadness, and anger.

[0456] Step 5:

[0457] The server analyzes voiceprint information to extract the unique characteristics of the user's voice. This allows for a more detailed understanding of the user's characteristics.

[0458] Step 6:

[0459] The server updates the user's profile using acquired sentiment data and voiceprint information. This includes emotional tendencies and preferences.

[0460] Step 7:

[0461] The server performs a database search based on profile data to select compatible partner candidates. The selection focuses on emotional compatibility and matching communication styles.

[0462] Step 8:

[0463] The server evaluates the selected partner candidates and calculates a compatibility score. This quantifies the compatibility between users.

[0464] Step 9:

[0465] The device receives information about potential partners along with compatibility scores from the server and provides it to the user. This information is displayed in a visually and intuitively easy-to-understand format.

[0466] Step 10:

[0467] Users review the information provided about potential partners and make a decision about matching. This is expressed as a contact request to the preferred candidate.

[0468] Step 11:

[0469] If a user match is successful, the server will continue to monitor the interaction between both parties in real time.

[0470] Step 12:

[0471] The server utilizes an emotion engine to track changes in emotions during conversations and generates communication advice as needed.

[0472] Step 13:

[0473] The device notifies the user of the generated advice, encouraging positive communication.

[0474] (Example 2)

[0475] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0476] In today's society, accurately understanding one's own emotional state and preferences, and finding a suitable partner based on that understanding, is crucial for building deeper relationships. However, conventional matching systems struggle to accurately analyze users' emotional states, limiting their ability to find compatible partner candidates. Furthermore, there is a lack of systems that provide appropriate advice to support communication after matching. Therefore, there is a need for a system that can more accurately and effectively select compatible partner candidates and support communication.

[0477] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0478] In this invention, the server includes means for acquiring voice data from a user, means for analyzing the voice data to generate the user's acoustic characteristic information and transcribed dialogue content, and means for analyzing the user's emotional state, preferences, and dialogue tendencies using the generated dialogue content and acoustic characteristic information. This makes it possible to accurately grasp the user's emotional state and preferences, select appropriate compatible relationship candidates based on them, and provide advice in real time that corresponds to emotional changes that occur during the conversation.

[0479] "Voice data" refers to information recorded in digital format from the voice signals spoken by the user.

[0480] "Acoustic feature information" refers to information that includes features such as spectrum and frequency extracted from audio data.

[0481] "Transcripted dialogue content" refers to linguistic information in text format converted from audio data.

[0482] A "generative AI model" refers to a technology that uses artificial intelligence algorithms learned from vast datasets to perform analysis and predictions that utilize the characteristics of emotions and preferences.

[0483] "Compatible social relationship candidates" refer to other users who are considered to have a high affinity with the user, selected based on the user's emotional state and preferences.

[0484] A "fit score" refers to a numerical value used to quantitatively evaluate the degree of compatibility between users.

[0485] "Monitoring" refers to the process of observing and recording user conversations in real time.

[0486] "Means of generating advice" refers to the ability to create suggestions to support user communication based on analysis results.

[0487] This invention relates to a system that precisely analyzes a user's emotions and preferences and recommends a compatible partner. This system primarily involves acquiring and analyzing the user's voice data, and performing a series of processes to support dialogue. Specific embodiments are described below.

[0488] Data acquisition and analysis

[0489] First, the user uses an application installed on their mobile device to input voice into the system. During this process, the user can freely express their thoughts and wishes.

[0490] The device instantly captures voice data from the user and transfers it to a cloud server using wireless communication. This transfer utilizes an internet connection.

[0491] Subsequently, the server converts the audio data into text data using automatic speech recognition software (e.g., an existing API for speech-to-text conversion). Based on the resulting text data and acoustic feature information, the server uses a generative AI model to analyze the user's emotional state, preferences, and dialogue tendencies.

[0492] Partner selection and provision

[0493] Based on the analysis results, the server selects candidates for highly compatible social relationships. Specifically, it matches emotional and preference data using prompt messages to find compatible candidates.

[0494] Information on selected candidates is presented visually to the user via their device. This allows the user to easily review the details of the candidates and choose those they are interested in.

[0495] Specific example

[0496] For example, if a user says, "I'm tired from work today, but I want to relax this weekend," the emotion engine will analyze the fatigue and desire for relaxation. Example prompt: "Recommend other users with relevant hobbies or activities when the user is seeking relaxation."

[0497] This embodiment enables sophisticated matching that takes into account the emotional characteristics of users, and also supports the building of good relationships by assisting with communication after the interaction.

[0498] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0499] Step 1:

[0500] The user launches the app on their mobile device and inputs voice data. What the user speaks is recorded in real time by the mobile device. Input: User's voice. Output: Recorded voice data.

[0501] Step 2:

[0502] The device acquires the recorded audio data and sends it to a cloud server via an internet connection. The audio data is converted to a compressed digital file format on the device. Input: Recorded audio data. Output: Audio data file sent to the server.

[0503] Step 3:

[0504] The server converts received audio data into text data using speech recognition software. In addition, it analyzes acoustic features to extract voiceprint information. Data processing includes analysis of audio waveform data and digital processing of the audio signal. Input: Audio data file sent to the server. Output: Transcripted dialogue content and voiceprint information.

[0505] Step 4:

[0506] The server utilizes a generative AI model to evaluate the user's emotional state, preferences, and communication style from transcribed dialogue content and voiceprint information. Natural language processing algorithms are used for data calculation to determine emotional tendencies. Input: Transcribed dialogue content, voiceprint information. Output: Analyzed emotional state, preferences, and communication style.

[0507] Step 5:

[0508] The server searches a database of past data based on the prompt message and selects suitable social relationship candidates. The server calculates a fit score and performs analysis to identify candidates. Input: Analyzed emotional state, preferences, and prompt message. Output: Suitable relationship candidates and fit score.

[0509] Step 6:

[0510] The terminal presents the user with information on potential relationships received from the server. It displays compatibility scores and candidate profiles in a visually easy-to-understand format. Input: Compatible relationship candidates, suitability scores. Output: Candidate information presented to the user.

[0511] Step 7:

[0512] The server monitors the content of the conversation between users after matching in real time and generates advice to facilitate communication using a generated AI model. It supports users by sending the advice to their devices in a timely manner. Input: Content of the conversation between users. Output: Generated communication advice.

[0513] (Application Example 2)

[0514] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0515] In today's world, a vast array of content exists, and there is a need to efficiently select and provide content that is suitable for each individual user. However, recommending appropriate content based on users' emotions and preferences is difficult, and users may not get the experience they desire.

[0516] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0517] In this invention, the server includes means for acquiring voice information from the user, means for analyzing the voice information to generate the user's voiceprint information and transcribed dialogue information, means for analyzing the user's emotions, preferences, and communication style from the generated dialogue information and voiceprint information, means for selecting appropriate information according to the emotional state, and means for providing the information to the user. This makes it possible to recommend optimal content based on the user's emotions.

[0518] "Audio information" refers to sound data obtained from the user, including information about speech and voice characteristics.

[0519] "Voiceprint information" refers to data that shows the unique vocal characteristics of an individual, extracted from audio information.

[0520] "Textualized dialogue information" refers to the content of the dialogue in string format, which is obtained by analyzing and converting the user's voice.

[0521] "Emotions" refer to the psychological state analyzed from the user's voice and dialogue content.

[0522] "Preferences" refer to information that indicates the subjects or things that a user is interested in or concerned with.

[0523] "Communication style" refers to information about the methods and styles of communication used by users.

[0524] "Emotional state" refers to a user's temporary psychological state, including the type and intensity of their emotions.

[0525] "Means of selecting information" refers to a method of selecting data appropriate for the user based on their analyzed emotional state.

[0526] "Means of providing information" refers to the methods and interfaces used to present selected data to users.

[0527] As an application example of this invention, a system is realized that recommends content based on the user's emotional state. The system uses a smartphone or smart glasses. First, the user inputs voice information into the smart device. This voice information is sent from the device to a server via the internet. The server uses speech recognition software (e.g., Google Cloud Speech-to-Text API) to convert the voice into textual dialogue information. Furthermore, the server uses sentiment analysis software (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state from the text.

[0528] Based on the identified emotional state, the server selects content appropriate for the user. This selection uses a relevance score based on past user data and data from similar users. Once content selection is complete, the server sends the selection results to the terminal, and the user receives the content. The received data is then displayed visually, allowing the user to experience it.

[0529] As a concrete example, suppose a user voice-inputs "I'm tired today." In this case, the system determines that the emotional state is "fatigue" and recommends relaxing music or stress-relieving videos to alleviate that state.

[0530] An example of a prompt is, "Please recommend music that can help me relax when I feel stressed." By inputting this prompt into a generating AI model, the system can quickly provide content that is suitable for the user.

[0531] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0532] Step 1:

[0533] The user inputs voice information using a smart device. This voice information is collected by the device and transmitted directly to the server. The input is voice data, and the output from the device to the server is also voice data.

[0534] Step 2:

[0535] The server uses speech recognition software to convert the received audio information into text. During this process, data processing is performed to convert the audio data into string data. The output is the textualized dialogue information.

[0536] Step 3:

[0537] The server processes the transcribed dialogue information using sentiment analysis software to extract the user's emotional state. The data calculation performed here involves assigning emotion labels through text analysis. The input is text information, and the output is emotional state data.

[0538] Step 4:

[0539] The server selects appropriate content from its content database based on the user's emotional state. The selection process uses accumulated data to calculate a relevance score. The input is emotional state data, and the output is information about the recommended content.

[0540] Step 5:

[0541] The server transmits selected content information to the terminal, which then visualizes and presents that information to the user. Here, visualization is used to communicate the data to the user. The input is content information, and the output is display data for the user.

[0542] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0543] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0544] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0545] [Third Embodiment]

[0546] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0547] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0548] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0549] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0550] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0551] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0552] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0553] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0554] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0555] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0556] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0557] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0558] This invention is a next-generation matchmaking system that uses the user's voice data to recommend compatible partners. This system acquires the user's voice data and, through analysis, derives their emotions, preferences, and communication style, thereby suggesting a partner who will help them maintain a happy life after marriage.

[0559] Program processing

[0560] 1. Data Acquisition and Analysis

[0561] The user uses the app to perform voice input. Through a series of conversations, the app elicits the user's thoughts and intentions.

[0562] The device sends the user's voice data to the server. The voice is recorded in real time and transferred to the server via streaming technology.

[0563] The server uses speech recognition technology to convert the voice data into text. Based on this text, a generating AI analyzes the user's voiceprint information to determine their emotions, preferences, and communication style.

[0564] 2. Compatibility assessment and partner selection

[0565] The server uses these analysis results to enrich user profiles. Furthermore, it compares them with information on other users stored in the database to select suitable partner candidates.

[0566] The server calculates a compatibility score and lists the candidates with the highest scores.

[0567] 3. Providing suggestions and advice

[0568] The server proposes selected partner candidates to the user via the terminal.

[0569] Users can review the candidates and decide whether or not to interact with the suggested individuals.

[0570] The server continuously monitors user interaction patterns and feedback, and generates advice for improvement, thereby supporting the deepening of relationships.

[0571] Specific example

[0572] Example 1: Proposal for a potential partner

[0573] Input an audio recording of User A happily discussing their vacation plans.

[0574] The server analyzes user A's voiceprint and text to determine their characteristics, such as their love of travel and enjoyment of conversation.

[0575] The server identifies user B, who has similar hobbies and communication styles, as a candidate and suggests it to user A via the terminal.

[0576] User A is interested in the candidate, and a match is made.

[0577] Example 2: Communication support

[0578] The server analyzes the conversation between user A and user B after matching.

[0579] Based on the conversation, the server sends advice to user A's terminal recommending that they "try to get the other person to talk a little more."

[0580] User A puts the advice into practice, and communication proceeds smoothly.

[0581] In this way, this system analyzes the user's psychological characteristics and recommends an ideal partner. Furthermore, it supports the creation of happy relationships by improving continuous communication.

[0582] The following describes the processing flow.

[0583] Step 1:

[0584] Users register through the application. They enter basic information such as their name, age, gender, and hobbies, and then press the submit button.

[0585] Step 2:

[0586] The terminal formats the input data from the user and sends it to the server via secure communication.

[0587] Step 3:

[0588] The server stores the received user information in a database. This registers the user's basic profile within the system.

[0589] Step 4:

[0590] The user activates the device's voice input function to begin an initial dialogue session with the AI.

[0591] Step 5:

[0592] The device records the user's voice in real time and sends the audio data to the server in streaming format.

[0593] Step 6:

[0594] The server processes the received audio data through a speech recognition engine and converts it into text data.

[0595] Step 7:

[0596] The server uses a generation AI to analyze the transcribed dialogue content and voiceprint information to understand the user's emotions, preferences, and communication style.

[0597] Step 8:

[0598] Based on the analysis results, the server updates the user's profile and stores in-depth information in the database.

[0599] Step 9:

[0600] The server runs an algorithm that compares the database with other users' databases to select compatible partner candidates.

[0601] Step 10:

[0602] The server calculates a compatibility score, generates a list of potential partners with high scores, and sends it to the terminal.

[0603] Step 11:

[0604] The device displays proposed partner candidates to the user and provides an interface that allows for visual confirmation.

[0605] Step 12:

[0606] Users review the list of candidates and accept or reconsider matching with those they are interested in.

[0607] Step 13:

[0608] The server continuously monitors the conversations of matched pairs and generates advice to improve the quality of their communication.

[0609] Step 14:

[0610] The device periodically notifies users with advice to help them achieve better communication.

[0611] (Example 1)

[0612] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0613] In modern society, there is a growing need for more efficient relationship building between individuals. In particular, selecting the optimal partner while understanding each person's personality and preferences is not easy. Against this backdrop, there is a need to develop a system that utilizes voice information to provide more personalized suggestions.

[0614] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0615] In this invention, the server includes a device for acquiring voice information, a device for analyzing the voice information to generate voice characteristic information and written communication content, and a device for analyzing feelings, preferences, and communication styles from the generated communication content and voice characteristic information. This enables sophisticated selection of potential partners based on the individual characteristics of the user.

[0616] "Audio information" refers to sound data acquired for the purpose of recording and analyzing the user's speech.

[0617] "Voice feature information" refers to data that indicates the characteristics of a voice, obtained by analyzing voice information.

[0618] "Transcripted communication content" refers to the content of a conversation that has been generated from audio information and expressed in written form.

[0619] "Feelings" refers to the emotional state or sensations of the user.

[0620] "Preferences" is a concept that refers to the user's likes and interests.

[0621] "Communication style" refers to the methods and styles in which users communicate with others.

[0622] "Potential partners" refers to other users selected by the system who may be a good match for the user.

[0623] A "fitness index" is a numerical indicator that quantifies the degree of compatibility between a user's characteristics and the system.

[0624] "Advice" refers to suggestions and guidance provided to improve the relationship between the user and the potential partner.

[0625] This system is configured to suggest the most suitable partner based on the user's voice information and to support the continuous improvement of the relationship. Specifically, the invention will be implemented in the following form.

[0626] Users install a dedicated application on their device and use the voice input function to naturally express their thoughts and feelings. The voice information is recorded in real time by the device and transmitted to a server using streaming technology. This process incorporates advanced security and privacy protection features.

[0627] The server uses speech recognition software, such as the Google Speech-to-Text API, to convert the spoken information into text. Next, a generative AI model is used to analyze the transcribed conversation content and speech feature information. This analysis identifies the user's feelings, preferences, and communication style. The generative AI model extracts these characteristics by utilizing natural language processing techniques and sentiment analysis algorithms.

[0628] The analysis results are stored in a database on the server and used to enrich user profiles. These results are then compared with data from other users to calculate a suitability index. Based on this index, the server selects the most suitable candidates and sends the candidate list to the user's terminal.

[0629] Users can review candidate information displayed on their device and select those they are interested in. After selection, the server continuously monitors the user's interactions and provides advice using AI-generated content. This advice is designed to improve the quality of conversations and help deepen relationships.

[0630] For example, if a user inputs a voice message "talking happily about travel," the server will extract characteristics such as a love of travel and enjoyment of conversation, and suggest partners with similar interests. An example of a prompt message is, "Analyze the user's voice data and suggest the best partner based on emotions and preferences. Provide advice based on that and teach me how to improve communication." This system promotes interaction with others and supports the building of richer relationships.

[0631] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0632] Step 1:

[0633] The user launches a dedicated app and inputs voice information into the device. The voice information consists of words expressing the user's thoughts and feelings. The device records this in real time and sends the voice data to the server. In this process, the input is voice data, and the output is the transfer of voice data to the server.

[0634] Step 2:

[0635] The server analyzes the received audio data. First, it uses speech recognition software to convert the audio data into text data. This conversion yields audio feature information and the transcribed content of the conversation. Here, the input is audio data, and the output is text data.

[0636] Step 3:

[0637] The server uses a generative AI model to analyze text data. This analysis employs natural language processing techniques to extract user feelings, preferences, and communication styles from the text. The input is text data, and the output is the analyzed user profile information.

[0638] Step 4:

[0639] The server updates the user profile based on the analysis results. It then compares this profile with data from other users and calculates a compatibility index. This involves database searches and statistical analysis techniques. The input is the analyzed user profile information, and the output is the compatibility score.

[0640] Step 5:

[0641] The server selects the most suitable partner candidate based on the compatibility score. It then lists the information of the selected candidates and sends it to the user's terminal. The input is the compatibility score, and the output is the candidate list.

[0642] Step 6:

[0643] The user views a list of suggested candidates on their device and selects those they are interested in. The selections are fed back to the server. The input here is the candidate list, and the output is the feedback from the selected individuals.

[0644] Step 7:

[0645] After selection, the server continuously monitors user communication content. It utilizes generative AI to analyze the communication content and generate specific advice. This provides advice to improve the quality of user interactions. Inputs are feedback and communication content, while output is advice for communication improvement.

[0646] (Application Example 1)

[0647] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0648] In recent years, with the advancement of digital communication technology, there has been a growing demand from users for real-time, personalized experiences. However, existing systems struggle to analyze users' emotions and interests in real time and provide immediate, appropriate suggestions. As a result, users face the challenge of a reduced quality of experience due to unnatural pauses in communication and inappropriate suggestions.

[0649] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0650] In this invention, the server includes means for collecting voice data from the user in real time, means for analyzing the user's emotions and interests from the voice data in real time, and means for making appropriate suggestions according to the situation. As a result, the user can receive situation-appropriate suggestions in real time and enjoy a more personalized and high-quality communication experience.

[0651] A "user" is an individual or group that utilizes a particular service or system.

[0652] "Audio data" refers to information recorded in digital format from the voice spoken by a user.

[0653] "Means of real-time data collection" refers to technologies or methods that enable the immediate acquisition of information provided by users.

[0654] "Voiceprint information" refers to biometric authentication-related information obtained by analyzing the characteristics of a user's voice.

[0655] "Transcripted dialogue content" refers to information in text format extracted from audio data.

[0656] "Emotions" refer to information that indicates the user's psychological state.

[0657] "Preferences" refer to information about a user's likes and interests regarding a particular subject.

[0658] "Communication style" refers to the unique behaviors and patterns that users exhibit when exchanging information with others.

[0659] A "compatible partner" is another individual or digital agent that possesses characteristics that match the user's requirements and features.

[0660] "A means of presenting candidates to users in real time and enabling them to access dialogue" refers to a technology that provides users with immediately visible information on selected candidates, allowing them to interact with them directly.

[0661] "A means of analyzing a user's emotions and interests in real time from the user's voice data during a conversation" refers to a technology that instantly interprets voice information obtained during interaction and identifies the user's emotional and interest tendencies.

[0662] "Means of making appropriate suggestions according to the situation" refers to technologies or methods that present beneficial options or actions according to the environment and the user's state.

[0663] The system of the present invention acquires voice data from a smart device worn by the user or a terminal used by the user, analyzes it on a server, and provides the user with appropriate suggestions in real time.

[0664] The server uses streaming technology to acquire user voice data and converts it to text using a speech recognition API. Specific hardware includes smart glasses and ear-mounted microphones, which collect voice with high accuracy. The software can utilize commonly used speech recognition APIs such as Google Cloud Speech-to-Text and Azure Cognitive Services for sentiment analysis. This allows for real-time analysis of user emotions, preferences, and communication styles from their voice, enabling appropriate suggestions tailored to the user's behavior in the digital environment.

[0665] When a user initiates a conversation in the digital space, the server collects the voice data and analyzes their emotions and interests. This allows it to make real-time suggestions to help them achieve their desired experience. These suggestions are continuously improved based on user feedback, and it's even possible to suggest compatible digital partners by comparing them with past data.

[0666] For example, if a user is talking about action movies during a virtual date, the server can analyze the audio and recommend new action movies that might interest them. The generative AI model used in this process includes the ability to analyze the flow of the conversation and generate appropriate options, and examples of prompts include the following:

[0667] "A user is talking about action movies during a virtual date. Analyze the user's comments and figure out how to recommend new action movies that might interest them."

[0668] Thus, the present invention enables high-quality, real-time interaction by proposing personalized digital experiences tailored to the user.

[0669] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0670] Step 1:

[0671] The user inputs voice data via a smart device. The device collects this voice data in real time and sends it to a server using streaming technology. In this process, the voice signal is converted into digital data and transmitted instantly.

[0672] Step 2:

[0673] The server inputs the received audio data into a speech recognition API, which converts the audio into text. This speech recognition API quickly converts the audio data into a string of characters. The output is the user's spoken content formatted as text.

[0674] Step 3:

[0675] The server uses a generative AI model to analyze the user's emotions, preferences, and communication style from the transcribed data. The input is the text data obtained in step 2, and the output is the result of analyzing its emotional and semantic nuances. This analysis clearly reveals the user's psychological state and interests.

[0676] Step 4:

[0677] Based on the analysis results, the server selects compatible digital content and partners from the available database and generates specific content. The input is the analysis results from step 3, and the output is content and action plans to be presented to the user. The generative AI model is also used here to derive options based on the prompt text.

[0678] Step 5:

[0679] The server instantly delivers selected content and information about the other party to the user's smart device and presents it as a user interface. Options for the next action or experience are displayed on the device, and the user uses these to proceed with the interaction.

[0680] Step 6:

[0681] When a user selects or reacts to presented content, the device sends this feedback back to the server. Using the returned feedback data, the server further learns the user's preferences and continuously improves its analysis algorithms.

[0682] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0683] This invention is a next-generation matchmaking system that analyzes users' emotions in detail and recommends compatible partners. This system analyzes the user's voiceprint information and transcribed dialogue content, and uses an emotion engine to achieve more sophisticated matching.

[0684] Program processing

[0685] 1. Data acquisition and voice analysis

[0686] The user launches the application and inputs voice commands into the device. They then engage in real-time dialogue, providing the system with their thoughts and wishes.

[0687] The terminal immediately transfers the acquired audio data to the server.

[0688] The server converts the audio to text and extracts voiceprint information. The generated text is then subjected to analysis by an emotion engine.

[0689] 2. Sentiment analysis and profile updates

[0690] The server uses an emotion engine to identify the user's emotions from the text. It recognizes the specific emotional state and reflects it in the profile.

[0691] We collect user sentiment data and compare it with past history to understand long-term trends.

[0692] 3. Partner Selection

[0693] Based on the analysis results, the server selects compatible partner candidates from the database, taking into account emotional characteristics and communication style compatibility.

[0694] We utilize past data to calculate compatibility scores and narrow down the candidates.

[0695] 4. Information provision and two-way advice

[0696] The device presents the user with selected partner candidates. Information is provided in a visually easy-to-understand format.

[0697] The server monitors user interactions and generates advice based on real-time sentiment analysis. It also sends timely notifications to the device to support effective communication.

[0698] Specific example

[0699] Example 1: Matching based on sentiment analysis

[0700] User A talks about stress through the app. The server analyzes User A's anxiety levels from their voiceprint and words. It also identifies that User A likes relaxation.

[0701] The server selects User B, who has resourceful hobbies, as a suitable partner for this profile, suggesting a high level of compatibility.

[0702] The device provides this information to user A, and user A becomes interested in user B.

[0703] Example 2: Providing real-time advice

[0704] User A and User B begin a conversation after being matched. The server monitors the conversation.

[0705] During a conversation, the emotion engine detects changes in user A's emotions (e.g., tension) and suggests topics to help them relax.

[0706] The device informs user A of this and supports the smooth progress of the conversation.

[0707] This system allows for a detailed analysis of the user's psychological characteristics, recommends an ideal partner, and continuously improves the relationship.

[0708] The following describes the processing flow.

[0709] Step 1:

[0710] The user opens the application and inputs voice commands into the device. The user can then introduce themselves or talk about topics that interest them.

[0711] Step 2:

[0712] The terminal receives voice input and captures it as audio data. At the same time, it prepares this audio data for transmission to the server.

[0713] Step 3:

[0714] The server receives the voice data sent from the terminal and uses a speech recognition engine to convert the data into text.

[0715] Step 4:

[0716] The server inputs text data into the emotion engine, which then analyzes the user's emotional state. This process identifies specific emotions such as joy, sadness, and anger.

[0717] Step 5:

[0718] The server analyzes voiceprint information to extract the unique characteristics of the user's voice. This allows for a more detailed understanding of the user's characteristics.

[0719] Step 6:

[0720] The server updates the user's profile using acquired sentiment data and voiceprint information. This includes emotional tendencies and preferences.

[0721] Step 7:

[0722] The server performs a database search based on profile data to select compatible partner candidates. The selection focuses on emotional compatibility and matching communication styles.

[0723] Step 8:

[0724] The server evaluates the selected partner candidates and calculates a compatibility score. This quantifies the compatibility between users.

[0725] Step 9:

[0726] The device receives information about potential partners along with compatibility scores from the server and provides it to the user. This information is displayed in a visually and intuitively easy-to-understand format.

[0727] Step 10:

[0728] Users review the information provided about potential partners and make a decision about matching. This is expressed as a contact request to the preferred candidate.

[0729] Step 11:

[0730] If a user match is successful, the server will continue to monitor the interaction between both parties in real time.

[0731] Step 12:

[0732] The server utilizes an emotion engine to track changes in emotions during conversations and generates communication advice as needed.

[0733] Step 13:

[0734] The device notifies the user of the generated advice, encouraging positive communication.

[0735] (Example 2)

[0736] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0737] In today's society, accurately understanding one's own emotional state and preferences, and finding a suitable partner based on that understanding, is crucial for building deeper relationships. However, conventional matching systems struggle to accurately analyze users' emotional states, limiting their ability to find compatible partner candidates. Furthermore, there is a lack of systems that provide appropriate advice to support communication after matching. Therefore, there is a need for a system that can more accurately and effectively select compatible partner candidates and support communication.

[0738] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0739] In this invention, the server includes means for acquiring voice data from a user, means for analyzing the voice data to generate the user's acoustic characteristic information and transcribed dialogue content, and means for analyzing the user's emotional state, preferences, and dialogue tendencies using the generated dialogue content and acoustic characteristic information. This makes it possible to accurately grasp the user's emotional state and preferences, select appropriate compatible relationship candidates based on them, and provide advice in real time that corresponds to emotional changes that occur during the conversation.

[0740] "Voice data" refers to information recorded in digital format from the voice signals spoken by the user.

[0741] "Acoustic feature information" refers to information that includes features such as spectrum and frequency extracted from audio data.

[0742] "Transcripted dialogue content" refers to linguistic information in text format converted from audio data.

[0743] A "generative AI model" refers to a technology that uses artificial intelligence algorithms learned from vast datasets to perform analysis and predictions that utilize the characteristics of emotions and preferences.

[0744] "Compatible social relationship candidates" refer to other users who are considered to have a high affinity with the user, selected based on the user's emotional state and preferences.

[0745] A "fit score" refers to a numerical value used to quantitatively evaluate the degree of compatibility between users.

[0746] "Monitoring" refers to the process of observing and recording user conversations in real time.

[0747] "Means of generating advice" refers to the ability to create suggestions to support user communication based on analysis results.

[0748] This invention relates to a system that precisely analyzes a user's emotions and preferences and recommends a compatible partner. This system primarily involves acquiring and analyzing the user's voice data, and performing a series of processes to support dialogue. Specific embodiments are described below.

[0749] Data acquisition and analysis

[0750] First, the user uses an application installed on their mobile device to input voice into the system. During this process, the user can freely express their thoughts and wishes.

[0751] The device instantly captures voice data from the user and transfers it to a cloud server using wireless communication. This transfer utilizes an internet connection.

[0752] Subsequently, the server converts the audio data into text data using automatic speech recognition software (e.g., an existing API for speech-to-text conversion). Based on the resulting text data and acoustic feature information, the server uses a generative AI model to analyze the user's emotional state, preferences, and dialogue tendencies.

[0753] Partner selection and provision

[0754] Based on the analysis results, the server selects candidates for highly compatible social relationships. Specifically, it matches emotional and preference data using prompt messages to find compatible candidates.

[0755] Information on selected candidates is presented visually to the user via their device. This allows the user to easily review the details of the candidates and choose those they are interested in.

[0756] Specific example

[0757] For example, if a user says, "I'm tired from work today, but I want to relax this weekend," the emotion engine will analyze the fatigue and desire for relaxation. Example prompt: "Recommend other users with relevant hobbies or activities when the user is seeking relaxation."

[0758] This embodiment enables sophisticated matching that takes into account the emotional characteristics of users, and also supports the building of good relationships by assisting with communication after the interaction.

[0759] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0760] Step 1:

[0761] The user launches the app on their mobile device and inputs voice data. What the user speaks is recorded in real time by the mobile device. Input: User's voice. Output: Recorded voice data.

[0762] Step 2:

[0763] The device acquires the recorded audio data and sends it to a cloud server via an internet connection. The audio data is converted to a compressed digital file format on the device. Input: Recorded audio data. Output: Audio data file sent to the server.

[0764] Step 3:

[0765] The server converts received audio data into text data using speech recognition software. In addition, it analyzes acoustic features to extract voiceprint information. Data processing includes analysis of audio waveform data and digital processing of the audio signal. Input: Audio data file sent to the server. Output: Transcripted dialogue content and voiceprint information.

[0766] Step 4:

[0767] The server utilizes a generative AI model to evaluate the user's emotional state, preferences, and communication style from transcribed dialogue content and voiceprint information. Natural language processing algorithms are used for data calculation to determine emotional tendencies. Input: Transcribed dialogue content, voiceprint information. Output: Analyzed emotional state, preferences, and communication style.

[0768] Step 5:

[0769] The server searches a database of past data based on the prompt message and selects suitable social relationship candidates. The server calculates a fit score and performs analysis to identify candidates. Input: Analyzed emotional state, preferences, and prompt message. Output: Suitable relationship candidates and fit score.

[0770] Step 6:

[0771] The terminal presents the user with information on potential relationships received from the server. It displays compatibility scores and candidate profiles in a visually easy-to-understand format. Input: Compatible relationship candidates, suitability scores. Output: Candidate information presented to the user.

[0772] Step 7:

[0773] The server monitors the content of the conversation between users after matching in real time and generates advice to facilitate communication using a generated AI model. It supports users by sending the advice to their devices in a timely manner. Input: Content of the conversation between users. Output: Generated communication advice.

[0774] (Application Example 2)

[0775] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0776] In today's world, a vast array of content exists, and there is a need to efficiently select and provide content that is suitable for each individual user. However, recommending appropriate content based on users' emotions and preferences is difficult, and users may not get the experience they desire.

[0777] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0778] In this invention, the server includes means for acquiring voice information from the user, means for analyzing the voice information to generate the user's voiceprint information and transcribed dialogue information, means for analyzing the user's emotions, preferences, and communication style from the generated dialogue information and voiceprint information, means for selecting appropriate information according to the emotional state, and means for providing the information to the user. This makes it possible to recommend optimal content based on the user's emotions.

[0779] "Audio information" refers to sound data obtained from the user, including information about speech and voice characteristics.

[0780] "Voiceprint information" refers to data that shows the unique vocal characteristics of an individual, extracted from audio information.

[0781] "Textualized dialogue information" refers to the content of the dialogue in string format, which is obtained by analyzing and converting the user's voice.

[0782] "Emotions" refer to the psychological state analyzed from the user's voice and dialogue content.

[0783] "Preferences" refer to information that indicates the subjects or things that a user is interested in or concerned with.

[0784] "Communication style" refers to information about the methods and styles of communication used by users.

[0785] "Emotional state" refers to a user's temporary psychological state, including the type and intensity of their emotions.

[0786] "Means of selecting information" refers to a method of selecting data appropriate for the user based on their analyzed emotional state.

[0787] "Means of providing information" refers to the methods and interfaces used to present selected data to users.

[0788] As an application example of this invention, a system is realized that recommends content based on the user's emotional state. The system uses a smartphone or smart glasses. First, the user inputs voice information into the smart device. This voice information is sent from the device to a server via the internet. The server uses speech recognition software (e.g., Google Cloud Speech-to-Text API) to convert the voice into textual dialogue information. Furthermore, the server uses sentiment analysis software (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state from the text.

[0789] Based on the identified emotional state, the server selects content appropriate for the user. This selection uses a relevance score based on past user data and data from similar users. Once content selection is complete, the server sends the selection results to the terminal, and the user receives the content. The received data is then displayed visually, allowing the user to experience it.

[0790] As a concrete example, suppose a user voice-inputs "I'm tired today." In this case, the system determines that the emotional state is "fatigue" and recommends relaxing music or stress-relieving videos to alleviate that state.

[0791] An example of a prompt is, "Please recommend music that can help me relax when I feel stressed." By inputting this prompt into a generating AI model, the system can quickly provide content that is suitable for the user.

[0792] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0793] Step 1:

[0794] The user inputs voice information using a smart device. This voice information is collected by the device and transmitted directly to the server. The input is voice data, and the output from the device to the server is also voice data.

[0795] Step 2:

[0796] The server uses speech recognition software to convert the received audio information into text. During this process, data processing is performed to convert the audio data into string data. The output is the textualized dialogue information.

[0797] Step 3:

[0798] The server processes the transcribed dialogue information using sentiment analysis software to extract the user's emotional state. The data calculation performed here involves assigning emotion labels through text analysis. The input is text information, and the output is emotional state data.

[0799] Step 4:

[0800] The server selects appropriate content from its content database based on the user's emotional state. The selection process uses accumulated data to calculate a relevance score. The input is emotional state data, and the output is information about the recommended content.

[0801] Step 5:

[0802] The server transmits selected content information to the terminal, which then visualizes and presents that information to the user. Here, visualization is used to communicate the data to the user. The input is content information, and the output is display data for the user.

[0803] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0804] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0805] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0806] [Fourth Embodiment]

[0807] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0808] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0809] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0810] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0811] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0812] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0813] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0814] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0815] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0816] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0817] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0818] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0819] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0820] This invention is a next-generation matchmaking system that uses the user's voice data to recommend compatible partners. This system acquires the user's voice data and, through analysis, derives their emotions, preferences, and communication style, thereby suggesting a partner who will help them maintain a happy life after marriage.

[0821] Program processing

[0822] 1. Data Acquisition and Analysis

[0823] The user uses the app to perform voice input. Through a series of conversations, the app elicits the user's thoughts and intentions.

[0824] The device sends the user's voice data to the server. The voice is recorded in real time and transferred to the server via streaming technology.

[0825] The server uses speech recognition technology to convert the voice data into text. Based on this text, a generating AI analyzes the user's voiceprint information to determine their emotions, preferences, and communication style.

[0826] 2. Compatibility assessment and partner selection

[0827] The server uses these analysis results to enrich user profiles. Furthermore, it compares them with information on other users stored in the database to select suitable partner candidates.

[0828] The server calculates a compatibility score and lists the candidates with the highest scores.

[0829] 3. Providing suggestions and advice

[0830] The server proposes selected partner candidates to the user via the terminal.

[0831] Users can review the candidates and decide whether or not to interact with the suggested individuals.

[0832] The server continuously monitors user interaction patterns and feedback, and generates advice for improvement, thereby supporting the deepening of relationships.

[0833] Specific example

[0834] Example 1: Proposal for a potential partner

[0835] Input an audio recording of User A happily discussing their vacation plans.

[0836] The server analyzes user A's voiceprint and text to determine their characteristics, such as their love of travel and enjoyment of conversation.

[0837] The server identifies user B, who has similar hobbies and communication styles, as a candidate and suggests it to user A via the terminal.

[0838] User A is interested in the candidate, and a match is made.

[0839] Example 2: Communication support

[0840] The server analyzes the conversation between user A and user B after matching.

[0841] Based on the conversation, the server sends advice to user A's terminal recommending that they "try to get the other person to talk a little more."

[0842] User A puts the advice into practice, and communication proceeds smoothly.

[0843] In this way, this system analyzes the user's psychological characteristics and recommends an ideal partner. Furthermore, it supports the creation of happy relationships by improving continuous communication.

[0844] The following describes the processing flow.

[0845] Step 1:

[0846] Users register through the application. They enter basic information such as their name, age, gender, and hobbies, and then press the submit button.

[0847] Step 2:

[0848] The terminal formats the input data from the user and sends it to the server via secure communication.

[0849] Step 3:

[0850] The server stores the received user information in a database. This registers the user's basic profile within the system.

[0851] Step 4:

[0852] The user activates the device's voice input function to begin an initial dialogue session with the AI.

[0853] Step 5:

[0854] The device records the user's voice in real time and sends the audio data to the server in streaming format.

[0855] Step 6:

[0856] The server processes the received audio data through a speech recognition engine and converts it into text data.

[0857] Step 7:

[0858] The server uses a generation AI to analyze the transcribed dialogue content and voiceprint information to understand the user's emotions, preferences, and communication style.

[0859] Step 8:

[0860] Based on the analysis results, the server updates the user's profile and stores in-depth information in the database.

[0861] Step 9:

[0862] The server runs an algorithm that compares the database with other users' databases to select compatible partner candidates.

[0863] Step 10:

[0864] The server calculates a compatibility score, generates a list of potential partners with high scores, and sends it to the terminal.

[0865] Step 11:

[0866] The device displays proposed partner candidates to the user and provides an interface that allows for visual confirmation.

[0867] Step 12:

[0868] Users review the list of candidates and accept or reconsider matching with those they are interested in.

[0869] Step 13:

[0870] The server continuously monitors the conversations of matched pairs and generates advice to improve the quality of their communication.

[0871] Step 14:

[0872] The device periodically notifies users with advice to help them achieve better communication.

[0873] (Example 1)

[0874] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0875] In modern society, there is a growing need for more efficient relationship building between individuals. In particular, selecting the optimal partner while understanding each person's personality and preferences is not easy. Against this backdrop, there is a need to develop a system that utilizes voice information to provide more personalized suggestions.

[0876] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0877] In this invention, the server includes a device for acquiring voice information, a device for analyzing the voice information to generate voice characteristic information and written communication content, and a device for analyzing feelings, preferences, and communication styles from the generated communication content and voice characteristic information. This enables sophisticated selection of potential partners based on the individual characteristics of the user.

[0878] "Audio information" refers to sound data acquired for the purpose of recording and analyzing the user's speech.

[0879] "Voice feature information" refers to data that indicates the characteristics of a voice, obtained by analyzing voice information.

[0880] "Transcripted communication content" refers to the content of a conversation that has been generated from audio information and expressed in written form.

[0881] "Feelings" refers to the emotional state or sensations of the user.

[0882] "Preferences" is a concept that refers to the user's likes and interests.

[0883] "Communication style" refers to the methods and styles in which users communicate with others.

[0884] "Potential partners" refers to other users selected by the system who may be a good match for the user.

[0885] A "fitness index" is a numerical indicator that quantifies the degree of compatibility between a user's characteristics and the system.

[0886] "Advice" refers to suggestions and guidance provided to improve the relationship between the user and the potential partner.

[0887] This system is configured to suggest the most suitable partner based on the user's voice information and to support the continuous improvement of the relationship. Specifically, the invention will be implemented in the following form.

[0888] Users install a dedicated application on their device and use the voice input function to naturally express their thoughts and feelings. The voice information is recorded in real time by the device and transmitted to a server using streaming technology. This process incorporates advanced security and privacy protection features.

[0889] The server uses speech recognition software, such as the Google Speech-to-Text API, to convert the spoken information into text. Next, a generative AI model is used to analyze the transcribed conversation content and speech feature information. This analysis identifies the user's feelings, preferences, and communication style. The generative AI model extracts these characteristics by utilizing natural language processing techniques and sentiment analysis algorithms.

[0890] The analysis results are stored in a database on the server and used to enrich user profiles. These results are then compared with data from other users to calculate a suitability index. Based on this index, the server selects the most suitable candidates and sends the candidate list to the user's terminal.

[0891] Users can review candidate information displayed on their device and select those they are interested in. After selection, the server continuously monitors the user's interactions and provides advice using AI-generated content. This advice is designed to improve the quality of conversations and help deepen relationships.

[0892] For example, if a user inputs a voice message "talking happily about travel," the server will extract characteristics such as a love of travel and enjoyment of conversation, and suggest partners with similar interests. An example of a prompt message is, "Analyze the user's voice data and suggest the best partner based on emotions and preferences. Provide advice based on that and teach me how to improve communication." This system promotes interaction with others and supports the building of richer relationships.

[0893] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0894] Step 1:

[0895] The user launches a dedicated app and inputs voice information into the device. The voice information consists of words expressing the user's thoughts and feelings. The device records this in real time and sends the voice data to the server. In this process, the input is voice data, and the output is the transfer of voice data to the server.

[0896] Step 2:

[0897] The server analyzes the received audio data. First, it uses speech recognition software to convert the audio data into text data. This conversion yields audio feature information and the transcribed content of the conversation. Here, the input is audio data, and the output is text data.

[0898] Step 3:

[0899] The server uses a generative AI model to analyze text data. This analysis employs natural language processing techniques to extract user feelings, preferences, and communication styles from the text. The input is text data, and the output is the analyzed user profile information.

[0900] Step 4:

[0901] The server updates the user profile based on the analysis results. It then compares this profile with data from other users and calculates a compatibility index. This involves database searches and statistical analysis techniques. The input is the analyzed user profile information, and the output is the compatibility score.

[0902] Step 5:

[0903] The server selects the most suitable partner candidate based on the compatibility score. It then lists the information of the selected candidates and sends it to the user's terminal. The input is the compatibility score, and the output is the candidate list.

[0904] Step 6:

[0905] The user views a list of suggested candidates on their device and selects those they are interested in. The selections are fed back to the server. The input here is the candidate list, and the output is the feedback from the selected individuals.

[0906] Step 7:

[0907] After selection, the server continuously monitors user communication content. It utilizes generative AI to analyze the communication content and generate specific advice. This provides advice to improve the quality of user interactions. Inputs are feedback and communication content, while output is advice for communication improvement.

[0908] (Application Example 1)

[0909] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0910] In recent years, with the advancement of digital communication technology, there has been a growing demand from users for real-time, personalized experiences. However, existing systems struggle to analyze users' emotions and interests in real time and provide immediate, appropriate suggestions. As a result, users face the challenge of a reduced quality of experience due to unnatural pauses in communication and inappropriate suggestions.

[0911] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0912] In this invention, the server includes means for collecting voice data from the user in real time, means for analyzing the user's emotions and interests from the voice data in real time, and means for making appropriate suggestions according to the situation. As a result, the user can receive situation-appropriate suggestions in real time and enjoy a more personalized and high-quality communication experience.

[0913] A "user" is an individual or group that utilizes a particular service or system.

[0914] "Audio data" refers to information recorded in digital format from the voice spoken by a user.

[0915] "Means of real-time data collection" refers to technologies or methods that enable the immediate acquisition of information provided by users.

[0916] "Voiceprint information" refers to biometric authentication-related information obtained by analyzing the characteristics of a user's voice.

[0917] "Transcripted dialogue content" refers to information in text format extracted from audio data.

[0918] "Emotions" refer to information that indicates the user's psychological state.

[0919] "Preferences" refer to information about a user's likes and interests regarding a particular subject.

[0920] "Communication style" refers to the unique behaviors and patterns that users exhibit when exchanging information with others.

[0921] A "compatible partner" is another individual or digital agent that possesses characteristics that match the user's requirements and features.

[0922] "A means of presenting candidates to users in real time and enabling them to access dialogue" refers to a technology that provides users with immediately visible information on selected candidates, allowing them to interact with them directly.

[0923] "A means of analyzing a user's emotions and interests in real time from the user's voice data during a conversation" refers to a technology that instantly interprets voice information obtained during interaction and identifies the user's emotional and interest tendencies.

[0924] "Means of making appropriate suggestions according to the situation" refers to technologies or methods that present beneficial options or actions according to the environment and the user's state.

[0925] The system of the present invention acquires voice data from a smart device worn by the user or a terminal used by the user, analyzes it on a server, and provides the user with appropriate suggestions in real time.

[0926] The server uses streaming technology to acquire user voice data and converts it to text using a speech recognition API. Specific hardware includes smart glasses and ear-mounted microphones, which collect voice with high accuracy. The software can utilize commonly used speech recognition APIs such as Google Cloud Speech-to-Text and Azure Cognitive Services for sentiment analysis. This allows for real-time analysis of user emotions, preferences, and communication styles from their voice, enabling appropriate suggestions tailored to the user's behavior in the digital environment.

[0927] When a user initiates a conversation in the digital space, the server collects the voice data and analyzes their emotions and interests. This allows it to make real-time suggestions to help them achieve their desired experience. These suggestions are continuously improved based on user feedback, and it's even possible to suggest compatible digital partners by comparing them with past data.

[0928] For example, if a user is talking about action movies during a virtual date, the server can analyze the audio and recommend new action movies that might interest them. The generative AI model used in this process includes the ability to analyze the flow of the conversation and generate appropriate options, and examples of prompts include the following:

[0929] "A user is talking about action movies during a virtual date. Analyze the user's comments and figure out how to recommend new action movies that might interest them."

[0930] Thus, the present invention enables high-quality, real-time interaction by proposing personalized digital experiences tailored to the user.

[0931] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0932] Step 1:

[0933] The user inputs voice data via a smart device. The device collects this voice data in real time and sends it to a server using streaming technology. In this process, the voice signal is converted into digital data and transmitted instantly.

[0934] Step 2:

[0935] The server inputs the received audio data into a speech recognition API, which converts the audio into text. This speech recognition API quickly converts the audio data into a string of characters. The output is the user's spoken content formatted as text.

[0936] Step 3:

[0937] The server uses a generative AI model to analyze the user's emotions, preferences, and communication style from the transcribed data. The input is the text data obtained in step 2, and the output is the result of analyzing its emotional and semantic nuances. This analysis clearly reveals the user's psychological state and interests.

[0938] Step 4:

[0939] Based on the analysis results, the server selects compatible digital content and partners from the available database and generates specific content. The input is the analysis results from step 3, and the output is content and action plans to be presented to the user. The generative AI model is also used here to derive options based on the prompt text.

[0940] Step 5:

[0941] The server instantly delivers selected content and information about the other party to the user's smart device and presents it as a user interface. Options for the next action or experience are displayed on the device, and the user uses these to proceed with the interaction.

[0942] Step 6:

[0943] When a user selects or reacts to presented content, the device sends this feedback back to the server. Using the returned feedback data, the server further learns the user's preferences and continuously improves its analysis algorithms.

[0944] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0945] This invention is a next-generation matchmaking system that analyzes users' emotions in detail and recommends compatible partners. This system analyzes the user's voiceprint information and transcribed dialogue content, and uses an emotion engine to achieve more sophisticated matching.

[0946] Program processing

[0947] 1. Data acquisition and voice analysis

[0948] The user launches the application and inputs voice commands into the device. They then engage in real-time dialogue, providing the system with their thoughts and wishes.

[0949] The terminal immediately transfers the acquired audio data to the server.

[0950] The server converts the audio to text and extracts voiceprint information. The generated text is then subjected to analysis by an emotion engine.

[0951] 2. Sentiment analysis and profile updates

[0952] The server uses an emotion engine to identify the user's emotions from the text. It recognizes the specific emotional state and reflects it in the profile.

[0953] We collect user sentiment data and compare it with past history to understand long-term trends.

[0954] 3. Partner Selection

[0955] Based on the analysis results, the server selects compatible partner candidates from the database, taking into account emotional characteristics and communication style compatibility.

[0956] We utilize past data to calculate compatibility scores and narrow down the candidates.

[0957] 4. Information provision and two-way advice

[0958] The device presents the user with selected partner candidates. Information is provided in a visually easy-to-understand format.

[0959] The server monitors user interactions and generates advice based on real-time sentiment analysis. It also sends timely notifications to the device to support effective communication.

[0960] Specific example

[0961] Example 1: Matching based on sentiment analysis

[0962] User A talks about stress through the app. The server analyzes User A's anxiety levels from their voiceprint and words. It also identifies that User A likes relaxation.

[0963] The server selects User B, who has resourceful hobbies, as a suitable partner for this profile, suggesting a high level of compatibility.

[0964] The device provides this information to user A, and user A becomes interested in user B.

[0965] Example 2: Providing real-time advice

[0966] User A and User B begin a conversation after being matched. The server monitors the conversation.

[0967] During a conversation, the emotion engine detects changes in user A's emotions (e.g., tension) and suggests topics to help them relax.

[0968] The device informs user A of this and supports the smooth progress of the conversation.

[0969] This system allows for a detailed analysis of the user's psychological characteristics, recommends an ideal partner, and continuously improves the relationship.

[0970] The following describes the processing flow.

[0971] Step 1:

[0972] The user opens the application and inputs voice commands into the device. The user can then introduce themselves or talk about topics that interest them.

[0973] Step 2:

[0974] The terminal receives voice input and captures it as audio data. At the same time, it prepares this audio data for transmission to the server.

[0975] Step 3:

[0976] The server receives the voice data sent from the terminal and uses a speech recognition engine to convert the data into text.

[0977] Step 4:

[0978] The server inputs text data into the emotion engine, which then analyzes the user's emotional state. This process identifies specific emotions such as joy, sadness, and anger.

[0979] Step 5:

[0980] The server analyzes voiceprint information to extract the unique characteristics of the user's voice. This allows for a more detailed understanding of the user's characteristics.

[0981] Step 6:

[0982] The server updates the user's profile using acquired sentiment data and voiceprint information. This includes emotional tendencies and preferences.

[0983] Step 7:

[0984] The server performs a database search based on profile data to select compatible partner candidates. The selection focuses on emotional compatibility and matching communication styles.

[0985] Step 8:

[0986] The server evaluates the selected partner candidates and calculates a compatibility score. This quantifies the compatibility between users.

[0987] Step 9:

[0988] The device receives information about potential partners along with compatibility scores from the server and provides it to the user. This information is displayed in a visually and intuitively easy-to-understand format.

[0989] Step 10:

[0990] Users review the information provided about potential partners and make a decision about matching. This is expressed as a contact request to the preferred candidate.

[0991] Step 11:

[0992] If a user match is successful, the server will continue to monitor the interaction between both parties in real time.

[0993] Step 12:

[0994] The server utilizes an emotion engine to track changes in emotions during conversations and generates communication advice as needed.

[0995] Step 13:

[0996] The device notifies the user of the generated advice, encouraging positive communication.

[0997] (Example 2)

[0998] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0999] In today's society, accurately understanding one's own emotional state and preferences, and finding a suitable partner based on that understanding, is crucial for building deeper relationships. However, conventional matching systems struggle to accurately analyze users' emotional states, limiting their ability to find compatible partner candidates. Furthermore, there is a lack of systems that provide appropriate advice to support communication after matching. Therefore, there is a need for a system that can more accurately and effectively select compatible partner candidates and support communication.

[1000] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1001] In this invention, the server includes means for acquiring voice data from a user, means for analyzing the voice data to generate the user's acoustic characteristic information and transcribed dialogue content, and means for analyzing the user's emotional state, preferences, and dialogue tendencies using the generated dialogue content and acoustic characteristic information. This makes it possible to accurately grasp the user's emotional state and preferences, select appropriate compatible relationship candidates based on them, and provide advice in real time that corresponds to emotional changes that occur during the conversation.

[1002] "Voice data" refers to information recorded in digital format from the voice signals spoken by the user.

[1003] "Acoustic feature information" refers to information that includes features such as spectrum and frequency extracted from audio data.

[1004] "Transcripted dialogue content" refers to linguistic information in text format converted from audio data.

[1005] A "generative AI model" refers to a technology that uses artificial intelligence algorithms learned from vast datasets to perform analysis and predictions that utilize the characteristics of emotions and preferences.

[1006] "Compatible social relationship candidates" refer to other users who are considered to have a high affinity with the user, selected based on the user's emotional state and preferences.

[1007] A "fit score" refers to a numerical value used to quantitatively evaluate the degree of compatibility between users.

[1008] "Monitoring" refers to the process of observing and recording user conversations in real time.

[1009] "Means of generating advice" refers to the ability to create suggestions to support user communication based on analysis results.

[1010] This invention relates to a system that precisely analyzes a user's emotions and preferences and recommends a compatible partner. This system primarily involves acquiring and analyzing the user's voice data, and performing a series of processes to support dialogue. Specific embodiments are described below.

[1011] Data acquisition and analysis

[1012] First, the user uses an application installed on their mobile device to input voice into the system. During this process, the user can freely express their thoughts and wishes.

[1013] The device instantly captures voice data from the user and transfers it to a cloud server using wireless communication. This transfer utilizes an internet connection.

[1014] Subsequently, the server converts the audio data into text data using automatic speech recognition software (e.g., an existing API for speech-to-text conversion). Based on the resulting text data and acoustic feature information, the server uses a generative AI model to analyze the user's emotional state, preferences, and dialogue tendencies.

[1015] Partner selection and provision

[1016] Based on the analysis results, the server selects candidates for highly compatible social relationships. Specifically, it matches emotional and preference data using prompt messages to find compatible candidates.

[1017] Information on selected candidates is presented visually to the user via their device. This allows the user to easily review the details of the candidates and choose those they are interested in.

[1018] Specific example

[1019] For example, if a user says, "I'm tired from work today, but I want to relax this weekend," the emotion engine will analyze the fatigue and desire for relaxation. Example prompt: "Recommend other users with relevant hobbies or activities when the user is seeking relaxation."

[1020] This embodiment enables sophisticated matching that takes into account the emotional characteristics of users, and also supports the building of good relationships by assisting with communication after the interaction.

[1021] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1022] Step 1:

[1023] The user launches the app on their mobile device and inputs voice data. What the user speaks is recorded in real time by the mobile device. Input: User's voice. Output: Recorded voice data.

[1024] Step 2:

[1025] The device acquires the recorded audio data and sends it to a cloud server via an internet connection. The audio data is converted to a compressed digital file format on the device. Input: Recorded audio data. Output: Audio data file sent to the server.

[1026] Step 3:

[1027] The server converts received audio data into text data using speech recognition software. In addition, it analyzes acoustic features to extract voiceprint information. Data processing includes analysis of audio waveform data and digital processing of the audio signal. Input: Audio data file sent to the server. Output: Transcripted dialogue content and voiceprint information.

[1028] Step 4:

[1029] The server utilizes a generative AI model to evaluate the user's emotional state, preferences, and communication style from transcribed dialogue content and voiceprint information. Natural language processing algorithms are used for data calculation to determine emotional tendencies. Input: Transcribed dialogue content, voiceprint information. Output: Analyzed emotional state, preferences, and communication style.

[1030] Step 5:

[1031] The server searches a database of past data based on the prompt message and selects suitable social relationship candidates. The server calculates a fit score and performs analysis to identify candidates. Input: Analyzed emotional state, preferences, and prompt message. Output: Suitable relationship candidates and fit score.

[1032] Step 6:

[1033] The terminal presents the user with information on potential relationships received from the server. It displays compatibility scores and candidate profiles in a visually easy-to-understand format. Input: Compatible relationship candidates, suitability scores. Output: Candidate information presented to the user.

[1034] Step 7:

[1035] The server monitors the content of the conversation between users after matching in real time and generates advice to facilitate communication using a generated AI model. It supports users by sending the advice to their devices in a timely manner. Input: Content of the conversation between users. Output: Generated communication advice.

[1036] (Application Example 2)

[1037] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1038] In today's world, a vast array of content exists, and there is a need to efficiently select and provide content that is suitable for each individual user. However, recommending appropriate content based on users' emotions and preferences is difficult, and users may not get the experience they desire.

[1039] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1040] In this invention, the server includes means for acquiring voice information from the user, means for analyzing the voice information to generate the user's voiceprint information and transcribed dialogue information, means for analyzing the user's emotions, preferences, and communication style from the generated dialogue information and voiceprint information, means for selecting appropriate information according to the emotional state, and means for providing the information to the user. This makes it possible to recommend optimal content based on the user's emotions.

[1041] "Audio information" refers to sound data obtained from the user, including information about speech and voice characteristics.

[1042] "Voiceprint information" refers to data that shows the unique vocal characteristics of an individual, extracted from audio information.

[1043] "Textualized dialogue information" refers to the content of the dialogue in string format, which is obtained by analyzing and converting the user's voice.

[1044] "Emotions" refer to the psychological state analyzed from the user's voice and dialogue content.

[1045] "Preferences" refer to information that indicates the subjects or things that a user is interested in or concerned with.

[1046] "Communication style" refers to information about the methods and styles of communication used by users.

[1047] "Emotional state" refers to a user's temporary psychological state, including the type and intensity of their emotions.

[1048] "Means of selecting information" refers to a method of selecting data appropriate for the user based on their analyzed emotional state.

[1049] "Means of providing information" refers to the methods and interfaces used to present selected data to users.

[1050] As an application example of this invention, a system is realized that recommends content based on the user's emotional state. The system uses a smartphone or smart glasses. First, the user inputs voice information into the smart device. This voice information is sent from the device to a server via the internet. The server uses speech recognition software (e.g., Google Cloud Speech-to-Text API) to convert the voice into textual dialogue information. Furthermore, the server uses sentiment analysis software (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state from the text.

[1051] Based on the identified emotional state, the server selects content appropriate for the user. This selection uses a relevance score based on past user data and data from similar users. Once content selection is complete, the server sends the selection results to the terminal, and the user receives the content. The received data is then displayed visually, allowing the user to experience it.

[1052] As a concrete example, suppose a user voice-inputs "I'm tired today." In this case, the system determines that the emotional state is "fatigue" and recommends relaxing music or stress-relieving videos to alleviate that state.

[1053] An example of a prompt is, "Please recommend music that can help me relax when I feel stressed." By inputting this prompt into a generating AI model, the system can quickly provide content that is suitable for the user.

[1054] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1055] Step 1:

[1056] The user inputs voice information using a smart device. This voice information is collected by the device and transmitted directly to the server. The input is voice data, and the output from the device to the server is also voice data.

[1057] Step 2:

[1058] The server uses speech recognition software to convert the received audio information into text. During this process, data processing is performed to convert the audio data into string data. The output is the textualized dialogue information.

[1059] Step 3:

[1060] The server processes the transcribed dialogue information using sentiment analysis software to extract the user's emotional state. The data calculation performed here involves assigning emotion labels through text analysis. The input is text information, and the output is emotional state data.

[1061] Step 4:

[1062] The server selects appropriate content from its content database based on the user's emotional state. The selection process uses accumulated data to calculate a relevance score. The input is emotional state data, and the output is information about the recommended content.

[1063] Step 5:

[1064] The server transmits selected content information to the terminal, which then visualizes and presents that information to the user. Here, visualization is used to communicate the data to the user. The input is content information, and the output is display data for the user.

[1065] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1066] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1067] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1068] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1069] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1070] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1071] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1072] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1073] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1074] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1075] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1076] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1077] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1078] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1079] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1080] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1081] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1082] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1083] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1084] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1085] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1086] The following is further disclosed regarding the embodiments described above.

[1087] (Claim 1)

[1088] A means of obtaining voice data from the user,

[1089] A means for analyzing the aforementioned audio data to generate user voiceprint information and transcribed dialogue content,

[1090] A means of analyzing the user's emotions, preferences, and communication style from the generated dialogue content and voiceprint information,

[1091] Based on the aforementioned analysis results, a means for selecting compatible partner candidates,

[1092] A means of providing the user with information on the aforementioned potential partners,

[1093] A system that includes this.

[1094] (Claim 2)

[1095] The system according to claim 1, further comprising means for calculating a compatibility score using data from other users accumulated in the past when selecting the aforementioned partner candidate.

[1096] (Claim 3)

[1097] The system according to claim 1, further comprising means for continuously monitoring the content of the user's dialogue and generating advice for improving the relationship.

[1098] "Example 1"

[1099] (Claim 1)

[1100] A device for acquiring voice information,

[1101] A device that analyzes the aforementioned audio information to generate audio feature information and written communication content,

[1102] A device that analyzes feelings, preferences, and communication styles from generated communication content and voice characteristic information,

[1103] Based on the aforementioned analysis results, a device is provided to select highly compatible partner candidates,

[1104] A device that provides information on the aforementioned potential opponents to the user,

[1105] A structure that includes this.

[1106] (Claim 2)

[1107] The structure according to claim 1, further comprising a device for calculating a goodness-of-fit index using records of other users accumulated in advance when selecting the aforementioned candidate partner.

[1108] (Claim 3)

[1109] The structure according to claim 1, further comprising a device for continuously monitoring the content of the user's communications and generating advice for improving the relationship.

[1110] "Application Example 1"

[1111] (Claim 1)

[1112] A terminal equipped with the function to acquire voice data from users provides a means for collecting voice data in real time,

[1113] A system that analyzes the aforementioned audio data to generate user voiceprint information and transcribed dialogue content,

[1114] A technology that analyzes the user's emotions, preferences, and communication style from the generated dialogue content and voiceprint information,

[1115] Based on the aforementioned analysis results, a means is provided to select a suitable partner, present candidates to the user in real time, and enable access to the dialogue.

[1116] A method for analyzing the user's emotions and interests in real time from the user's voice data during a conversation and making appropriate suggestions according to the situation,

[1117] A system that includes this.

[1118] (Claim 2)

[1119] The system according to claim 1, further comprising a function to calculate a compatibility score using data from other users accumulated in the past when selecting the candidate, and a means to generate metadata along with the analyzed partner information and provide it to the user.

[1120] (Claim 3)

[1121] The system according to claim 1, further comprising the function of continuously monitoring the content of the user's dialogue, analyzing the user's feedback in real time, generating specific advice to facilitate the relationship, and providing it to the user in real time.

[1122] "Example 2 of combining an emotion engine"

[1123] (Claim 1)

[1124] A means of obtaining voice data from the user,

[1125] A means for analyzing the aforementioned audio data to generate user acoustic characteristic information and transcribed dialogue content,

[1126] A means for analyzing the user's emotional state, preferences, and dialogue tendencies using generated dialogue content and acoustic feature information,

[1127] A method for selecting highly suitable social relationship candidates based on analysis results using a generative AI model,

[1128] A display means for providing users with information on selected candidates for social relationships,

[1129] A system that includes this.

[1130] (Claim 2)

[1131] The system according to claim 1, further comprising means for calculating a goodness of fit score based on information of other users collected in the past when selecting the candidate social relationship.

[1132] (Claim 3)

[1133] The system according to claim 1, further comprising means for continuously monitoring the content of the user's dialogue and generating advice for maintaining or improving a good relationship.

[1134] "Application example 2 when combining with an emotional engine"

[1135] (Claim 1)

[1136] A means of obtaining voice information from the user,

[1137] A means for analyzing the aforementioned audio information to generate user voiceprint information and transcribed dialogue information,

[1138] A means for analyzing the user's emotions, preferences, and communication style from generated dialogue information and voiceprint information,

[1139] Based on the aforementioned analysis results, a means for selecting appropriate information according to the emotional state,

[1140] Means for providing the aforementioned information to the user,

[1141] A system that includes this.

[1142] (Claim 2)

[1143] The system according to claim 1, further comprising means for calculating a relevance score using data from other users accumulated in the past when selecting the aforementioned information.

[1144] (Claim 3)

[1145] The system according to claim 1, further comprising means for continuously monitoring the user's interaction information and generating guidelines for improving the experience. [Explanation of Symbols]

[1146] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of obtaining voice data from the user, A means for analyzing the aforementioned audio data to generate user voiceprint information and transcribed dialogue content, A means of analyzing the user's emotions, preferences, and communication style from the generated dialogue content and voiceprint information, Based on the aforementioned analysis results, a means for selecting compatible partner candidates, A means of providing the user with information on the aforementioned potential partners, A system that includes this.

2. The system according to claim 1, further comprising means for calculating a compatibility score using data from other users accumulated in the past when selecting the aforementioned partner candidate.

3. The system according to claim 1, further comprising means for continuously monitoring the content of the user's dialogue and generating advice for improving the relationship.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A