system

The system addresses the inefficiencies of conventional text-input reviews by allowing voice-based input, sentiment analysis, and personalized summaries to enhance consumer decision-making.

JP2026074868APending Publication Date: 2026-05-07SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Consumers face challenges in accessing reliable word-of-mouth information due to the time-consuming nature of conventional text-input reviews, which often contain false information, making it difficult for enterprises to gain consumer trust and hindering efficient decision-making.

Method used

A system that allows consumers to post reviews using voice input, converting it to text data, performing sentiment analysis, evaluating reliability, and providing personalized summaries based on user preferences.

Benefits of technology

Enables efficient and reliable information delivery by converting voice reviews to text, analyzing sentiment, and generating concise summaries, thereby supporting informed purchasing decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074868000001_ABST
    Figure 2026074868000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Voice input method, A speech recognition means for converting the aforementioned voice input into text data, A sentiment analysis means for analyzing emotions based on the aforementioned text data, A truthfulness evaluation method for evaluating the reliability of word-of-mouth based on the aforementioned text data and sentiment analysis results, A summary generation means that summarizes the aforementioned text data and presents it to the user, Output means for providing the user with the aforementioned evaluation and summary results, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In online shopping and service use, consumers can access a lot of word-of-mouth information, but there is a problem that reliable information is insufficient. Conventional text-input word-of-mouth is time-consuming and contains mixed false reviews, which hinders consumers' decision-making. Due to such a situation, consumers spend time on purchase judgments, and there is also a problem that it is difficult for the enterprise side to gain consumers' trust.

Means for Solving the Problems

[0005] This invention solves the above problems by providing a system that allows consumers to easily post reviews using voice input. Specifically, it converts the input voice into text data using a voice recognition means. Furthermore, it provides highly reliable feedback by performing sentiment analysis based on the text data and evaluating the truthfulness of the review content. In addition, it presents the generated summary to the user and provides personalized information that takes into account the user's past preferences, thereby creating an environment in which consumers can make purchase decisions efficiently.

[0006] "Voice input means" refers to a device or software for acquiring digital voice data used by users to post reviews via voice.

[0007] "Speech recognition means" refers to a technology or process for converting voice-input data into text data.

[0008] "Sentiment analysis methods" are techniques that extract user emotions from transcribed reviews and classify them as positive, negative, neutral, etc.

[0009] A "truthfulness evaluation method" is a system or method for evaluating the reliability of word-of-mouth text data by comparing it with past data and external information, and for calculating a reliability score.

[0010] A "summary generation method" is a generation AI technology that concisely summarizes the content of reviews and highlights and presents information that is important to consumers.

[0011] "Output means" refers to a device or interface that provides users with a summary or evaluation result of processed reviews. [Brief explanation of the drawing]

[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0013] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0020] [First Embodiment]

[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0033] The system of the present invention is a platform that combines speech recognition technology and generative AI, allowing consumers to easily post reviews using voice and automatically processing them. An embodiment thereof is described below.

[0034] First, the user records their review via a voice input device. The device temporarily stores this audio data and, if necessary, sends it to a server in an encrypted format over the internet. The audio data received by the server is converted into text data by speech recognition. This text data is then analyzed for sentiment using Natural Language Processing (NLP) to determine whether the emotion of the review is positive, negative, or neutral.

[0035] Next, the server uses a truthfulness evaluation tool to assess the reliability of the converted text. This evaluation calculates a reliability score by utilizing past similar review data and information from external sources. For reviews whose reliability has been confirmed, the content is summarized by a generation AI, which extracts important information and expresses it in a concise form.

[0036] Finally, the generated summary and trust score are returned to the device. This includes personalized information based on the user's past review history and preferences, and other similar reviews are recommended. The device displays this information through the user interface, supporting consumers in making efficient purchasing decisions.

[0037] For example, if a user posts a voice review about a new home appliance saying, "It has many functions and is easy to use," the system transcribes the voice into text and detects positive sentiment. If it is confirmed that other users have shared similar experiences, the review is highly valued, and a summary such as "Highly rated for being highly functional and easy to use" is provided. In this way, it is possible to efficiently provide users with useful and reliable information.

[0038] The following describes the processing flow.

[0039] Step 1:

[0040] The user presses the voice input button on their device to record their review as voice. The device converts this voice data into a digital signal and stores it temporarily.

[0041] Step 2:

[0042] The device encrypts the stored audio data and sends it to the server using a secure communication protocol.

[0043] Step 3:

[0044] The server receives the audio data, processes it through a speech recognition system, and converts the audio into text data. The converted text is then stored in an internal database.

[0045] Step 4:

[0046] The server inputs text data into a natural language processing module and performs sentiment analysis. As a result of the analysis, sentiment scores of positive, negative, and neutral are calculated.

[0047] Step 5:

[0048] The server uses an authenticity assessment tool to evaluate the reliability of the text data. It compares the content with past databases and external sources and calculates a confidence score.

[0049] Step 6:

[0050] The server uses AI to summarize text data, extracting key information and presenting it in a concise and easy-to-read format.

[0051] Step 7:

[0052] The server sends a summary and confidence score back to the device. Personalized reviews and recommendations based on the user's preferences are also added.

[0053] Step 8:

[0054] The user interface displays information about the returned device. Users can review this information and use it as a reference when making a purchase decision.

[0055] (Example 1)

[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0057] Consumers need to be able to quickly and efficiently obtain reliable word-of-mouth information to support their decision-making regarding products and services. However, traditional methods have challenges such as difficulty in evaluating the authenticity and relevance of word-of-mouth, resulting in low convenience for users.

[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0059] In this invention, the server includes means for acquiring voice data via a voice input device, voice conversion means for converting the voice data into text data, and emotion analysis means for analyzing the text data and identifying emotional states. This enables users to efficiently acquire reliable word-of-mouth information.

[0060] A "voice input device" is a device used to acquire voice data and has the function of recording the user's voice.

[0061] "Voice conversion means" refers to a technology or device that converts acquired voice data into text data, and uses voice recognition technology to convert voice information into text information.

[0062] "Emotional analysis means" refers to a technology or device that analyzes textual data, identifies the emotions contained therein, and classifies them as positive, negative, or neutral.

[0063] A "reliability evaluation method" is a technology or device that uses past data and external information sources to evaluate the truthfulness and reliability of word-of-mouth information.

[0064] A "summary generation means" is a technology or device that extracts important parts of text information and summarizes them concisely, providing information in a format that is easy for the user to understand.

[0065] "Display means" refers to a technology or device that provides users with evaluation content and summary results, and displays the information visually.

[0066] To implement this invention, a system is constructed that combines voice input technology and a generative AI model. Users can input word-of-mouth information in voice format using a voice input device. The voice input device can be implemented using a smartphone or a dedicated recording device.

[0067] The device temporarily stores the recorded audio data. The audio data is transmitted to the server via the internet using AES encryption technology. The server converts the audio data into text data using speech recognition technology such as Google® Cloud Speech-to-Text. This extracts the content of the audio as specific textual information.

[0068] The server then uses natural language processing (NLP) techniques to analyze the text data and identifies the type of emotion using sentiment analysis tools. By utilizing sentiment analysis software such as TextBlob, it can be classified as either positive, negative, or neutral.

[0069] Furthermore, the server uses reliability evaluation tools to assess the reliability of the reviews. This evaluation includes comparison with similar historical data and referencing external information databases. Database management systems such as MySQL® and machine learning algorithms such as TENSORFLOW® are used.

[0070] If a user review is deemed reliable, a generative AI model is used to summarize the text. OpenAI's GPT model is used for summarization, concisely presenting information important to the user. Finally, the server returns the summary and confidence score to the device, which then displays the information visually to the user.

[0071] For example, if a user posts a voice comment saying, "The new laptop is fast and convenient," the system converts that information into text data and analyzes it as positive sentiment. If its reliability is confirmed, a summary such as "Highly rated for its speed and convenience" will be displayed. In this way, users can make decisions based on reliable summarized information.

[0072] An example of a prompt is, "Please provide example phrases for providing accurate reviews of home appliances." This serves as a reference for verbalizing the user's experience.

[0073] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0074] Step 1:

[0075] Users record reviews using a voice input device.

[0076] Input: Voice data provided by the user.

[0077] In practice, users operate their smartphones or recording devices to provide voice-based reviews. This voice data is temporarily stored on the device.

[0078] Step 2:

[0079] The device sends voice data to the server.

[0080] Input: Audio data stored on the device.

[0081] Output: Encrypted audio data sent to the server.

[0082] The device performs the specific action of protecting the voice data using AES encryption and sending it to the server over the internet.

[0083] Step 3:

[0084] The server converts the audio data into text data.

[0085] Input: Audio data sent to the server.

[0086] Output: Character data.

[0087] The server uses speech recognition software such as Google Cloud Speech-to-Text to analyze speech and convert it into text.

[0088] Step 4:

[0089] The server analyzes the content of the text data.

[0090] Input: Text data converted by speech recognition.

[0091] Output: Analysis data including sentiment analysis results.

[0092] The server analyzes the text data and uses TextBlob or similar tools to classify emotions as positive, negative, or neutral.

[0093] Step 5:

[0094] The server evaluates the reliability of the reviews.

[0095] Input: Analysis data including emotion analysis results.

[0096] Output: Evaluation data with confidence scores.

[0097] The server references historical databases and uses machine learning algorithms (such as TensorFlow) to perform specific calculations to evaluate the reliability of reviews. External information sources are also utilized in this evaluation.

[0098] Step 6:

[0099] The server generates a summary using an AI model.

[0100] Input: Evaluation data with confidence scores.

[0101] Output: Summary text.

[0102] The server uses OpenAI's GPT model to perform specific processing, such as extracting important information and generating a concise summary.

[0103] Step 7:

[0104] The server sends a summary and reliability information back to the terminal.

[0105] Input: Summary text and confidence score.

[0106] Output: Information displayed on the terminal.

[0107] The server sends a summary and reliability score to the terminal, which then performs the specific actions to display it appropriately on the user interface. Based on the information received, the user can understand the content of the reviews and use it to help make decisions.

[0108] (Application Example 1)

[0109] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0110] Traditional review platforms required consumers to manually read through a large volume of reviews, necessitating significant time and effort to make informed decisions. Furthermore, the insufficient availability of voice-based review submissions highlighted the need for efficient information delivery utilizing smart devices. This made it difficult for consumers to quickly obtain reliable information and make smooth purchasing decisions.

[0111] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0112] In this invention, the server includes voice input means, voice recognition means, sentiment analysis means, reliability evaluation means, summary generation means, and user interface means. This allows consumers to easily post reviews by voice, have that information processed quickly through a smart device, and receive it as reliable summary information.

[0113] A "voice input device" is a device that acquires voice as digital data and inputs that information into a system.

[0114] "Speech recognition means" refers to technology for converting input speech data into text data.

[0115] "Sentiment analysis tool" refers to a device or program that performs a process to classify the emotions contained in text data as positive, negative, or neutral.

[0116] A "reliability evaluation method" is a process for evaluating the veracity of word-of-mouth data and calculating a reliability score.

[0117] A "summary generation method" is a technology for extracting important information from text data and summarizing it concisely.

[0118] A "user interface means" is an interface that provides functions that allow a user to interact with the system and view and manipulate information.

[0119] This invention is a system that allows consumers to easily post reviews using voice, and that processes that information quickly and provides it in a useful format. It consists of a server, a voice input device, and a smart device.

[0120] Program processing

[0121] Users input reviews via voice using an application on their smart devices. A voice input means acquires this voice data and converts it into text data using a speech recognition means. The converted text is analyzed by a sentiment analysis means to determine whether its content is positive, negative, or neutral. A reliability evaluation means compares the obtained text data with an external data source to evaluate the truthfulness of the review and calculate a reliability score. Next, a summary generation means summarizes the important information from the text data and generates data to present it in an easy-to-understand manner for the user. These processes are managed and processed on a server, and ultimately, the information is presented to the user through a user interface means.

[0122] Specific example

[0123] For example, if a user posts a voice review about a new home appliance saying, "This rice cooker cooks rice deliciously and is easy to use," the voice is immediately transcribed into text and sentimentally analyzed as "positive." After verifying its reliability by cross-referencing it with other reviews, a summary is generated stating, "Highly rated for cooking delicious rice easily." Users can receive this summarized information through their smart devices and make decisions more efficiently.

[0124] Example of a prompt

[0125] "Perform a sentiment analysis on the following text, assess its reliability, and summarize it: 'This rice cooker cooks delicious rice and is easy to use.'"

[0126] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0127] Step 1:

[0128] Users input reviews by voice using an application on their smart devices. The input voice data is temporarily stored in the device's storage. At this stage, the input is voice data, and the output is a temporarily stored voice file.

[0129] Step 2:

[0130] The device sends the stored audio file to a speech recognition API. The audio data is converted to text via the server. This API analyzes the audio data and outputs it as text data in real time. The input is audio data, and the output is the converted text data.

[0131] Step 3:

[0132] The server passes the acquired text data to a sentiment analysis tool. The server uses a generative AI model to analyze the sentiment of the text and classify its content as positive, negative, or neutral. The input is text data, and the output is the result of the sentiment classification.

[0133] Step 4:

[0134] The server uses a reliability assessment function to evaluate the reliability of the review based on the results of the sentiment analysis. Here, a reliability score is calculated by cross-referencing with external information sources. The input is sentiment-classified text data, and the output is the reliability score.

[0135] Step 5:

[0136] The server inputs verified text data into its summary generation function. The summary generation function extracts key points from the text data and creates a concise summary. The input is verified text data, and the output is a summary.

[0137] Step 6:

[0138] The terminal displays the summary text and confidence score received from the server to the user through a user interface. This allows the user to efficiently obtain information and support their decision-making. The input is the summary text and confidence score, and the output is the information presented to the user.

[0139] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0140] The present invention provides a platform that combines speech recognition technology, natural language processing, and an emotion engine, enabling consumers to easily post reviews via voice, which are then automatically processed to provide personalized content. This embodiment is described below.

[0141] Users record reviews using a voice input device, and the audio data is converted into a digital signal by the terminal. After temporarily storing the signal, the terminal sends the audio data to the server via a secure protocol. The server transcribes the received data into text using speech recognition and then classifies the emotions using sentiment analysis. This identifies positive, negative, or neutral emotions.

[0142] Furthermore, a key feature of this invention is the inclusion of an emotion engine. The emotion engine recognizes the user's emotions in real time from voice data and transcribed reviews, and is responsible for generating personalized information that takes that emotional state into account. This allows for analysis of the user's emotional changes and long-term emotional trends, resulting in more personalized feedback.

[0143] Next, the server verifies the text data using a truthfulness evaluation system and scores the reliability of the reviews. This evaluation process is carried out using historical data and external reliability data sources. The evaluated data is passed to a summary generation system, where important information is extracted. The information provided to the user includes the summary information, the reliability score, and recommendation information based on the user's preferences, which is displayed on the terminal.

[0144] For example, if user A records a product review stating, "This product has a good design and excellent functionality," the emotion engine will highly rate user A's positive emotions. The server will then assign a confidence score, summarize the reasons for the high rating, and present it as, "Highly rated for its design and functionality." In this way, the present invention enables advanced analysis to improve the user experience and support purchasing decisions.

[0145] The following describes the processing flow.

[0146] Step 1:

[0147] The user presses the voice input button on the device to record their review of a product or service. The device converts this audio data into a digital signal and temporarily stores it on the device.

[0148] Step 2:

[0149] The device encrypts the stored audio data using a secure communication protocol and sends it to the server.

[0150] Step 3:

[0151] The server receives the audio data and passes it through a speech recognition system to convert the audio into text data. This conversion process generates the text data.

[0152] Step 4:

[0153] The server inputs text data into a sentiment analysis system, which then performs the analysis. This system identifies emotions as positive, negative, or neutral.

[0154] Step 5:

[0155] The emotion engine analyzes the emotions recognized from the user's voice input in detail and generates personalized information in real time based on that analysis. This analysis result is used as feedback.

[0156] Step 6:

[0157] The server uses a truth-of-fact evaluation tool to assess the reliability of the text data. This evaluation calculates a confidence score by referring to historical data and external data sources.

[0158] Step 7:

[0159] The server uses AI generation technology to summarize text data using a summarization generation mechanism. Important information is extracted and compiled into a concise overview.

[0160] Step 8:

[0161] The terminal displays summary information, confidence scores, and recommendations returned from the server in its user interface. This allows users to review detailed information and streamline their purchasing decisions.

[0162] (Example 2)

[0163] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0164] Voice-based word-of-mouth information presents challenges in judging emotional nuances and reliability, and in adequately addressing situations where information tailored to individual user characteristics is required. Furthermore, if the summarization and evaluation of transcribed information are not performed quickly and accurately, it will be insufficient in improving the user experience and supporting purchasing decisions.

[0165] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0166] In this invention, the server includes speech recognition means for converting voice input into text data, sentiment analysis means for classifying and analyzing emotions based on the text data, and personalized information generation means for individualizing the information. This enables accurate analysis of emotions and trustworthiness from voice data, and provides personalized information tailored to the user's preferences.

[0167] "Voice input means" refers to devices or methods for users to input voice as digital data.

[0168] "Speech recognition means" refers to the technology or process of converting speech data into text data.

[0169] "Sentiment analysis tools" refer to systems for evaluating and classifying emotions based on text data.

[0170] "Personalized information generation means" refers to a system that generates customized information for each user, taking into account the results of sentiment analysis.

[0171] "Truthfulness evaluation methods" refer to methods and systems for evaluating the reliability of information based on text data and sentiment analysis results.

[0172] "Summary generation means" refers to technology for extracting important information from text data and creating a summary.

[0173] "Output means" refers to devices or methods for displaying or providing evaluation and summary results to the user.

[0174] This invention includes a system that converts information acquired by a user through voice input from voice to text, analyzes that text to classify emotions, and provides personalized information. This system is implemented as follows:

[0175] Users record reviews and comments using a voice input device, such as a smartphone or tablet. The recorded audio data is converted into a digital signal and temporarily stored on the device. The device then sends the audio data to a server. This transmission uses a secure protocol, such as HTTPS, to ensure security.

[0176] The server converts the received audio data into text data using speech recognition software. For example, it utilizes speech recognition technology such as the Google Cloud Speech-to-Text API. Once the audio is converted to text, natural language processing (NLP) tools on the server analyze this text and perform sentiment analysis. This analysis uses a sentiment analysis engine to identify positive, negative, and neutral sentiment categories.

[0177] Furthermore, based on the sentiment analysis results, the system generates personalized information tailored to the user's emotional state. This allows users to receive information that is appropriate to their emotions and preferences. The server then evaluates the reliability of the word-of-mouth data. It compares it with past data and external sources to calculate a reliability score.

[0178] For example, if a user records a positive review of a product, for instance, saying, "This product has a good design and is very easy to use," the server will determine this review is positive and generate a summary of the product's design and ease of use. It can also provide the user with a list of other related products as recommendations.

[0179] An example of a prompt to input into a generative AI model might be: "A user has submitted a product review in audio format. The review states, 'This product has a good design and excellent functionality.' Convert the audio data into text, analyze the sentiment, and summarize its reliability and key points."

[0180] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0181] Step 1:

[0182] The user records their review using a voice input device. During this process, the user's voice is input. The device converts the recorded audio data into a digital signal. This digital conversion transforms the analog audio into digital data, which is then temporarily stored on the device.

[0183] Step 2:

[0184] The terminal transmits the digitally converted audio data to the server using a secure protocol such as HTTPS. The input is the converted audio data, and data communication takes place to securely transfer this data to the server.

[0185] Step 3:

[0186] The server converts received audio data into text data using speech recognition software. The input is digital audio data, and the output is speech-recognized text data. This conversion utilizes a speech recognition engine. Specifically, it involves analyzing the audio waveform data and converting it into a string based on a language model.

[0187] Step 4:

[0188] The server passes the converted text data to a sentiment analysis engine to classify the emotions. The input is text data, and the output is an emotion class such as positive, negative, or neutral. Here, natural language processing techniques are used to extract emotional characteristics from the text and classify them into emotion categories.

[0189] Step 5:

[0190] The server uses the results of sentiment analysis and text data to generate information tailored to the user through a personalized information generation system. The input is the sentiment analysis results and text data, and the output is personalized information that corresponds to the user's emotions and preferences. Specifically, this involves using a generative AI model to create personalized content that also takes into account the user's past behavioral data and preferences.

[0191] Step 6:

[0192] The server evaluates the reliability of the data obtained in the previous step. It verifies the authenticity of the reviews by comparing them with historical data and external reliability data sources. The input is text data and the results of sentiment analysis, and the output is a confidence score. This process involves data matching using data mining and machine learning.

[0193] Step 7:

[0194] The server summarizes text data using a summarization generation mechanism and presents it to the user. The input is text data with reliability ratings and sentiment results, and the output is summarized information. An algorithm is used for summarization, which extracts important information and condenses it into a shortened form.

[0195] Step 8:

[0196] The terminal displays summary information, confidence scores, and personalized recommendations received from the server on the user interface. Input is data from the server, and output is visually represented, user-friendly information. This step involves UI / UX design regarding how the information is displayed.

[0197] (Application Example 2)

[0198] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0199] Conventional voice-based word-of-mouth systems have struggled to accurately analyze emotional feedback from users and recommend personalized information. Furthermore, there has been a lack of means to properly evaluate the reliability of word-of-mouth, raising concerns about the provision of misleading information. This invention aims to solve these problems and provide users with highly reliable information based on their emotional tendencies.

[0200] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0201] In this invention, the server includes an input means for receiving voice, a voice recognition means for converting the voice into text information, an emotion analysis means for analyzing emotions based on the text information, a truthfulness evaluation means for evaluating the reliability of information based on the text information and the results of the emotion analysis, a summary generation means for summarizing the text information and presenting it to the user, a presentation means for supplying the evaluation and summary results to the user, and a recommendation means for recommending products based on the user's emotional tendencies. This makes it possible to smoothly convert user voice reviews into text and provide highly reliable information tailored to the individual user's preferences through emotion-based analysis.

[0202] A "voice input means" is a device that provides the function of collecting voice information emitted by a user and inputting it into the system.

[0203] "Speech recognition means" refers to technology that analyzes collected speech data and processes it to convert it into text information.

[0204] "Emotional analysis methods" are techniques for analyzing a user's emotions from textual information and classifying them as positive, negative, or neutral.

[0205] A "truthfulness evaluation method" is a function that implements a process for evaluating the reliability of word-of-mouth information and measuring the sincerity of the information.

[0206] A "summary generation method" is a technology that analyzes vast amounts of text data, extracts important information, and summarizes it concisely.

[0207] "Output means" refers to a device for displaying or notifying the user of processed evaluation or summary information.

[0208] A "recommendation system" is a system that selects and provides the most suitable products and services based on the user's emotional tendencies and preferences.

[0209] The present invention provides a platform for effectively collecting, analyzing, and recommending word-of-mouth information via voice. The hardware and software configurations for realizing this system are described in detail below.

[0210] First, users use voice input devices such as smartphones to input their reviews of products and services via voice. This voice data is temporarily stored by the device. The voice input devices used here include smartphones, tablets, and voice recognition-enabled devices equipped with a standard microphone.

[0211] The device converts the audio data into a digital signal and sends it to the server via a secure protocol. The server uses the Google Cloud Speech-to-Text API to convert the audio data into text. This process transcribes the audio message into text, which is then available for use in the next parsing step.

[0212] The converted text information is analyzed using a sentiment analysis method based on TensorFlow and classified as positive, negative, or neutral. This analysis result is then used to provide personalized information based on the product's characteristics.

[0213] Furthermore, a truthfulness evaluation tool compares the word-of-mouth information with past user data and reliable information sources to assess its reliability. Using the Django framework, a summary generation tool extracts key information based on this evaluation data and provides it in a format that is easily understandable to the user.

[0214] The recommendation system suggests relevant products based on the user's emotional tendencies and past preference data. This information is then presented to the user via the display of a smartphone or tablet.

[0215] For example, if a user enters "The image quality and ease of use of this camera are excellent," the system will interpret this review as positive and recommend similar high-quality cameras.

[0216] Examples of prompt statements to input into the generating AI model are as follows:

[0217] "Please provide an audio review of the product you purchased. Based on your review, we will share the product's features and reasons for recommending it with other potential customers."

[0218] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0219] Step 1:

[0220] Users speak their reviews using a voice input device. The input is an analog voice signal, which the device converts into a digital signal for speech recognition. The digitized voice data is then transmitted to the server via a secure protocol.

[0221] Step 2:

[0222] The server sends the received digital audio data to the Google Cloud Speech-to-Text API to convert it into text. The input is audio data, which is analyzed by an algorithm, and the output is text data.

[0223] Step 3:

[0224] The server processes text data using a sentiment analysis tool powered by TensorFlow. The input is text data, which is classified as positive, negative, or neutral based on a sentiment classification model. The output is data indicating the user's emotional state.

[0225] Step 4:

[0226] The server evaluates the reliability of text data using a truthfulness assessment tool. The input consists of text data and external reliability data, which are compared against past review history, and a reliability score is generated as output. This process improves the quality of information presented to the user.

[0227] Step 5:

[0228] The server summarizes text data using a summarization generation mechanism. The input consists of detailed text data and confidence scores, and the summarization algorithm extracts important information, resulting in simplified summary information as output.

[0229] Step 6:

[0230] The server recommends products based on user sentiment and preference data using recommendation methods. Inputs are past user data and current sentiment analysis results. Data mining techniques are used to select appropriate product candidates, and a customized list of recommended products is generated as output.

[0231] Step 7:

[0232] The terminal presents the user with recommended products and summary data. Input is data transmitted from the server, and the information is displayed on the screen, allowing the user to visually and audibly confirm the information as output.

[0233] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0234] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0235] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0236] [Second Embodiment]

[0237] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0238] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0239] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0240] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0241] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0242] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0243] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0244] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0245] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0246] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0247] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0248] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0249] The system of the present invention is a platform that combines speech recognition technology and generative AI, allowing consumers to easily post reviews using voice and automatically processing them. An embodiment thereof is described below.

[0250] First, the user records their review via a voice input device. The device temporarily stores this audio data and, if necessary, sends it to a server in an encrypted format over the internet. The audio data received by the server is converted into text data by speech recognition. This text data is then analyzed for sentiment using Natural Language Processing (NLP) to determine whether the emotion of the review is positive, negative, or neutral.

[0251] Next, the server uses a truthfulness evaluation tool to assess the reliability of the converted text. This evaluation calculates a reliability score by utilizing past similar review data and information from external sources. For reviews whose reliability has been confirmed, the content is summarized by a generation AI, which extracts important information and expresses it in a concise form.

[0252] Finally, the generated summary and trust score are returned to the device. This includes personalized information based on the user's past review history and preferences, and other similar reviews are recommended. The device displays this information through the user interface, supporting consumers in making efficient purchasing decisions.

[0253] For example, if a user posts a voice review about a new home appliance saying, "It has many functions and is easy to use," the system transcribes the voice into text and detects positive sentiment. If it is confirmed that other users have shared similar experiences, the review is highly valued, and a summary such as "Highly rated for being highly functional and easy to use" is provided. In this way, it is possible to efficiently provide users with useful and reliable information.

[0254] The following describes the processing flow.

[0255] Step 1:

[0256] The user presses the voice input button on their device to record their review as voice. The device converts this voice data into a digital signal and stores it temporarily.

[0257] Step 2:

[0258] The device encrypts the stored audio data and sends it to the server using a secure communication protocol.

[0259] Step 3:

[0260] The server receives the audio data, processes it through a speech recognition system, and converts the audio into text data. The converted text is then stored in an internal database.

[0261] Step 4:

[0262] The server inputs text data into a natural language processing module and performs sentiment analysis. As a result of the analysis, sentiment scores of positive, negative, and neutral are calculated.

[0263] Step 5:

[0264] The server uses an authenticity assessment tool to evaluate the reliability of the text data. It compares the content with past databases and external sources and calculates a confidence score.

[0265] Step 6:

[0266] The server uses AI to summarize text data, extracting key information and presenting it in a concise and easy-to-read format.

[0267] Step 7:

[0268] The server sends a summary and confidence score back to the device. Personalized reviews and recommendations based on the user's preferences are also added.

[0269] Step 8:

[0270] The user interface displays information about the returned device. Users can review this information and use it as a reference when making a purchase decision.

[0271] (Example 1)

[0272] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0273] Consumers need to be able to quickly and efficiently obtain reliable word-of-mouth information to support their decision-making regarding products and services. However, traditional methods have challenges such as difficulty in evaluating the authenticity and relevance of word-of-mouth, resulting in low convenience for users.

[0274] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0275] In this invention, the server includes means for acquiring voice data via a voice input device, voice conversion means for converting the voice data into text data, and emotion analysis means for analyzing the text data and identifying emotional states. This enables users to efficiently acquire reliable word-of-mouth information.

[0276] A "voice input device" is a device used to acquire voice data and has the function of recording the user's voice.

[0277] "Voice conversion means" refers to a technology or device that converts acquired voice data into text data, and uses voice recognition technology to convert voice information into text information.

[0278] "Emotional analysis means" refers to a technology or device that analyzes textual data, identifies the emotions contained therein, and classifies them as positive, negative, or neutral.

[0279] The "reliability evaluation means" is a technology or device that evaluates the authenticity and reliability of word-of-mouth using past data and external information sources.

[0280] The "summary generation means" is a technology or device that extracts important parts of text information and compiles them compactly, providing information in a form that is easy for users to understand.

[0281] The "display means" is a technology or device that provides evaluation content and summary results to users and visually displays information.

[0282] To implement this invention, a system combining voice input technology and a generative AI model is constructed. The user can input word-of-mouth information in voice format using a voice input device. The voice input device is realized using a smartphone or a dedicated recording device.

[0283] The terminal temporarily stores the recorded voice data. The voice data is transmitted to the server via the Internet using AES encryption technology. The server converts the voice data into text data using a voice recognition technology such as Google Cloud Speech-to-Text. Thereby, the content of the voice is extracted as specific character information.

[0284] After that, the server analyzes the text data using natural language processing (NLP) technology and identifies the type of emotion using emotion analysis means. By utilizing emotion analysis software such as TextBlob, it can be classified into any of positive, negative, or neutral.

[0285] Furthermore, the server evaluates the reliability of the word-of-mouth using the reliability evaluation means. This evaluation includes comparison with past similar data and reference to an external information database. A database management system such as MySQL and a machine learning algorithm such as TensorFlow are used.

[0286] When the review is determined to be reliable, a text summary is performed using the generative AI model. OpenAI's GPT model is used for summary creation, which concisely represents information important to the user. Finally, the server returns the summary and the trust score to the terminal, and the terminal visually displays the information to the user.

[0287] As a specific example, when a user posts a review such as "The new laptop is fast and convenient" by voice, the system converts the information into text data and analyzes it as a positive sentiment. If the reliability is confirmed, "High-speed processing and high convenience are highly evaluated" will be displayed as a summary. In this way, the user can make decisions based on reliable summary information.

[0288] An example of a prompt sentence is "Please give examples of phrases for providing accurate reviews of household appliances." This serves as a reference when verbalizing the user's experience.

[0289] The flow of the specific process in Example 1 will be described using FIG. 11.

[0290] Step 1:

[0291] The user records a review using the voice input device.

[0292] Input: Voice data by the user.

[0293] As a specific operation, the user operates a smartphone or a recording device and provides a review by voice. This voice data is temporarily stored in the terminal.

[0294] Step 2:

[0295] The terminal sends the voice data to the server.

[0296] Input: Voice data stored in the terminal.

[0297] Output: Encrypted voice data sent to the server.

[0298] The terminal performs specific operations to protect the voice data using AES encryption and send it to the server via the Internet.

[0299] Step 3:

[0300] The server converts the voice data into character data.

[0301] Input: Voice data sent to the server.

[0302] Output: Character data.

[0303] The server uses speech recognition software such as Google Cloud Speech-to-Text to perform specific operations to analyze the voice and convert it into text.

[0304] Step 4:

[0305] The server analyzes the content of the character data.

[0306] Input: Character data converted by speech recognition.

[0307] Output: Analysis data including sentiment analysis results.

[0308] The server analyzes the text data and performs specific operations to classify the sentiment as positive, negative, or neutral using tools such as TextBlob or similar.

[0309] Step 5:

[0310] The server evaluates the reliability of the review.

[0311] Input: Analysis data including sentiment analysis results.

[0312] Output: Evaluation data with a trust score.

[0313] The server references historical databases and uses machine learning algorithms (such as TensorFlow) to perform specific calculations to evaluate the reliability of reviews. External information sources are also utilized in this evaluation.

[0314] Step 6:

[0315] The server generates a summary using an AI model.

[0316] Input: Evaluation data with confidence scores.

[0317] Output: Summary text.

[0318] The server uses OpenAI's GPT model to perform specific processing, such as extracting important information and generating a concise summary.

[0319] Step 7:

[0320] The server sends a summary and reliability information back to the terminal.

[0321] Input: Summary text and confidence score.

[0322] Output: Information displayed on the terminal.

[0323] The server sends a summary and reliability score to the terminal, which then performs the specific actions to display it appropriately on the user interface. Based on the information received, the user can understand the content of the reviews and use it to help make decisions.

[0324] (Application Example 1)

[0325] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0326] Traditional review platforms required consumers to manually read through a large volume of reviews, necessitating significant time and effort to make informed decisions. Furthermore, the insufficient availability of voice-based review submissions highlighted the need for efficient information delivery utilizing smart devices. This made it difficult for consumers to quickly obtain reliable information and make smooth purchasing decisions.

[0327] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0328] In this invention, the server includes voice input means, voice recognition means, sentiment analysis means, reliability evaluation means, summary generation means, and user interface means. This allows consumers to easily post reviews by voice, have that information processed quickly through a smart device, and receive it as reliable summary information.

[0329] A "voice input device" is a device that acquires voice as digital data and inputs that information into a system.

[0330] "Speech recognition means" refers to technology for converting input speech data into text data.

[0331] "Sentiment analysis tool" refers to a device or program that performs a process to classify the emotions contained in text data as positive, negative, or neutral.

[0332] A "reliability evaluation method" is a process for evaluating the veracity of word-of-mouth data and calculating a reliability score.

[0333] A "summary generation method" is a technology for extracting important information from text data and summarizing it concisely.

[0334] A "user interface means" is an interface that provides functions that allow a user to interact with the system and view and manipulate information.

[0335] This invention is a system that allows consumers to easily post reviews using voice, and that processes that information quickly and provides it in a useful format. It consists of a server, a voice input device, and a smart device.

[0336] Program processing

[0337] Users input reviews via voice using an application on their smart devices. A voice input means acquires this voice data and converts it into text data using a speech recognition means. The converted text is analyzed by a sentiment analysis means to determine whether its content is positive, negative, or neutral. A reliability evaluation means compares the obtained text data with an external data source to evaluate the truthfulness of the review and calculate a reliability score. Next, a summary generation means summarizes the important information from the text data and generates data to present it in an easy-to-understand manner for the user. These processes are managed and processed on a server, and ultimately, the information is presented to the user through a user interface means.

[0338] Specific example

[0339] For example, if a user posts a voice review about a new home appliance saying, "This rice cooker cooks rice deliciously and is easy to use," the voice is immediately transcribed into text and sentimentally analyzed as "positive." After verifying its reliability by cross-referencing it with other reviews, a summary is generated stating, "Highly rated for cooking delicious rice easily." Users can receive this summarized information through their smart devices and make decisions more efficiently.

[0340] Example of a prompt

[0341] "Perform a sentiment analysis on the following text, assess its reliability, and summarize it: 'This rice cooker cooks delicious rice and is easy to use.'"

[0342] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0343] Step 1:

[0344] Users input reviews by voice using an application on their smart devices. The input voice data is temporarily stored in the device's storage. At this stage, the input is voice data, and the output is a temporarily stored voice file.

[0345] Step 2:

[0346] The device sends the stored audio file to a speech recognition API. The audio data is converted to text via the server. This API analyzes the audio data and outputs it as text data in real time. The input is audio data, and the output is the converted text data.

[0347] Step 3:

[0348] The server passes the acquired text data to a sentiment analysis tool. The server uses a generative AI model to analyze the sentiment of the text and classify its content as positive, negative, or neutral. The input is text data, and the output is the result of the sentiment classification.

[0349] Step 4:

[0350] The server uses a reliability assessment function to evaluate the reliability of the review based on the results of the sentiment analysis. Here, a reliability score is calculated by cross-referencing with external information sources. The input is sentiment-classified text data, and the output is the reliability score.

[0351] Step 5:

[0352] The server inputs verified text data into its summary generation function. The summary generation function extracts key points from the text data and creates a concise summary. The input is verified text data, and the output is a summary.

[0353] Step 6:

[0354] The terminal displays the summary text and confidence score received from the server to the user through a user interface. This allows the user to efficiently obtain information and support their decision-making. The input is the summary text and confidence score, and the output is the information presented to the user.

[0355] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0356] The present invention provides a platform that combines speech recognition technology, natural language processing, and an emotion engine, enabling consumers to easily post reviews via voice, which are then automatically processed to provide personalized content. This embodiment is described below.

[0357] Users record reviews using a voice input device, and the audio data is converted into a digital signal by the terminal. After temporarily storing the signal, the terminal sends the audio data to the server via a secure protocol. The server transcribes the received data into text using speech recognition and then classifies the emotions using sentiment analysis. This identifies positive, negative, or neutral emotions.

[0358] Furthermore, a key feature of this invention is the inclusion of an emotion engine. The emotion engine recognizes the user's emotions in real time from voice data and transcribed reviews, and is responsible for generating personalized information that takes that emotional state into account. This allows for analysis of the user's emotional changes and long-term emotional trends, resulting in more personalized feedback.

[0359] Next, the server verifies the text data using a truthfulness evaluation system and scores the reliability of the reviews. This evaluation process is carried out using historical data and external reliability data sources. The evaluated data is passed to a summary generation system, where important information is extracted. The information provided to the user includes the summary information, the reliability score, and recommendation information based on the user's preferences, which is displayed on the terminal.

[0360] For example, if user A records a product review stating, "This product has a good design and excellent functionality," the emotion engine will highly rate user A's positive emotions. The server will then assign a confidence score, summarize the reasons for the high rating, and present it as, "Highly rated for its design and functionality." In this way, the present invention enables advanced analysis to improve the user experience and support purchasing decisions.

[0361] The following describes the processing flow.

[0362] Step 1:

[0363] The user presses the voice input button on the device to record their review of a product or service. The device converts this audio data into a digital signal and temporarily stores it on the device.

[0364] Step 2:

[0365] The device encrypts the stored audio data using a secure communication protocol and sends it to the server.

[0366] Step 3:

[0367] The server receives the audio data and passes it through a speech recognition system to convert the audio into text data. This conversion process generates the text data.

[0368] Step 4:

[0369] The server inputs text data into a sentiment analysis system, which then performs the analysis. This system identifies emotions as positive, negative, or neutral.

[0370] Step 5:

[0371] The emotion engine analyzes the emotions recognized from the user's voice input in detail and generates personalized information in real time based on that analysis. This analysis result is used as feedback.

[0372] Step 6:

[0373] The server uses a truth-of-fact evaluation tool to assess the reliability of the text data. This evaluation calculates a confidence score by referring to historical data and external data sources.

[0374] Step 7:

[0375] The server uses AI generation technology to summarize text data using a summarization generation mechanism. Important information is extracted and compiled into a concise overview.

[0376] Step 8:

[0377] The terminal displays summary information, confidence scores, and recommendations returned from the server in its user interface. This allows users to review detailed information and streamline their purchasing decisions.

[0378] (Example 2)

[0379] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0380] Voice-based word-of-mouth information presents challenges in judging emotional nuances and reliability, and in adequately addressing situations where information tailored to individual user characteristics is required. Furthermore, if the summarization and evaluation of transcribed information are not performed quickly and accurately, it will be insufficient in improving the user experience and supporting purchasing decisions.

[0381] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0382] In this invention, the server includes speech recognition means for converting voice input into text data, sentiment analysis means for classifying and analyzing emotions based on the text data, and personalized information generation means for individualizing the information. This enables accurate analysis of emotions and trustworthiness from voice data, and provides personalized information tailored to the user's preferences.

[0383] "Voice input means" refers to devices or methods for users to input voice as digital data.

[0384] "Speech recognition means" refers to the technology or process of converting speech data into text data.

[0385] "Sentiment analysis tools" refer to systems for evaluating and classifying emotions based on text data.

[0386] "Personalized information generation means" refers to a system that generates customized information for each user, taking into account the results of sentiment analysis.

[0387] "Truthfulness evaluation methods" refer to methods and systems for evaluating the reliability of information based on text data and sentiment analysis results.

[0388] "Summary generation means" refers to technology for extracting important information from text data and creating a summary.

[0389] "Output means" refers to devices or methods for displaying or providing evaluation and summary results to the user.

[0390] This invention includes a system that converts information acquired by a user through voice input from voice to text, analyzes that text to classify emotions, and provides personalized information. This system is implemented as follows:

[0391] Users record reviews and comments using a voice input device, such as a smartphone or tablet. The recorded audio data is converted into a digital signal and temporarily stored on the device. The device then sends the audio data to a server. This transmission uses a secure protocol, such as HTTPS, to ensure security.

[0392] The server converts the received audio data into text data using speech recognition software. For example, it utilizes speech recognition technology such as the Google Cloud Speech-to-Text API. Once the audio is converted to text, natural language processing (NLP) tools on the server analyze this text and perform sentiment analysis. This analysis uses a sentiment analysis engine to identify positive, negative, and neutral sentiment categories.

[0393] Furthermore, based on the sentiment analysis results, the system generates personalized information tailored to the user's emotional state. This allows users to receive information that is appropriate to their emotions and preferences. The server then evaluates the reliability of the word-of-mouth data. It compares it with past data and external sources to calculate a reliability score.

[0394] For example, if a user records a positive review of a product, for instance, saying, "This product has a good design and is very easy to use," the server will determine this review is positive and generate a summary of the product's design and ease of use. It can also provide the user with a list of other related products as recommendations.

[0395] An example of a prompt to input into a generative AI model might be: "A user has submitted a product review in audio format. The review states, 'This product has a good design and excellent functionality.' Convert the audio data into text, analyze the sentiment, and summarize its reliability and key points."

[0396] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0397] Step 1:

[0398] The user records their review using a voice input device. During this process, the user's voice is input. The device converts the recorded audio data into a digital signal. This digital conversion transforms the analog audio into digital data, which is then temporarily stored on the device.

[0399] Step 2:

[0400] The terminal transmits the digitally converted audio data to the server using a secure protocol such as HTTPS. The input is the converted audio data, and data communication takes place to securely transfer this data to the server.

[0401] Step 3:

[0402] The server converts received audio data into text data using speech recognition software. The input is digital audio data, and the output is speech-recognized text data. This conversion utilizes a speech recognition engine. Specifically, it involves analyzing the audio waveform data and converting it into a string based on a language model.

[0403] Step 4:

[0404] The server passes the converted text data to a sentiment analysis engine to classify the emotions. The input is text data, and the output is an emotion class such as positive, negative, or neutral. Here, natural language processing techniques are used to extract emotional characteristics from the text and classify them into emotion categories.

[0405] Step 5:

[0406] The server uses the results of sentiment analysis and text data to generate information tailored to the user through a personalized information generation system. The input is the sentiment analysis results and text data, and the output is personalized information that corresponds to the user's emotions and preferences. Specifically, this involves using a generative AI model to create personalized content that also takes into account the user's past behavioral data and preferences.

[0407] Step 6:

[0408] The server evaluates the reliability of the data obtained in the previous step. It verifies the authenticity of the reviews by comparing them with historical data and external reliability data sources. The input is text data and the results of sentiment analysis, and the output is a confidence score. This process involves data matching using data mining and machine learning.

[0409] Step 7:

[0410] The server summarizes text data using a summarization generation mechanism and presents it to the user. The input is text data with reliability ratings and sentiment results, and the output is summarized information. An algorithm is used for summarization, which extracts important information and condenses it into a shortened form.

[0411] Step 8:

[0412] The terminal displays summary information, confidence scores, and personalized recommendations received from the server on the user interface. Input is data from the server, and output is visually represented, user-friendly information. This step involves UI / UX design regarding how the information is displayed.

[0413] (Application Example 2)

[0414] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0415] Conventional voice-based word-of-mouth systems have struggled to accurately analyze emotional feedback from users and recommend personalized information. Furthermore, there has been a lack of means to properly evaluate the reliability of word-of-mouth, raising concerns about the provision of misleading information. This invention aims to solve these problems and provide users with highly reliable information based on their emotional tendencies.

[0416] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0417] In this invention, the server includes an input means for receiving voice, a voice recognition means for converting the voice into text information, an emotion analysis means for analyzing emotions based on the text information, a truthfulness evaluation means for evaluating the reliability of information based on the text information and the results of the emotion analysis, a summary generation means for summarizing the text information and presenting it to the user, a presentation means for supplying the evaluation and summary results to the user, and a recommendation means for recommending products based on the user's emotional tendencies. This makes it possible to smoothly convert user voice reviews into text and provide highly reliable information tailored to the individual user's preferences through emotion-based analysis.

[0418] A "voice input means" is a device that provides the function of collecting voice information emitted by a user and inputting it into the system.

[0419] "Speech recognition means" refers to technology that analyzes collected speech data and processes it to convert it into text information.

[0420] "Emotional analysis methods" are techniques for analyzing a user's emotions from textual information and classifying them as positive, negative, or neutral.

[0421] A "truthfulness evaluation method" is a function that implements a process for evaluating the reliability of word-of-mouth information and measuring the sincerity of the information.

[0422] A "summary generation method" is a technology that analyzes vast amounts of text data, extracts important information, and summarizes it concisely.

[0423] "Output means" refers to a device for displaying or notifying the user of processed evaluation or summary information.

[0424] A "recommendation system" is a system that selects and provides the most suitable products and services based on the user's emotional tendencies and preferences.

[0425] The present invention provides a platform for effectively collecting, analyzing, and recommending word-of-mouth information via voice. The hardware and software configurations for realizing this system are described in detail below.

[0426] First, users use voice input devices such as smartphones to input their reviews of products and services via voice. This voice data is temporarily stored by the device. The voice input devices used here include smartphones, tablets, and voice recognition-enabled devices equipped with a standard microphone.

[0427] The device converts the audio data into a digital signal and sends it to the server via a secure protocol. The server uses the Google Cloud Speech-to-Text API to convert the audio data into text. This process transcribes the audio message into text, which is then available for use in the next parsing step.

[0428] The converted text information is analyzed using a sentiment analysis method based on TensorFlow and classified as positive, negative, or neutral. This analysis result is then used to provide personalized information based on the product's characteristics.

[0429] Furthermore, a truthfulness evaluation tool compares the word-of-mouth information with past user data and reliable information sources to assess its reliability. Using the Django framework, a summary generation tool extracts key information based on this evaluation data and provides it in a format that is easily understandable to the user.

[0430] The recommendation system suggests relevant products based on the user's emotional tendencies and past preference data. This information is then presented to the user via the display of a smartphone or tablet.

[0431] For example, if a user enters "The image quality and ease of use of this camera are excellent," the system will interpret this review as positive and recommend similar high-quality cameras.

[0432] Examples of prompt statements to input into the generating AI model are as follows:

[0433] "Please provide an audio review of the product you purchased. Based on your review, we will share the product's features and reasons for recommending it with other potential customers."

[0434] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0435] Step 1:

[0436] Users speak their reviews using a voice input device. The input is an analog voice signal, which the device converts into a digital signal for speech recognition. The digitized voice data is then transmitted to the server via a secure protocol.

[0437] Step 2:

[0438] The server sends the received digital audio data to the Google Cloud Speech-to-Text API to convert it into text. The input is audio data, which is analyzed by an algorithm, and the output is text data.

[0439] Step 3:

[0440] The server processes text data using a sentiment analysis tool powered by TensorFlow. The input is text data, which is classified as positive, negative, or neutral based on a sentiment classification model. The output is data indicating the user's emotional state.

[0441] Step 4:

[0442] The server evaluates the reliability of text data using a truthfulness assessment tool. The input consists of text data and external reliability data, which are compared against past review history, and a reliability score is generated as output. This process improves the quality of information presented to the user.

[0443] Step 5:

[0444] The server summarizes text data using a summarization generation mechanism. The input consists of detailed text data and confidence scores, and the summarization algorithm extracts important information, resulting in simplified summary information as output.

[0445] Step 6:

[0446] The server recommends products based on user sentiment and preference data using recommendation methods. Inputs are past user data and current sentiment analysis results. Data mining techniques are used to select appropriate product candidates, and a customized list of recommended products is generated as output.

[0447] Step 7:

[0448] The terminal presents the user with recommended products and summary data. Input is data transmitted from the server, and the information is displayed on the screen, allowing the user to visually and audibly confirm the information as output.

[0449] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0450] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0451] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0452] [Third Embodiment]

[0453] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0454] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0455] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0456] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0457] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0458] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0459] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0460] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0461] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0462] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0463] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0464] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0465] The system of the present invention is a platform that combines speech recognition technology and generative AI, allowing consumers to easily post reviews using voice and automatically processing them. An embodiment thereof is described below.

[0466] First, the user records their review via a voice input device. The device temporarily stores this audio data and, if necessary, sends it to a server in an encrypted format over the internet. The audio data received by the server is converted into text data by speech recognition. This text data is then analyzed for sentiment using Natural Language Processing (NLP) to determine whether the emotion of the review is positive, negative, or neutral.

[0467] Next, the server uses a truthfulness evaluation tool to assess the reliability of the converted text. This evaluation calculates a reliability score by utilizing past similar review data and information from external sources. For reviews whose reliability has been confirmed, the content is summarized by a generation AI, which extracts important information and expresses it in a concise form.

[0468] Finally, the generated summary and trust score are returned to the device. This includes personalized information based on the user's past review history and preferences, and other similar reviews are recommended. The device displays this information through the user interface, supporting consumers in making efficient purchasing decisions.

[0469] For example, if a user posts a voice review about a new home appliance saying, "It has many functions and is easy to use," the system transcribes the voice into text and detects positive sentiment. If it is confirmed that other users have shared similar experiences, the review is highly valued, and a summary such as "Highly rated for being highly functional and easy to use" is provided. In this way, it is possible to efficiently provide users with useful and reliable information.

[0470] The following describes the processing flow.

[0471] Step 1:

[0472] The user presses the voice input button on their device to record their review as voice. The device converts this voice data into a digital signal and stores it temporarily.

[0473] Step 2:

[0474] The device encrypts the stored audio data and sends it to the server using a secure communication protocol.

[0475] Step 3:

[0476] The server receives the audio data, processes it through a speech recognition system, and converts the audio into text data. The converted text is then stored in an internal database.

[0477] Step 4:

[0478] The server inputs text data into a natural language processing module and performs sentiment analysis. As a result of the analysis, sentiment scores of positive, negative, and neutral are calculated.

[0479] Step 5:

[0480] The server uses an authenticity assessment tool to evaluate the reliability of the text data. It compares the content with past databases and external sources and calculates a confidence score.

[0481] Step 6:

[0482] The server uses AI to summarize text data, extracting key information and presenting it in a concise and easy-to-read format.

[0483] Step 7:

[0484] The server sends a summary and confidence score back to the device. Personalized reviews and recommendations based on the user's preferences are also added.

[0485] Step 8:

[0486] The user interface displays information about the returned device. Users can review this information and use it as a reference when making a purchase decision.

[0487] (Example 1)

[0488] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0489] Consumers need to be able to quickly and efficiently obtain reliable word-of-mouth information to support their decision-making regarding products and services. However, traditional methods have challenges such as difficulty in evaluating the authenticity and relevance of word-of-mouth, resulting in low convenience for users.

[0490] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0491] In this invention, the server includes means for acquiring voice data via a voice input device, voice conversion means for converting the voice data into text data, and emotion analysis means for analyzing the text data and identifying emotional states. This enables users to efficiently acquire reliable word-of-mouth information.

[0492] A "voice input device" is a device used to acquire voice data and has the function of recording the user's voice.

[0493] "Voice conversion means" refers to a technology or device that converts acquired voice data into text data, and uses voice recognition technology to convert voice information into text information.

[0494] "Emotional analysis means" refers to a technology or device that analyzes textual data, identifies the emotions contained therein, and classifies them as positive, negative, or neutral.

[0495] A "reliability evaluation method" is a technology or device that uses past data and external information sources to evaluate the truthfulness and reliability of word-of-mouth information.

[0496] A "summary generation means" is a technology or device that extracts important parts of text information and summarizes them concisely, providing information in a format that is easy for the user to understand.

[0497] "Display means" refers to a technology or device that provides users with evaluation content and summary results, and displays the information visually.

[0498] To implement this invention, a system is constructed that combines voice input technology and a generative AI model. Users can input word-of-mouth information in voice format using a voice input device. The voice input device can be implemented using a smartphone or a dedicated recording device.

[0499] The device temporarily stores the recorded audio data. The audio data is transmitted to the server via the internet using AES encryption technology. The server converts the audio data into text data using speech recognition technology such as Google Cloud Speech-to-Text. This extracts the content of the audio as specific textual information.

[0500] The server then uses natural language processing (NLP) techniques to analyze the text data and identifies the type of emotion using sentiment analysis tools. By utilizing sentiment analysis software such as TextBlob, it can be classified as either positive, negative, or neutral.

[0501] Furthermore, the server uses reliability evaluation tools to assess the reliability of the reviews. This evaluation includes comparison with similar historical data and referencing external information databases. Database management systems such as MySQL and machine learning algorithms such as TensorFlow are used.

[0502] If a user review is deemed reliable, a generative AI model is used to summarize the text. OpenAI's GPT model is used to create the summary, concisely representing information important to the user. Finally, the server returns the summary and confidence score to the device, which then displays the information visually to the user.

[0503] For example, if a user posts a voice comment saying, "The new laptop is fast and convenient," the system converts that information into text data and analyzes it as positive sentiment. If its reliability is confirmed, a summary such as "Highly rated for its speed and convenience" will be displayed. In this way, users can make decisions based on reliable summarized information.

[0504] An example of a prompt is, "Please provide example phrases for providing accurate reviews of home appliances." This serves as a reference for verbalizing the user's experience.

[0505] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0506] Step 1:

[0507] Users record reviews using a voice input device.

[0508] Input: Voice data provided by the user.

[0509] In practice, users operate their smartphones or recording devices to provide voice-based reviews. This voice data is temporarily stored on the device.

[0510] Step 2:

[0511] The device sends voice data to the server.

[0512] Input: Audio data stored on the device.

[0513] Output: Encrypted audio data sent to the server.

[0514] The device performs the specific action of protecting the voice data using AES encryption and sending it to the server over the internet.

[0515] Step 3:

[0516] The server converts the audio data into text data.

[0517] Input: Audio data sent to the server.

[0518] Output: Character data.

[0519] The server uses speech recognition software such as Google Cloud Speech-to-Text to analyze speech and convert it into text.

[0520] Step 4:

[0521] The server analyzes the content of the text data.

[0522] Input: Text data converted by speech recognition.

[0523] Output: Analysis data including sentiment analysis results.

[0524] The server analyzes the text data and uses TextBlob or similar tools to classify emotions as positive, negative, or neutral.

[0525] Step 5:

[0526] The server evaluates the reliability of the reviews.

[0527] Input: Analysis data including emotion analysis results.

[0528] Output: Evaluation data with confidence scores.

[0529] The server references historical databases and uses machine learning algorithms (such as TensorFlow) to perform specific calculations to evaluate the reliability of reviews. External information sources are also utilized in this evaluation.

[0530] Step 6:

[0531] The server generates a summary using an AI model.

[0532] Input: Evaluation data with confidence scores.

[0533] Output: Summary text.

[0534] The server uses OpenAI's GPT model to perform specific processing, such as extracting important information and generating a concise summary.

[0535] Step 7:

[0536] The server sends a summary and reliability information back to the terminal.

[0537] Input: Summary text and confidence score.

[0538] Output: Information displayed on the terminal.

[0539] The server sends a summary and reliability score to the terminal, which then performs the specific actions to display it appropriately on the user interface. Based on the information received, the user can understand the content of the reviews and use it to help make decisions.

[0540] (Application Example 1)

[0541] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0542] Traditional review platforms required consumers to manually read through a large volume of reviews, necessitating significant time and effort to make informed decisions. Furthermore, the insufficient availability of voice-based review submissions highlighted the need for efficient information delivery utilizing smart devices. This made it difficult for consumers to quickly obtain reliable information and make smooth purchasing decisions.

[0543] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0544] In this invention, the server includes voice input means, voice recognition means, sentiment analysis means, reliability evaluation means, summary generation means, and user interface means. This allows consumers to easily post reviews by voice, have that information processed quickly through a smart device, and receive it as reliable summary information.

[0545] A "voice input device" is a device that acquires voice as digital data and inputs that information into a system.

[0546] "Speech recognition means" refers to technology for converting input speech data into text data.

[0547] "Sentiment analysis tool" refers to a device or program that performs a process to classify the emotions contained in text data as positive, negative, or neutral.

[0548] A "reliability evaluation method" is a process for evaluating the veracity of word-of-mouth data and calculating a reliability score.

[0549] A "summary generation method" is a technology for extracting important information from text data and summarizing it concisely.

[0550] A "user interface means" is an interface that provides functions that allow a user to interact with the system and view and manipulate information.

[0551] This invention is a system that allows consumers to easily post reviews using voice, and that processes that information quickly and provides it in a useful format. It consists of a server, a voice input device, and a smart device.

[0552] Program processing

[0553] Users input reviews via voice using an application on their smart devices. A voice input means acquires this voice data and converts it into text data using a speech recognition means. The converted text is analyzed by a sentiment analysis means to determine whether its content is positive, negative, or neutral. A reliability evaluation means compares the obtained text data with an external data source to evaluate the truthfulness of the review and calculate a reliability score. Next, a summary generation means summarizes the important information from the text data and generates data to present it in an easy-to-understand manner for the user. These processes are managed and processed on a server, and ultimately, the information is presented to the user through a user interface means.

[0554] Specific example

[0555] For example, if a user posts a voice review about a new home appliance saying, "This rice cooker cooks rice deliciously and is easy to use," the voice is immediately transcribed into text and sentimentally analyzed as "positive." After verifying its reliability by cross-referencing it with other reviews, a summary is generated stating, "Highly rated for cooking delicious rice easily." Users can receive this summarized information through their smart devices and make decisions more efficiently.

[0556] Example of a prompt

[0557] "Perform a sentiment analysis on the following text, assess its reliability, and summarize it: 'This rice cooker cooks delicious rice and is easy to use.'"

[0558] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0559] Step 1:

[0560] Users input reviews by voice using an application on their smart devices. The input voice data is temporarily stored in the device's storage. At this stage, the input is voice data, and the output is a temporarily stored voice file.

[0561] Step 2:

[0562] The device sends the stored audio file to a speech recognition API. The audio data is converted to text via the server. This API analyzes the audio data and outputs it as text data in real time. The input is audio data, and the output is the converted text data.

[0563] Step 3:

[0564] The server passes the acquired text data to a sentiment analysis tool. The server uses a generative AI model to analyze the sentiment of the text and classify its content as positive, negative, or neutral. The input is text data, and the output is the result of the sentiment classification.

[0565] Step 4:

[0566] The server uses a reliability assessment function to evaluate the reliability of the review based on the results of the sentiment analysis. Here, a reliability score is calculated by cross-referencing with external information sources. The input is sentiment-classified text data, and the output is the reliability score.

[0567] Step 5:

[0568] The server inputs verified text data into its summary generation function. The summary generation function extracts key points from the text data and creates a concise summary. The input is verified text data, and the output is a summary.

[0569] Step 6:

[0570] The terminal displays the summary text and confidence score received from the server to the user through a user interface. This allows the user to efficiently obtain information and support their decision-making. The input is the summary text and confidence score, and the output is the information presented to the user.

[0571] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0572] The present invention provides a platform that combines speech recognition technology, natural language processing, and an emotion engine, enabling consumers to easily post reviews via voice, which are then automatically processed to provide personalized content. This embodiment is described below.

[0573] Users record reviews using a voice input device, and the audio data is converted into a digital signal by the terminal. After temporarily storing the signal, the terminal sends the audio data to the server via a secure protocol. The server transcribes the received data into text using speech recognition and then classifies the emotions using sentiment analysis. This identifies positive, negative, or neutral emotions.

[0574] Furthermore, a key feature of this invention is the inclusion of an emotion engine. The emotion engine recognizes the user's emotions in real time from voice data and transcribed reviews, and is responsible for generating personalized information that takes that emotional state into account. This allows for analysis of the user's emotional changes and long-term emotional trends, resulting in more personalized feedback.

[0575] Next, the server verifies the text data using a truthfulness evaluation system and scores the reliability of the reviews. This evaluation process is carried out using historical data and external reliability data sources. The evaluated data is passed to a summary generation system, where important information is extracted. The information provided to the user includes the summary information, the reliability score, and recommendation information based on the user's preferences, which is displayed on the terminal.

[0576] For example, if user A records a product review stating, "This product has a good design and excellent functionality," the emotion engine will highly rate user A's positive emotions. The server will then assign a confidence score, summarize the reasons for the high rating, and present it as, "Highly rated for its design and functionality." In this way, the present invention enables advanced analysis to improve the user experience and support purchasing decisions.

[0577] The following describes the processing flow.

[0578] Step 1:

[0579] The user presses the voice input button on the device to record their review of a product or service. The device converts this audio data into a digital signal and temporarily stores it on the device.

[0580] Step 2:

[0581] The device encrypts the stored audio data using a secure communication protocol and sends it to the server.

[0582] Step 3:

[0583] The server receives the audio data and passes it through a speech recognition system to convert the audio into text data. This conversion process generates the text data.

[0584] Step 4:

[0585] The server inputs text data into a sentiment analysis system, which then performs the analysis. This system identifies emotions as positive, negative, or neutral.

[0586] Step 5:

[0587] The emotion engine analyzes the emotions recognized from the user's voice input in detail and generates personalized information in real time based on that analysis. This analysis result is used as feedback.

[0588] Step 6:

[0589] The server uses a truth-of-fact evaluation tool to assess the reliability of the text data. This evaluation calculates a confidence score by referring to historical data and external data sources.

[0590] Step 7:

[0591] The server uses AI generation technology to summarize text data using a summarization generation mechanism. Important information is extracted and compiled into a concise overview.

[0592] Step 8:

[0593] The terminal displays summary information, confidence scores, and recommendations returned from the server in its user interface. This allows users to review detailed information and streamline their purchasing decisions.

[0594] (Example 2)

[0595] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0596] Voice-based word-of-mouth information presents challenges in judging emotional nuances and reliability, and in adequately addressing situations where information tailored to individual user characteristics is required. Furthermore, if the summarization and evaluation of transcribed information are not performed quickly and accurately, it will be insufficient in improving the user experience and supporting purchasing decisions.

[0597] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0598] In this invention, the server includes speech recognition means for converting voice input into text data, sentiment analysis means for classifying and analyzing emotions based on the text data, and personalized information generation means for individualizing the information. This enables accurate analysis of emotions and trustworthiness from voice data, and provides personalized information tailored to the user's preferences.

[0599] "Voice input means" refers to devices or methods for users to input voice as digital data.

[0600] "Speech recognition means" refers to the technology or process of converting speech data into text data.

[0601] "Sentiment analysis tools" refer to systems for evaluating and classifying emotions based on text data.

[0602] "Personalized information generation means" refers to a system that generates customized information for each user, taking into account the results of sentiment analysis.

[0603] "Truthfulness evaluation methods" refer to methods and systems for evaluating the reliability of information based on text data and sentiment analysis results.

[0604] "Summary generation means" refers to technology for extracting important information from text data and creating a summary.

[0605] "Output means" refers to devices or methods for displaying or providing evaluation and summary results to the user.

[0606] This invention includes a system that converts information acquired by a user through voice input from voice to text, analyzes that text to classify emotions, and provides personalized information. This system is implemented as follows:

[0607] Users record reviews and comments using a voice input device, such as a smartphone or tablet. The recorded audio data is converted into a digital signal and temporarily stored on the device. The device then sends the audio data to a server. This transmission uses a secure protocol, such as HTTPS, to ensure security.

[0608] The server converts the received audio data into text data using speech recognition software. For example, it utilizes speech recognition technology such as the Google Cloud Speech-to-Text API. Once the audio is converted to text, natural language processing (NLP) tools on the server analyze this text and perform sentiment analysis. This analysis uses a sentiment analysis engine to identify positive, negative, and neutral sentiment categories.

[0609] Furthermore, based on the sentiment analysis results, the system generates personalized information tailored to the user's emotional state. This allows users to receive information that is appropriate to their emotions and preferences. The server then evaluates the reliability of the word-of-mouth data. It compares it with past data and external sources to calculate a reliability score.

[0610] For example, if a user records a positive review of a product, for instance, saying, "This product has a good design and is very easy to use," the server will determine this review is positive and generate a summary of the product's design and ease of use. It can also provide the user with a list of other related products as recommendations.

[0611] An example of a prompt to input into a generative AI model might be: "A user has submitted a product review in audio format. The review states, 'This product has a good design and excellent functionality.' Convert the audio data into text, analyze the sentiment, and summarize its reliability and key points."

[0612] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0613] Step 1:

[0614] The user records their review using a voice input device. During this process, the user's voice is input. The device converts the recorded audio data into a digital signal. This digital conversion transforms the analog audio into digital data, which is then temporarily stored on the device.

[0615] Step 2:

[0616] The terminal transmits the digitally converted audio data to the server using a secure protocol such as HTTPS. The input is the converted audio data, and data communication takes place to securely transfer this data to the server.

[0617] Step 3:

[0618] The server converts received audio data into text data using speech recognition software. The input is digital audio data, and the output is speech-recognized text data. This conversion utilizes a speech recognition engine. Specifically, it involves analyzing the audio waveform data and converting it into a string based on a language model.

[0619] Step 4:

[0620] The server passes the converted text data to a sentiment analysis engine to classify the emotions. The input is text data, and the output is an emotion class such as positive, negative, or neutral. Here, natural language processing techniques are used to extract emotional characteristics from the text and classify them into emotion categories.

[0621] Step 5:

[0622] The server uses the results of sentiment analysis and text data to generate information tailored to the user through a personalized information generation system. The input is the sentiment analysis results and text data, and the output is personalized information that corresponds to the user's emotions and preferences. Specifically, this involves using a generative AI model to create personalized content that also takes into account the user's past behavioral data and preferences.

[0623] Step 6:

[0624] The server evaluates the reliability of the data obtained in the previous step. It verifies the authenticity of the reviews by comparing them with historical data and external reliability data sources. The input is text data and the results of sentiment analysis, and the output is a confidence score. This process involves data matching using data mining and machine learning.

[0625] Step 7:

[0626] The server summarizes text data using a summarization generation mechanism and presents it to the user. The input is text data with reliability ratings and sentiment results, and the output is summarized information. An algorithm is used for summarization, which extracts important information and condenses it into a shortened form.

[0627] Step 8:

[0628] The terminal displays summary information, confidence scores, and personalized recommendations received from the server on the user interface. Input is data from the server, and output is visually represented, user-friendly information. This step involves UI / UX design regarding how the information is displayed.

[0629] (Application Example 2)

[0630] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0631] Conventional voice-based word-of-mouth systems have struggled to accurately analyze emotional feedback from users and recommend personalized information. Furthermore, there has been a lack of means to properly evaluate the reliability of word-of-mouth, raising concerns about the provision of misleading information. This invention aims to solve these problems and provide users with highly reliable information based on their emotional tendencies.

[0632] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0633] In this invention, the server includes an input means for receiving voice, a voice recognition means for converting the voice into text information, an emotion analysis means for analyzing emotions based on the text information, a truthfulness evaluation means for evaluating the reliability of information based on the text information and the results of the emotion analysis, a summary generation means for summarizing the text information and presenting it to the user, a presentation means for supplying the evaluation and summary results to the user, and a recommendation means for recommending products based on the user's emotional tendencies. This makes it possible to smoothly convert user voice reviews into text and provide highly reliable information tailored to the individual user's preferences through emotion-based analysis.

[0634] A "voice input means" is a device that provides the function of collecting voice information emitted by a user and inputting it into the system.

[0635] "Speech recognition means" refers to technology that analyzes collected speech data and processes it to convert it into text information.

[0636] "Emotional analysis methods" are techniques for analyzing a user's emotions from textual information and classifying them as positive, negative, or neutral.

[0637] A "truthfulness evaluation method" is a function that implements a process for evaluating the reliability of word-of-mouth information and measuring the sincerity of the information.

[0638] A "summary generation method" is a technology that analyzes vast amounts of text data, extracts important information, and summarizes it concisely.

[0639] "Output means" refers to a device for displaying or notifying the user of processed evaluation or summary information.

[0640] A "recommendation system" is a system that selects and provides the most suitable products and services based on the user's emotional tendencies and preferences.

[0641] The present invention provides a platform for effectively collecting, analyzing, and recommending word-of-mouth information via voice. The hardware and software configurations for realizing this system are described in detail below.

[0642] First, users use voice input devices such as smartphones to input their reviews of products and services via voice. This voice data is temporarily stored by the device. The voice input devices used here include smartphones, tablets, and voice recognition-enabled devices equipped with a standard microphone.

[0643] The device converts the audio data into a digital signal and sends it to the server via a secure protocol. The server uses the Google Cloud Speech-to-Text API to convert the audio data into text. This process transcribes the audio message into text, which is then available for use in the next parsing step.

[0644] The converted text information is analyzed using a sentiment analysis method based on TensorFlow and classified as positive, negative, or neutral. This analysis result is then used to provide personalized information based on the product's characteristics.

[0645] Furthermore, a truthfulness evaluation tool compares the word-of-mouth information with past user data and reliable information sources to assess its reliability. Using the Django framework, a summary generation tool extracts key information based on this evaluation data and provides it in a format that is easily understandable to the user.

[0646] The recommendation system suggests relevant products based on the user's emotional tendencies and past preference data. This information is then presented to the user via the display of a smartphone or tablet.

[0647] For example, if a user enters "The image quality and ease of use of this camera are excellent," the system will interpret this review as positive and recommend similar high-quality cameras.

[0648] Examples of prompt statements to input into the generating AI model are as follows:

[0649] "Please provide an audio review of the product you purchased. Based on your review, we will share the product's features and reasons for recommending it with other potential customers."

[0650] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0651] Step 1:

[0652] Users speak their reviews using a voice input device. The input is an analog voice signal, which the device converts into a digital signal for speech recognition. The digitized voice data is then transmitted to the server via a secure protocol.

[0653] Step 2:

[0654] The server sends the received digital audio data to the Google Cloud Speech-to-Text API to convert it into text. The input is audio data, which is analyzed by an algorithm, and the output is text data.

[0655] Step 3:

[0656] The server processes text data using a sentiment analysis tool powered by TensorFlow. The input is text data, which is classified as positive, negative, or neutral based on a sentiment classification model. The output is data indicating the user's emotional state.

[0657] Step 4:

[0658] The server evaluates the reliability of text data using a truthfulness assessment tool. The input consists of text data and external reliability data, which are compared against past review history, and a reliability score is generated as output. This process improves the quality of information presented to the user.

[0659] Step 5:

[0660] The server summarizes text data using a summarization generation mechanism. The input consists of detailed text data and confidence scores, and the summarization algorithm extracts important information, resulting in simplified summary information as output.

[0661] Step 6:

[0662] The server recommends products based on user sentiment and preference data using recommendation methods. Inputs are past user data and current sentiment analysis results. Data mining techniques are used to select appropriate product candidates, and a customized list of recommended products is generated as output.

[0663] Step 7:

[0664] The terminal presents the user with recommended products and summary data. Input is data transmitted from the server, and the information is displayed on the screen, allowing the user to visually and audibly confirm the information as output.

[0665] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0666] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0667] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0668] [Fourth Embodiment]

[0669] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0670] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0671] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0672] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0673] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0674] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0675] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0676] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0677] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0678] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0679] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0680] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0681] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0682] The system of the present invention is a platform that combines speech recognition technology and generative AI, allowing consumers to easily post reviews using voice and automatically processing them. An embodiment thereof is described below.

[0683] First, the user records their review via a voice input device. The device temporarily stores this audio data and, if necessary, sends it to a server in an encrypted format over the internet. The audio data received by the server is converted into text data by speech recognition. This text data is then analyzed for sentiment using Natural Language Processing (NLP) to determine whether the emotion of the review is positive, negative, or neutral.

[0684] Next, the server uses a truthfulness evaluation tool to assess the reliability of the converted text. This evaluation calculates a reliability score by utilizing past similar review data and information from external sources. For reviews whose reliability has been confirmed, the content is summarized by a generation AI, which extracts important information and expresses it in a concise form.

[0685] Finally, the generated summary and trust score are returned to the device. This includes personalized information based on the user's past review history and preferences, and other similar reviews are recommended. The device displays this information through the user interface, supporting consumers in making efficient purchasing decisions.

[0686] For example, if a user posts a voice review about a new home appliance saying, "It has many functions and is easy to use," the system transcribes the voice into text and detects positive sentiment. If it is confirmed that other users have shared similar experiences, the review is highly valued, and a summary such as "Highly rated for being highly functional and easy to use" is provided. In this way, it is possible to efficiently provide users with useful and reliable information.

[0687] The following describes the processing flow.

[0688] Step 1:

[0689] The user presses the voice input button on their device to record their review as voice. The device converts this voice data into a digital signal and stores it temporarily.

[0690] Step 2:

[0691] The device encrypts the stored audio data and sends it to the server using a secure communication protocol.

[0692] Step 3:

[0693] The server receives the audio data, processes it through a speech recognition system, and converts the audio into text data. The converted text is then stored in an internal database.

[0694] Step 4:

[0695] The server inputs text data into a natural language processing module and performs sentiment analysis. As a result of the analysis, sentiment scores of positive, negative, and neutral are calculated.

[0696] Step 5:

[0697] The server uses an authenticity assessment tool to evaluate the reliability of the text data. It compares the content with past databases and external sources and calculates a confidence score.

[0698] Step 6:

[0699] The server uses AI to summarize text data, extracting key information and presenting it in a concise and easy-to-read format.

[0700] Step 7:

[0701] The server sends a summary and confidence score back to the device. Personalized reviews and recommendations based on the user's preferences are also added.

[0702] Step 8:

[0703] The user interface displays information about the returned device. Users can review this information and use it as a reference when making a purchase decision.

[0704] (Example 1)

[0705] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0706] Consumers need to be able to quickly and efficiently obtain reliable word-of-mouth information to support their decision-making regarding products and services. However, traditional methods have challenges such as difficulty in evaluating the authenticity and relevance of word-of-mouth, resulting in low convenience for users.

[0707] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0708] In this invention, the server includes means for acquiring voice data via a voice input device, voice conversion means for converting the voice data into text data, and emotion analysis means for analyzing the text data and identifying emotional states. This enables users to efficiently acquire reliable word-of-mouth information.

[0709] A "voice input device" is a device used to acquire voice data and has the function of recording the user's voice.

[0710] "Voice conversion means" refers to a technology or device that converts acquired voice data into text data, and uses voice recognition technology to convert voice information into text information.

[0711] "Emotional analysis means" refers to a technology or device that analyzes textual data, identifies the emotions contained therein, and classifies them as positive, negative, or neutral.

[0712] A "reliability evaluation method" is a technology or device that uses past data and external information sources to evaluate the truthfulness and reliability of word-of-mouth information.

[0713] A "summary generation means" is a technology or device that extracts important parts of text information and summarizes them concisely, providing information in a format that is easy for the user to understand.

[0714] "Display means" refers to a technology or device that provides users with evaluation content and summary results, and displays the information visually.

[0715] To implement this invention, a system is constructed that combines voice input technology and a generative AI model. Users can input word-of-mouth information in voice format using a voice input device. The voice input device can be implemented using a smartphone or a dedicated recording device.

[0716] The device temporarily stores the recorded audio data. The audio data is transmitted to the server via the internet using AES encryption technology. The server converts the audio data into text data using speech recognition technology such as Google Cloud Speech-to-Text. This extracts the content of the audio as specific textual information.

[0717] The server then uses natural language processing (NLP) techniques to analyze the text data and identifies the type of emotion using sentiment analysis tools. By utilizing sentiment analysis software such as TextBlob, it can be classified as either positive, negative, or neutral.

[0718] Furthermore, the server uses reliability evaluation tools to assess the reliability of the reviews. This evaluation includes comparison with similar historical data and referencing external information databases. Database management systems such as MySQL and machine learning algorithms such as TensorFlow are used.

[0719] If a user review is deemed reliable, a generative AI model is used to summarize the text. OpenAI's GPT model is used to create the summary, concisely representing information important to the user. Finally, the server returns the summary and confidence score to the device, which then displays the information visually to the user.

[0720] For example, if a user posts a voice comment saying, "The new laptop is fast and convenient," the system converts that information into text data and analyzes it as positive sentiment. If its reliability is confirmed, a summary such as "Highly rated for its speed and convenience" will be displayed. In this way, users can make decisions based on reliable summarized information.

[0721] An example of a prompt is, "Please provide example phrases for providing accurate reviews of home appliances." This serves as a reference for verbalizing the user's experience.

[0722] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0723] Step 1:

[0724] Users record reviews using a voice input device.

[0725] Input: Voice data provided by the user.

[0726] In practice, users operate their smartphones or recording devices to provide voice-based reviews. This voice data is temporarily stored on the device.

[0727] Step 2:

[0728] The device sends voice data to the server.

[0729] Input: Audio data stored on the device.

[0730] Output: Encrypted audio data sent to the server.

[0731] The device performs the specific action of protecting the voice data using AES encryption and sending it to the server over the internet.

[0732] Step 3:

[0733] The server converts the audio data into text data.

[0734] Input: Audio data sent to the server.

[0735] Output: Character data.

[0736] The server uses speech recognition software such as Google Cloud Speech-to-Text to analyze speech and convert it into text.

[0737] Step 4:

[0738] The server analyzes the content of the text data.

[0739] Input: Text data converted by speech recognition.

[0740] Output: Analysis data including sentiment analysis results.

[0741] The server analyzes the text data and uses TextBlob or similar tools to classify emotions as positive, negative, or neutral.

[0742] Step 5:

[0743] The server evaluates the reliability of the reviews.

[0744] Input: Analysis data including emotion analysis results.

[0745] Output: Evaluation data with confidence scores.

[0746] The server references historical databases and uses machine learning algorithms (such as TensorFlow) to perform specific calculations to evaluate the reliability of reviews. External information sources are also utilized in this evaluation.

[0747] Step 6:

[0748] The server generates a summary using an AI model.

[0749] Input: Evaluation data with confidence scores.

[0750] Output: Summary text.

[0751] The server uses OpenAI's GPT model to perform specific processing, such as extracting important information and generating a concise summary.

[0752] Step 7:

[0753] The server sends a summary and reliability information back to the terminal.

[0754] Input: Summary text and confidence score.

[0755] Output: Information displayed on the terminal.

[0756] The server sends a summary and reliability score to the terminal, which then performs the specific actions to display it appropriately on the user interface. Based on the information received, the user can understand the content of the reviews and use it to help make decisions.

[0757] (Application Example 1)

[0758] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0759] Traditional review platforms required consumers to manually read through a large volume of reviews, necessitating significant time and effort to make informed decisions. Furthermore, the insufficient availability of voice-based review submissions highlighted the need for efficient information delivery utilizing smart devices. This made it difficult for consumers to quickly obtain reliable information and make smooth purchasing decisions.

[0760] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0761] In this invention, the server includes voice input means, voice recognition means, sentiment analysis means, reliability evaluation means, summary generation means, and user interface means. This allows consumers to easily post reviews by voice, have that information processed quickly through a smart device, and receive it as reliable summary information.

[0762] A "voice input device" is a device that acquires voice as digital data and inputs that information into a system.

[0763] "Speech recognition means" refers to technology for converting input speech data into text data.

[0764] "Sentiment analysis tool" refers to a device or program that performs a process to classify the emotions contained in text data as positive, negative, or neutral.

[0765] A "reliability evaluation method" is a process for evaluating the veracity of word-of-mouth data and calculating a reliability score.

[0766] A "summary generation method" is a technology for extracting important information from text data and summarizing it concisely.

[0767] A "user interface means" is an interface that provides functions that allow a user to interact with the system and view and manipulate information.

[0768] This invention is a system that allows consumers to easily post reviews using voice, and that processes that information quickly and provides it in a useful format. It consists of a server, a voice input device, and a smart device.

[0769] Program processing

[0770] Users input reviews via voice using an application on their smart devices. A voice input means acquires this voice data and converts it into text data using a speech recognition means. The converted text is analyzed by a sentiment analysis means to determine whether its content is positive, negative, or neutral. A reliability evaluation means compares the obtained text data with an external data source to evaluate the truthfulness of the review and calculate a reliability score. Next, a summary generation means summarizes the important information from the text data and generates data to present it in an easy-to-understand manner for the user. These processes are managed and processed on a server, and ultimately, the information is presented to the user through a user interface means.

[0771] Specific example

[0772] For example, if a user posts a voice review about a new home appliance saying, "This rice cooker cooks rice deliciously and is easy to use," the voice is immediately transcribed into text and sentimentally analyzed as "positive." After verifying its reliability by cross-referencing it with other reviews, a summary is generated stating, "Highly rated for cooking delicious rice easily." Users can receive this summarized information through their smart devices and make decisions more efficiently.

[0773] Example of a prompt

[0774] "Perform a sentiment analysis on the following text, assess its reliability, and summarize it: 'This rice cooker cooks delicious rice and is easy to use.'"

[0775] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0776] Step 1:

[0777] Users input reviews by voice using an application on their smart devices. The input voice data is temporarily stored in the device's storage. At this stage, the input is voice data, and the output is a temporarily stored voice file.

[0778] Step 2:

[0779] The device sends the stored audio file to a speech recognition API. The audio data is converted to text via the server. This API analyzes the audio data and outputs it as text data in real time. The input is audio data, and the output is the converted text data.

[0780] Step 3:

[0781] The server passes the acquired text data to a sentiment analysis tool. The server uses a generative AI model to analyze the sentiment of the text and classify its content as positive, negative, or neutral. The input is text data, and the output is the result of the sentiment classification.

[0782] Step 4:

[0783] The server uses a reliability assessment function to evaluate the reliability of the review based on the results of the sentiment analysis. Here, a reliability score is calculated by cross-referencing with external information sources. The input is sentiment-classified text data, and the output is the reliability score.

[0784] Step 5:

[0785] The server inputs verified text data into its summary generation function. The summary generation function extracts key points from the text data and creates a concise summary. The input is verified text data, and the output is a summary.

[0786] Step 6:

[0787] The terminal displays the summary text and confidence score received from the server to the user through a user interface. This allows the user to efficiently obtain information and support their decision-making. The input is the summary text and confidence score, and the output is the information presented to the user.

[0788] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0789] The present invention provides a platform that combines speech recognition technology, natural language processing, and an emotion engine, enabling consumers to easily post reviews via voice, which are then automatically processed to provide personalized content. This embodiment is described below.

[0790] Users record reviews using a voice input device, and the audio data is converted into a digital signal by the terminal. After temporarily storing the signal, the terminal sends the audio data to the server via a secure protocol. The server transcribes the received data into text using speech recognition and then classifies the emotions using sentiment analysis. This identifies positive, negative, or neutral emotions.

[0791] Furthermore, a key feature of this invention is the inclusion of an emotion engine. The emotion engine recognizes the user's emotions in real time from voice data and transcribed reviews, and is responsible for generating personalized information that takes that emotional state into account. This allows for analysis of the user's emotional changes and long-term emotional trends, resulting in more personalized feedback.

[0792] Next, the server verifies the text data using a truthfulness evaluation system and scores the reliability of the reviews. This evaluation process is carried out using historical data and external reliability data sources. The evaluated data is passed to a summary generation system, where important information is extracted. The information provided to the user includes the summary information, the reliability score, and recommendation information based on the user's preferences, which is displayed on the terminal.

[0793] For example, if user A records a product review stating, "This product has a good design and excellent functionality," the emotion engine will highly rate user A's positive emotions. The server will then assign a confidence score, summarize the reasons for the high rating, and present it as, "Highly rated for its design and functionality." In this way, the present invention enables advanced analysis to improve the user experience and support purchasing decisions.

[0794] The following describes the processing flow.

[0795] Step 1:

[0796] The user presses the voice input button on the device to record their review of a product or service. The device converts this audio data into a digital signal and temporarily stores it on the device.

[0797] Step 2:

[0798] The device encrypts the stored audio data using a secure communication protocol and sends it to the server.

[0799] Step 3:

[0800] The server receives the audio data and passes it through a speech recognition system to convert the audio into text data. This conversion process generates the text data.

[0801] Step 4:

[0802] The server inputs text data into a sentiment analysis system, which then performs the analysis. This system identifies emotions as positive, negative, or neutral.

[0803] Step 5:

[0804] The emotion engine analyzes the emotions recognized from the user's voice input in detail and generates personalized information in real time based on that analysis. This analysis result is used as feedback.

[0805] Step 6:

[0806] The server uses a truth-of-fact evaluation tool to assess the reliability of the text data. This evaluation calculates a confidence score by referring to historical data and external data sources.

[0807] Step 7:

[0808] The server uses AI generation technology to summarize text data using a summarization generation mechanism. Important information is extracted and compiled into a concise overview.

[0809] Step 8:

[0810] The terminal displays summary information, confidence scores, and recommendations returned from the server in its user interface. This allows users to review detailed information and streamline their purchasing decisions.

[0811] (Example 2)

[0812] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0813] Voice-based word-of-mouth information presents challenges in judging emotional nuances and reliability, and in adequately addressing situations where information tailored to individual user characteristics is required. Furthermore, if the summarization and evaluation of transcribed information are not performed quickly and accurately, it will be insufficient in improving the user experience and supporting purchasing decisions.

[0814] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0815] In this invention, the server includes speech recognition means for converting voice input into text data, sentiment analysis means for classifying and analyzing emotions based on the text data, and personalized information generation means for individualizing the information. This enables accurate analysis of emotions and trustworthiness from voice data, and provides personalized information tailored to the user's preferences.

[0816] "Voice input means" refers to devices or methods for users to input voice as digital data.

[0817] "Speech recognition means" refers to the technology or process of converting speech data into text data.

[0818] "Sentiment analysis tools" refer to systems for evaluating and classifying emotions based on text data.

[0819] "Personalized information generation means" refers to a system that generates customized information for each user, taking into account the results of sentiment analysis.

[0820] "Truthfulness evaluation methods" refer to methods and systems for evaluating the reliability of information based on text data and sentiment analysis results.

[0821] "Summary generation means" refers to technology for extracting important information from text data and creating a summary.

[0822] "Output means" refers to devices or methods for displaying or providing evaluation and summary results to the user.

[0823] This invention includes a system that converts information acquired by a user through voice input from voice to text, analyzes that text to classify emotions, and provides personalized information. This system is implemented as follows:

[0824] Users record reviews and comments using a voice input device, such as a smartphone or tablet. The recorded audio data is converted into a digital signal and temporarily stored on the device. The device then sends the audio data to a server. This transmission uses a secure protocol, such as HTTPS, to ensure security.

[0825] The server converts the received audio data into text data using speech recognition software. For example, it utilizes speech recognition technology such as the Google Cloud Speech-to-Text API. Once the audio is converted to text, natural language processing (NLP) tools on the server analyze this text and perform sentiment analysis. This analysis uses a sentiment analysis engine to identify positive, negative, and neutral sentiment categories.

[0826] Furthermore, based on the sentiment analysis results, the system generates personalized information tailored to the user's emotional state. This allows users to receive information that is appropriate to their emotions and preferences. The server then evaluates the reliability of the word-of-mouth data. It compares it with past data and external sources to calculate a reliability score.

[0827] For example, if a user records a positive review of a product, for instance, saying, "This product has a good design and is very easy to use," the server will determine this review is positive and generate a summary of the product's design and ease of use. It can also provide the user with a list of other related products as recommendations.

[0828] An example of a prompt to input into a generative AI model might be: "A user has submitted a product review in audio format. The review states, 'This product has a good design and excellent functionality.' Convert the audio data into text, analyze the sentiment, and summarize its reliability and key points."

[0829] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0830] Step 1:

[0831] The user records their review using a voice input device. During this process, the user's voice is input. The device converts the recorded audio data into a digital signal. This digital conversion transforms the analog audio into digital data, which is then temporarily stored on the device.

[0832] Step 2:

[0833] The terminal transmits the digitally converted audio data to the server using a secure protocol such as HTTPS. The input is the converted audio data, and data communication takes place to securely transfer this data to the server.

[0834] Step 3:

[0835] The server converts received audio data into text data using speech recognition software. The input is digital audio data, and the output is speech-recognized text data. This conversion utilizes a speech recognition engine. Specifically, it involves analyzing the audio waveform data and converting it into a string based on a language model.

[0836] Step 4:

[0837] The server passes the converted text data to a sentiment analysis engine to classify the emotions. The input is text data, and the output is an emotion class such as positive, negative, or neutral. Here, natural language processing techniques are used to extract emotional characteristics from the text and classify them into emotion categories.

[0838] Step 5:

[0839] The server uses the results of sentiment analysis and text data to generate information tailored to the user through a personalized information generation system. The input is the sentiment analysis results and text data, and the output is personalized information that corresponds to the user's emotions and preferences. Specifically, this involves using a generative AI model to create personalized content that also takes into account the user's past behavioral data and preferences.

[0840] Step 6:

[0841] The server evaluates the reliability of the data obtained in the previous step. It verifies the authenticity of the reviews by comparing them with historical data and external reliability data sources. The input is text data and the results of sentiment analysis, and the output is a confidence score. This process involves data matching using data mining and machine learning.

[0842] Step 7:

[0843] The server summarizes text data using a summarization generation mechanism and presents it to the user. The input is text data with reliability ratings and sentiment results, and the output is summarized information. An algorithm is used for summarization, which extracts important information and condenses it into a shortened form.

[0844] Step 8:

[0845] The terminal displays summary information, confidence scores, and personalized recommendations received from the server on the user interface. Input is data from the server, and output is visually represented, user-friendly information. This step involves UI / UX design regarding how the information is displayed.

[0846] (Application Example 2)

[0847] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0848] Conventional voice-based word-of-mouth systems have struggled to accurately analyze emotional feedback from users and recommend personalized information. Furthermore, there has been a lack of means to properly evaluate the reliability of word-of-mouth, raising concerns about the provision of misleading information. This invention aims to solve these problems and provide users with highly reliable information based on their emotional tendencies.

[0849] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0850] In this invention, the server includes an input means for receiving voice, a voice recognition means for converting the voice into text information, an emotion analysis means for analyzing emotions based on the text information, a truthfulness evaluation means for evaluating the reliability of information based on the text information and the results of the emotion analysis, a summary generation means for summarizing the text information and presenting it to the user, a presentation means for supplying the evaluation and summary results to the user, and a recommendation means for recommending products based on the user's emotional tendencies. This makes it possible to smoothly convert user voice reviews into text and provide highly reliable information tailored to the individual user's preferences through emotion-based analysis.

[0851] A "voice input means" is a device that provides the function of collecting voice information emitted by a user and inputting it into the system.

[0852] "Speech recognition means" refers to technology that analyzes collected speech data and processes it to convert it into text information.

[0853] "Emotional analysis methods" are techniques for analyzing a user's emotions from textual information and classifying them as positive, negative, or neutral.

[0854] A "truthfulness evaluation method" is a function that implements a process for evaluating the reliability of word-of-mouth information and measuring the sincerity of the information.

[0855] A "summary generation method" is a technology that analyzes vast amounts of text data, extracts important information, and summarizes it concisely.

[0856] "Output means" refers to a device for displaying or notifying the user of processed evaluation or summary information.

[0857] A "recommendation system" is a system that selects and provides the most suitable products and services based on the user's emotional tendencies and preferences.

[0858] The present invention provides a platform for effectively collecting, analyzing, and recommending word-of-mouth information via voice. The hardware and software configurations for realizing this system are described in detail below.

[0859] First, users use voice input devices such as smartphones to input their reviews of products and services via voice. This voice data is temporarily stored by the device. The voice input devices used here include smartphones, tablets, and voice recognition-enabled devices equipped with a standard microphone.

[0860] The device converts the audio data into a digital signal and sends it to the server via a secure protocol. The server uses the Google Cloud Speech-to-Text API to convert the audio data into text. This process transcribes the audio message into text, which is then available for use in the next parsing step.

[0861] The converted text information is analyzed using a sentiment analysis method based on TensorFlow and classified as positive, negative, or neutral. This analysis result is then used to provide personalized information based on the product's characteristics.

[0862] Furthermore, a truthfulness evaluation tool compares the word-of-mouth information with past user data and reliable information sources to assess its reliability. Using the Django framework, a summary generation tool extracts key information based on this evaluation data and provides it in a format that is easily understandable to the user.

[0863] The recommendation system suggests relevant products based on the user's emotional tendencies and past preference data. This information is then presented to the user via the display of a smartphone or tablet.

[0864] For example, if a user enters "The image quality and ease of use of this camera are excellent," the system will interpret this review as positive and recommend similar high-quality cameras.

[0865] Examples of prompt statements to input into the generating AI model are as follows:

[0866] "Please provide an audio review of the product you purchased. Based on your review, we will share the product's features and reasons for recommending it with other potential customers."

[0867] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0868] Step 1:

[0869] Users speak their reviews using a voice input device. The input is an analog voice signal, which the device converts into a digital signal for speech recognition. The digitized voice data is then transmitted to the server via a secure protocol.

[0870] Step 2:

[0871] The server sends the received digital audio data to the Google Cloud Speech-to-Text API to convert it into text. The input is audio data, which is analyzed by an algorithm, and the output is text data.

[0872] Step 3:

[0873] The server processes text data using a sentiment analysis tool powered by TensorFlow. The input is text data, which is classified as positive, negative, or neutral based on a sentiment classification model. The output is data indicating the user's emotional state.

[0874] Step 4:

[0875] The server evaluates the reliability of text data using a truthfulness assessment tool. The input consists of text data and external reliability data, which are compared against past review history, and a reliability score is generated as output. This process improves the quality of information presented to the user.

[0876] Step 5:

[0877] The server summarizes text data using a summarization generation mechanism. The input consists of detailed text data and confidence scores, and the summarization algorithm extracts important information, resulting in simplified summary information as output.

[0878] Step 6:

[0879] The server recommends products based on user sentiment and preference data using recommendation methods. Inputs are past user data and current sentiment analysis results. Data mining techniques are used to select appropriate product candidates, and a customized list of recommended products is generated as output.

[0880] Step 7:

[0881] The terminal presents the user with recommended products and summary data. Input is data transmitted from the server, and the information is displayed on the screen, allowing the user to visually and audibly confirm the information as output.

[0882] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0883] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0884] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0885] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0886] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0887] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0888] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0889] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0890] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0891] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0892] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0893] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0894] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0895] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0896] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0897] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0898] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0899] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0900] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0901] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0902] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0903] The following is further disclosed regarding the embodiments described above.

[0904] (Claim 1)

[0905] Voice input method,

[0906] A speech recognition means for converting the aforementioned voice input into text data,

[0907] A sentiment analysis means for analyzing emotions based on the aforementioned text data,

[0908] A truthfulness evaluation method for evaluating the reliability of word-of-mouth based on the aforementioned text data and sentiment analysis results,

[0909] A summary generation means that summarizes the aforementioned text data and presents it to the user,

[0910] Output means for providing the user with the aforementioned evaluation and summary results,

[0911] A system that includes this.

[0912] (Claim 2)

[0913] The system according to claim 1, characterized in that the truthfulness evaluation means includes a function to calculate a confidence score by comparing it with an external reliability data source.

[0914] (Claim 3)

[0915] The system according to claim 1, characterized in that the summary generation means has a function to personalize word-of-mouth information based on the user's past preference information and recommend relevant information.

[0916] "Example 1"

[0917] (Claim 1)

[0918] A means for acquiring audio data via an audio input device,

[0919] A speech conversion means for converting the aforementioned speech data into text data,

[0920] An emotion analysis means for analyzing the aforementioned character data and identifying the emotional state,

[0921] A reliability evaluation method that evaluates the reliability of word-of-mouth by referring to past data and external information sources,

[0922] A summarization generation means that summarizes text information and generates content suitable for the user,

[0923] A display means for providing the user with the aforementioned evaluation content and summary results,

[0924] A system that includes this.

[0925] (Claim 2)

[0926] The system according to claim 1, characterized in that the reliability evaluation means includes a function for calculating evaluation indicators by referring to an external data source.

[0927] (Claim 3)

[0928] The system according to claim 1, characterized in that the summary generation means has a function to personalize information based on the user's past history information and recommend relevant information.

[0929] "Application Example 1"

[0930] (Claim 1)

[0931] Voice input method,

[0932] A speech recognition means for converting the aforementioned voice input into text data,

[0933] A sentiment analysis means for analyzing emotions based on the aforementioned text data,

[0934] A reliability evaluation method for evaluating the reliability of word-of-mouth based on the aforementioned text data and sentiment analysis results,

[0935] A summary generation means that summarizes the aforementioned text data and presents it to the user,

[0936] Output means for providing the aforementioned evaluation and summary results to the user,

[0937] A user interface means installed on a smart device that quickly processes voice reviews and optimizes the information provided,

[0938] A system that includes this.

[0939] (Claim 2)

[0940] The system according to claim 1, characterized in that the reliability evaluation means has a function to calculate a reliability score by comparing it with an external data source and has a function to generate information that supports the decision-making of buyers on smart devices.

[0941] (Claim 3)

[0942] The system according to claim 1, characterized in that the summary generation means has a function to personalize word-of-mouth information based on the user's past preference information, recommend related information, and provide information to assist in the user's purchase decision.

[0943] "Example 2 of combining an emotion engine"

[0944] (Claim 1)

[0945] Voice input method,

[0946] A speech recognition means for converting the aforementioned voice input into text data,

[0947] A sentiment analysis means for classifying and analyzing emotions based on the aforementioned text data,

[0948] A personalized information generation means that personalizes the information based on the aforementioned text data and the results of sentiment analysis,

[0949] A truthfulness evaluation method for evaluating the reliability of word-of-mouth based on the aforementioned text data and sentiment analysis results,

[0950] A summary generation means that summarizes the aforementioned text data and presents it to the user,

[0951] Output means for providing the user with the aforementioned evaluation and summary results,

[0952] A system that includes this.

[0953] (Claim 2)

[0954] The system according to claim 1, wherein the truthfulness evaluation means has a function to calculate a confidence score by comparing it with an external reliability data source.

[0955] (Claim 3)

[0956] The system according to claim 1, wherein the summary generation means has a function to personalize word-of-mouth information based on the user's past preference information and recommend relevant information.

[0957] "Application example 2 when combining with an emotional engine"

[0958] (Claim 1)

[0959] An input means for receiving voice,

[0960] A speech recognition means for converting the aforementioned speech into text information,

[0961] A means for analyzing emotions based on the aforementioned textual information,

[0962] A means for evaluating the reliability of information based on the aforementioned textual information and the results of sentiment analysis,

[0963] A summary generation means for summarizing the aforementioned textual information and presenting it to the user,

[0964] A presentation means for supplying the aforementioned evaluation and summary results to the user,

[0965] A recommendation method for recommending products based on the emotional tendencies of the aforementioned users,

[0966] A system that includes this.

[0967] (Claim 2)

[0968] The system according to claim 1, characterized in that the truthfulness evaluation means includes a function to calculate the reliability by comparing it with an external evaluation source.

[0969] (Claim 3)

[0970] The system according to claim 1, characterized in that the summary generation means has a function to personalize information based on the user's past preference information and suggest related products. [Explanation of symbols]

[0971] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Voice input method, A speech recognition means for converting the aforementioned voice input into text data, A sentiment analysis means for analyzing emotions based on the aforementioned text data, A truthfulness evaluation method for evaluating the reliability of word-of-mouth based on the aforementioned text data and sentiment analysis results, A summary generation means that summarizes the aforementioned text data and presents it to the user, Output means for providing the user with the aforementioned evaluation and summary results, A system that includes this.

2. The system according to claim 1, characterized in that the truthfulness evaluation means includes a function to calculate a confidence score by comparing it with an external reliability data source.

3. The system according to claim 1, characterized in that the summary generation means has a function to personalize word-of-mouth information based on the user's past preference information and recommend relevant information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A