system
The system addresses the challenge of processing and discussing diverse perspectives from press conferences by converting audio to text, generating summaries, and facilitating comment-based discussions, enhancing information exchange and discourse quality.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Journalists and the general public face challenges in efficiently processing and discussing diverse perspectives from press conferences due to the overwhelming amount of information and lack of a high-quality two-way discussion platform.
A system that includes transcription, generation, and distribution means to convert audio data into text, automatically generate summaries and key points, and publish them on an information distribution platform with a comment function, enabling efficient information processing and multifaceted discussions.
Enables efficient information processing and promotes high-quality discourse by allowing journalists and citizens to exchange opinions effectively, transcending traditional media boundaries.
Smart Images

Figure 2026073429000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the modern reporting environment, journalists are dealing with a huge amount of information, and the load of dealing with various press conferences is increasing. In addition, the general public is also in a situation where it is difficult to obtain diverse perspectives and background information, and there is a lack of a high-quality two-way discussion platform. Against this background, it is necessary to realize a platform that can efficiently and effectively summarize information, provide important perspectives, and enable journalists and citizens to exchange opinions with each other.
Means for Solving the Problems
[0005] This invention includes a transcription means for receiving audio data and converting that data into text. Furthermore, it includes a generation means for automatically generating a summary and key points based on this text. The generated information is published on an information distribution platform, and users can exchange opinions through a comment function. This system allows journalists to process information efficiently and citizens to participate in multifaceted discussions.
[0006] "Receiving means" refers to a device or function for acquiring audio data from an external source.
[0007] "Transcription method" refers to a device or function that analyzes audio data and converts it into text information.
[0008] "Generation means" refers to a device or function that automatically generates summaries and key perspectives based on text data.
[0009] "Distribution means" refers to a device or structure that has the function of reposting generated information to an information distribution platform.
[0010] "User interface means" refers to a device or function that provides an interface for a user to directly interact with, input, or view information.
[0011] The "comment function" is a feature that allows users to input their opinions and views and share information with other users. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4]It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] [[ID=2�]]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] The system according to the present invention provides a specific implementation method for efficiently processing audio data and distributing information from press conferences and the like. A specific embodiment is described below.
[0034] This system consists of a server, terminals (users' PCs and smartphones), and users, all connected via a network. The server has the function of receiving audio data from an external source and transcribing it. The received audio data is automatically converted into text data using a speech recognition algorithm. This makes it easy for journalists and citizens who are not physically present at the press conference to understand the content.
[0035] A server equipped with generation AI takes text data as input and generates summaries and key points. The generated summaries are automatically published on an information distribution platform. The platform comes standard with a comment function, through which users (journalists and the general public) can express their opinions and reply to comments from other users. This function allows journalists to organize information efficiently and citizens to be exposed to diverse perspectives.
[0036] As a concrete example, if a political press conference is held, the server immediately receives the audio data and simultaneously begins transcribing it into text. The generation AI creates a summary from the text in real time, and that summary is quickly published. Users can comment on the summary, thereby promoting multifaceted discussion.
[0037] This system provides a platform for information sharing and discussion that transcends the boundaries between mass media and online media, enabling high-quality discourse.
[0038] The following describes the processing flow.
[0039] Step 1:
[0040] The server receives audio data from press conferences from an external source. Even when the audio arrives in real time, the buffering function is used to ensure stable data reception.
[0041] Step 2:
[0042] The server uses a speech recognition engine to convert the received audio data into text data. This records the content of the audio as text information.
[0043] Step 3:
[0044] The server inputs the transcribed text into the generating AI. The generating AI then summarizes the text and automatically generates key points based on a specified algorithm.
[0045] Step 4:
[0046] The server formats the generated summaries and perspectives and publishes them on the information distribution platform. The platform makes them accessible to users online.
[0047] Step 5:
[0048] The user's device (PC or smartphone) accesses the platform and views the published summaries and perspectives. The user accesses the information through their device.
[0049] Step 6:
[0050] Users can express their opinions through the comment function. Journalists can provide supplementary information based on the information they have gathered, and ordinary citizens can leave their own perceptions and questions as comments.
[0051] Step 7:
[0052] The server manages the exchange of comments between users and filters or moderates them as needed, thereby maintaining a healthy environment for discussion.
[0053] (Example 1)
[0054] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0055] In today's information society, there is a need to quickly and efficiently grasp audio information and accurately share necessary information. However, journalists and citizens who are not physically present at the scene face the challenge of not being able to easily acquire audio information and grasp the important points. Furthermore, there is a lack of effective means to promote discussion from diverse perspectives.
[0056] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0057] In this invention, the server includes receiving means for receiving acoustic information, transcription means for converting it into a written representation, and generation means for generating summaries and key points. This makes it possible for users who are not physically present at the site to quickly and accurately grasp the audio information and engage in discussions from diverse perspectives.
[0058] "Acoustic information" refers to all data acquired as sound, and is typically a signal acquired through devices such as microphones.
[0059] A "receiving means" is an element that has the function of receiving acoustic information from an external source and transferring it to the system for processing.
[0060] A "transcription means" is an element that performs the process of converting acquired acoustic information into a string of characters and generating a written representation.
[0061] "Written representation" refers to the representation of acoustic information in text format, visually showing the content of audio data.
[0062] "Generative means" refers to elements that perform processing to create a summary based on written expression and extract important information.
[0063] An "information distribution platform" is a system that provides users with generated summaries and important information, and enables interaction.
[0064] The "opinion function" is a feature that provides an interface for users to add their own opinions and comments to generated information and to communicate with others.
[0065] "User operation means" refers to the means used by users to view information or add opinions on the information distribution platform.
[0066] The "response function" is a feature that allows users to reply to comments in order to facilitate communication between them.
[0067] The "evaluation function" is a feature that allows other users to evaluate the quality of other users' opinions and comments, thereby promoting more active discussion.
[0068] This system achieves efficient processing and distribution of acoustic information by automating multiple processes. The system consists of servers, terminals (user computers and mobile devices), and users, all connected via a network.
[0069] The server first receives acoustic information. This acoustic information is typically audio data sent over the internet. Secure data transfer is ensured using a secure protocol. The received acoustic information is then converted into written form using speech recognition software. Specific software used for this process includes commonly available speech recognition APIs.
[0070] Next, the server uses a generative AI model to generate a summary and key points from the written text. The AI models used here include models widely known in the field of natural language processing. Specifically, the generative AI model is run by inputting a prompt such as "Please summarize this text and list the five key points."
[0071] The generated summaries and key information are immediately published through the information distribution platform. This platform is designed to update information in real time over the internet and enables two-way communication between users and servers.
[0072] Users of the device can view published summaries and key points, and express their own opinions on the platform through the opinion function. This enables discussions from diverse perspectives and stimulates information exchange among users.
[0073] As a concrete example, this system can receive audio information from a public press conference, input written statements in real time using pre-prepared prompts, generate summaries, and efficiently process the entire process up to publication. Thus, the ability to quickly access high-quality information regardless of location is a key feature of this system.
[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0075] Step 1:
[0076] The server receives audio data over the network. It takes audio data as input and transfers it to the server using a secure protocol, ensuring data integrity and confidentiality. The output is the storage of audio data, which forms the basis for subsequent processing.
[0077] Step 2:
[0078] The server transcribes the received audio data using speech recognition software. The input is audio data, and the speech recognition engine analyzes the audio signal and converts it into text data. Specific operations include the extraction of acoustic features and language analysis using a probabilistic model. The output is text data representing the content of the audio in written form.
[0079] Step 3:
[0080] The server generates a summary and key points from transcribed text data using a generative AI model. The input is text data, and the generative model is given the prompt "Summarize this text and list five key points" for processing. Specifically, it uses natural language processing techniques to understand the overall structure of the text and extract important context. The output is a summary and a list of key points.
[0081] Step 4:
[0082] The server publishes the generated summary and key points to the information distribution platform. The input consists of the summary and key points, which are formatted into an appropriate format and uploaded to the information distribution platform. Specifically, this involves format conversion and data transmission to the web server, ensuring real-time publication. The output is text information made public to users.
[0083] Step 5:
[0084] Users using the terminal can access publicly available information and add their opinions. Input consists of information on the distribution platform and user comments. Users enter comments, which the system saves to a database and displays to other users. Output is updated content, enabling the exchange of opinions from diverse perspectives.
[0085] (Application Example 1)
[0086] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0087] Traditional information distribution services face challenges in the efficiency of their information gathering and distribution processes, particularly regarding press conferences and other events. They lack features for quickly summarizing important information and facilitating the exchange of opinions from multiple perspectives. To address these challenges, there is a need to provide a system that facilitates real-time information conversion and sharing, as well as smooth dialogue among users.
[0088] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0089] In this invention, the server includes an acquisition means for acquiring information, a conversion means for converting the information into a string, and a generation means for inputting the string and generating a summary and key points. This enables efficient information collection and distribution, as well as smooth exchange of opinions among users.
[0090] "Acquisition means" refers to a mechanism for receiving information from an external source and incorporating it into the system.
[0091] A "conversion method" is a process for analyzing acquired information and converting it into a string in a useful format.
[0092] The "generation method" refers to a function that uses the converted string to summarize information and automatically extract key perspectives.
[0093] A "distribution method" is a mechanism for publishing the generated summaries and perspectives on an information distribution platform and sharing them with a large number of users.
[0094] A "user interface means" is an interface that allows users to add opinions through the system and exchange opinions with others.
[0095] The "opinion exchange function" is a feature that allows users to post their own opinions and engage in dialogue with the opinions of other users.
[0096] The "response function" is a feature that allows multiple users to reply to or comment on opinions posted by other users.
[0097] The "rating function" is a function that provides a means to evaluate the quality and usefulness of opinions posted by users.
[0098] This invention is an information processing system comprising acquisition means, conversion means, generation means, distribution means, and user interface means.
[0099] The server receives audio data from press conferences, events, and other sources using information acquisition methods. The received audio data is converted into text data by conversion methods. This process utilizes speech recognition software such as Google® Cloud Speech-to-Text.
[0100] Text data is summarized using a generation AI model (e.g., OpenAI® GPT model) by a generation method, and key perspectives are extracted. This summarized data is then published in real time on an information distribution platform by a distribution method.
[0101] Users can access the user interface using their devices (smartphones or PCs) and add comments and exchange opinions with other users through the provided opinion exchange function. This utilizes a real-time database such as Firebase.
[0102] As a concrete example, when a user opens a news app to check the breaking news of a politician's press conference, they can immediately view summarized information, see comments from other users, and share their own opinions. An example of a prompt for the generative AI model could be an instruction such as, "Summarize the following press conference text and identify three key points."
[0103] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0104] Step 1:
[0105] The server receives audio data from the information source. The input here is either real-time or recorded audio data. The server acquires this data via a network connection and sends it to the conversion means.
[0106] Step 2:
[0107] The server uses a conversion mechanism to convert received audio data into text data. The input is audio data, and speech recognition software such as Google Cloud Speech-to-Text is used to output text data. Feature extraction from the audio data and transcription from the audio are performed using deep learning.
[0108] Step 3:
[0109] The server utilizes generation methods to summarize text data and extract key perspectives. The input is transformed text data, and the output is a summary and perspective information. Generative AI models such as the OpenAI GPT model are used. Text analysis and natural language processing techniques are employed to generate the summary and perspectives.
[0110] Step 4:
[0111] The server publishes the generated summaries and perspectives to the information distribution platform via a distribution method. The input is the generated summary data, and the output is the online distribution content. Cloud storage or databases are used for publication.
[0112] Step 5:
[0113] Users access the information distribution platform using their terminals and utilize the provided opinion exchange function. Input consists of distributed information and user comments, while output is the user's opinion. Interactive opinion sharing is achieved through the user interface.
[0114] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0115] The system according to the present invention provides specific technologies for considering user emotions in information provision and achieving more effective communication.
[0116] In one embodiment of the invention, this system consists of a server, a terminal (the user's PC or smartphone), and the user. In addition to conventional voice data reception and transcription functions, the server has an emotion engine. This emotion engine analyzes the user's input comment text and recognizes the user's underlying emotions. The recognized emotions are then used in subsequent interactions with the user.
[0117] As a concrete example, this sentiment engine is activated when a user reads an article on an information distribution platform and posts a comment. If a user posts a stressful comment, the server can analyze it and generate an interface that soothes the user's emotions by providing additional information and diagrams to help them relax. The server also uses a sentiment-based feedback system to encourage constructive discussion and facilitate the exchange of opinions among multiple users.
[0118] Furthermore, the device displays pre-adjusted content sent from the server, allowing users to receive information tailored to their emotions. In this way, the emotion engine understands the user's feelings and adjusts the content accordingly, enabling users to enjoy a more satisfying experience. This system aims to create high-quality information delivery and communication by taking emotions into consideration.
[0119] The following describes the processing flow.
[0120] Step 1:
[0121] The server receives audio data from the press conference from an external source. The audio files are stored for processing according to the format.
[0122] Step 2:
[0123] The server uses a speech recognition engine to convert the received audio data into text data. This conversion transcribes the content.
[0124] Step 3:
[0125] The server passes the transcribed text data to a generation AI model, which generates a summary and key points. The generated information is then used for subsequent processing.
[0126] Step 4:
[0127] The server publishes information so that users can view it on the platform. Additionally, a sentiment engine monitors user interactions.
[0128] Step 5:
[0129] The terminal displays published summaries and perspectives. An interface is provided for users to view the information and enter comments.
[0130] Step 6:
[0131] When a user enters and submits a comment, the server analyzes the comment using a sentiment engine. This analysis identifies the sentiment contained in the comment.
[0132] Step 7:
[0133] The server adjusts the information presentation based on the sentiment analysis results. Additional information and feedback adapted to the user's emotions are generated and provided to the user through the user interface.
[0134] Step 8:
[0135] Users receive responses from the system based on their own reactions, allowing them to provide context-appropriate feedback and additional interactions. This process is repeated to continuously enhance the user experience.
[0136] (Example 2)
[0137] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0138] Conventional information distribution systems provide information uniformly without considering users' emotions, resulting in a lack of emotionally-driven interaction and challenges in information acceptance and satisfaction. Furthermore, they lacked appropriate functions to facilitate constructive exchange of opinions among users, leading to insufficient communication quality.
[0139] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0140] In this invention, the server includes receiving means for receiving voice information, conversion means for converting it into text information, and emotion recognition means for analyzing the text information to identify emotions. This makes it possible to design interactions that respond to the user's emotions, and further enables the provision of emotion-based, tailored information and the promotion of constructive exchange of opinions among users.
[0141] "Audio information" refers to data that represents spoken language and other audio signals in digital or analog format.
[0142] "Receiving means" refers to a device or program for receiving information from an external source, and in this case, it refers to the function of receiving audio information.
[0143] "Textual information" refers to data obtained by converting audio information into text format, which is interpreted as sentences or words.
[0144] "Conversion means" refers to a device or program for converting data in one format to another, and in this case, it refers to the function of converting audio information into text information.
[0145] "Emotion recognition means" refers to a technology or program that analyzes textual information and identifies the user's emotions from it.
[0146] "Interaction design" is the process of dynamically adjusting the content and user interface provided by a system in response to the user's emotions and needs.
[0147] An "information sharing medium" is a platform or network for providing and sharing information with a wide range of users.
[0148] "Interactive features" refer to functions that enable users to input information into the system and exchange information with other users.
[0149] A "response function" is a function that provides appropriate responses or feedback to received information.
[0150] A "feedback function" is a feature that incorporates user input and comments into the system or service.
[0151] This system is comprised of a server, terminals (e.g., the user's personal computer or mobile device), and the user themselves. The server, at the core of the invention, is equipped with a speech-to-text engine for receiving speech information and converting it into text information. Specifically, speech information is transmitted to the server via a network and converted into text data using the conversion engine. General-purpose speech recognition software is used for this engine. The core operation of this system is described below.
[0152] After the server converts the audio information to text, it uses an emotion recognition engine to identify the emotions. This engine analyzes the text using a natural language processing library to identify the emotions contained within it. The results of the emotion analysis are then used to design subsequent interactions. For example, the server adjusts the information and content it provides according to specific emotions (e.g., stress or joy). To this end, a content management system works in conjunction to dynamically generate and provide content.
[0153] The terminal is responsible for displaying the adjusted content sent from the server to the user. The displayed content is visually formatted using HTML and CSS. This allows users to receive information that aligns with their emotions, enabling them to enjoy a more satisfying user experience.
[0154] As a concrete example, consider a scenario where a user reads a news article and enters a comment on it. In this case, the entered comment is analyzed by the server, and if the emotion is determined to be "anger," the server provides content that addresses that emotion. This may include information on relaxation techniques or detailed illustrations on a specific topic.
[0155] An example of a prompt for a generative AI model is, "Please show a method for identifying emotions from user comments and providing appropriate feedback." This prompt allows the AI model to provide insights into emotion analysis and appropriate content generation.
[0156] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0157] Step 1:
[0158] The server receives voice information from the user. A voice data file is passed as input, and the server uses a speech recognition engine to analyze the data. This converts the voice into text information, generating text data. The output is text information.
[0159] Step 2:
[0160] The server inputs the generated text information into the sentiment recognition engine. Here, natural language processing algorithms are used to analyze keywords and context within the text and identify the user's emotions. During this sentiment analysis process, the text is assigned a sentiment score, such as positive, negative, or neutral. The output is the analysis result, including the sentiment score.
[0161] Step 3:
[0162] The server adjusts the content it provides based on the results of sentiment analysis. Sentiment scores are used as input, and a content management system is utilized to select information corresponding to specific emotions. For example, if the user is identified as stressed, relaxation information and specific visual content will be selected. The adjusted content is then generated as output.
[0163] Step 4:
[0164] The device receives the formatted content from the server and displays it to the user. The input is HTML and CSS code sent from the server, which is then rendered in a web browser. The device provides a visual interface through the browser, allowing the user to view the content and send feedback as needed. The output is the provision of visual information to the user.
[0165] Step 5:
[0166] Users enter feedback and new comments through the displayed interface. The input is sent to the server as text data, which then re-analyzes the data to initiate a new analysis cycle. This facilitates constructive exchange of ideas among users and deepens the discussion. Further refined content is provided as output from the re-analysis.
[0167] (Application Example 2)
[0168] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0169] In today's information delivery landscape, users offer comments and feedback with a wide range of emotions, but the lack of information delivery that adequately considers these emotions results in insufficient improvement of the user experience. To address this issue, there is a need for a means to accurately capture users' emotions and provide appropriate information that responds to those emotions.
[0170] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0171] In this invention, the server includes an acquisition means for receiving audio data, a conversion means for converting the audio data into text information, and an emotion recognition means for inputting the text information and analyzing emotions. This enables the provision of information and recommendation of appropriate content based on the user's emotions.
[0172] "Audio data" refers to data recorded in digital format, which is fundamental information used to analyze or process audio.
[0173] "Acquisition means" refers to functions or devices for receiving or collecting audio data.
[0174] A "conversion means" is a mechanism for processing received audio data into text information.
[0175] "Textual information" refers to information that is expressed as text based on audio data, and is written in text format.
[0176] "Emotion recognition means" refers to algorithms and technologies for analyzing and identifying a user's emotions from text information.
[0177] A "generation means" is a device or process that has the function of generating related information or content based on analyzed emotional information.
[0178] "Information provision means" refers to the means of presenting generated information to the user, and includes user interfaces and platforms.
[0179] "User interface means" refers to an interface that allows users to receive and manipulate information, such as a device screen or application interface.
[0180] To realize this invention, a server plays a central role. The server first has a receiving means for acquiring audio data and utilizes a conversion means to convert this audio data into text information. By using a widely used solution as speech recognition software, such as the Google Cloud Speech-to-Text API, it is possible to convert speech to text with high accuracy.
[0181] Next, the server applies emotion recognition tools to analyze the emotions based on the converted text information. In this process, the Google Cloud Natural Language API is used as the software tool for emotion analysis. This allows the server to determine positive, negative, or neutral emotions from user comments and feedback.
[0182] Based on the analyzed sentiment information, the server generates relevant information and content tailored to the user's emotions through a generation mechanism. This generation includes summarizing the transcribed content and algorithms to provide information best suited to the user's current emotions. The generated information is presented to the user through an information delivery mechanism. The user interface operates on devices such as smartphones and tablets, displaying the information in a format that the user can view and interact with.
[0183] As a concrete example, suppose a user watches a movie review video and comments, "This ending is confusing." In this case, the server's emotion engine identifies the emotion of "confusion," generates content recommending similar movie commentary articles or videos, and displays it in the user interface.
[0184] An example of a prompt to input into a generative AI model is as follows: "A user posted the comment 'I didn't understand the ending of this movie.' Analyze the sentiment of this comment and suggest appropriate content."
[0185] This allows users to receive information that takes their feelings into consideration, enabling them to enjoy a better experience.
[0186] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0187] Step 1:
[0188] The server receives audio data. Using the received audio data as input, it performs speech recognition using a conversion method and converts it into text information. By utilizing the Google Cloud Speech-to-Text API, it outputs highly accurate transcription results. This is for use as basic data for information provision.
[0189] Step 2:
[0190] The server passes text information as input to an emotion recognition system to analyze the user's emotions. This process utilizes the Google Cloud Natural Language API to extract emotions from various parts of the text. Specifically, it generates emotional labels such as positive, negative, and neutral, and outputs the analysis results. This information is used to generate content optimized for each individual user.
[0191] Step 3:
[0192] The server generates relevant content using a generation mechanism based on the analyzed emotional information. The generated content must be in harmony with the user's emotions. For example, if negative emotions are analyzed, relaxing information or explanations will be selected. This allows for the output of customized information tailored to the user.
[0193] Step 4:
[0194] The server transmits the generated information to the terminal through an information delivery system. The terminal displays this information via a user interface. Users can receive and view information that corresponds to their own emotions. This can improve the user experience.
[0195] Step 5:
[0196] After receiving information that resonates with their emotions, users can add further feedback and comments as needed. This allows for even more optimized content suggestions in the future.
[0197] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0198] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0199] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0200] [Second Embodiment]
[0201] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0202] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0203] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0204] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0205] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0206] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0207] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0208] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0209] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0210] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0211] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0212] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0213] The system according to the present invention provides a specific implementation method for efficiently processing audio data and distributing information from press conferences and the like. A specific embodiment is described below.
[0214] This system consists of a server, terminals (users' PCs and smartphones), and users, all connected via a network. The server has the function of receiving audio data from an external source and transcribing it. The received audio data is automatically converted into text data using a speech recognition algorithm. This makes it easy for journalists and citizens who are not physically present at the press conference to understand the content.
[0215] A server equipped with generation AI takes text data as input and generates summaries and key points. The generated summaries are automatically published on an information distribution platform. The platform comes standard with a comment function, through which users (journalists and the general public) can express their opinions and reply to comments from other users. This function allows journalists to organize information efficiently and citizens to be exposed to diverse perspectives.
[0216] As a concrete example, if a political press conference is held, the server immediately receives the audio data and simultaneously begins transcribing it into text. The generation AI creates a summary from the text in real time, and that summary is quickly published. Users can comment on the summary, thereby promoting multifaceted discussion.
[0217] This system provides a platform for information sharing and discussion that transcends the boundaries between mass media and online media, enabling high-quality discourse.
[0218] The following describes the processing flow.
[0219] Step 1:
[0220] The server receives audio data from press conferences from an external source. Even when the audio arrives in real time, the buffering function is used to ensure stable data reception.
[0221] Step 2:
[0222] The server uses a speech recognition engine to convert the received audio data into text data. This records the content of the audio as text information.
[0223] Step 3:
[0224] The server inputs the transcribed text into the generating AI. The generating AI then summarizes the text and automatically generates key points based on a specified algorithm.
[0225] Step 4:
[0226] The server formats the generated summaries and perspectives and publishes them on the information distribution platform. The platform makes them accessible to users online.
[0227] Step 5:
[0228] The user's device (PC or smartphone) accesses the platform and views the published summaries and perspectives. The user accesses the information through their device.
[0229] Step 6:
[0230] Users can express their opinions through the comment function. Journalists can provide supplementary information based on the information they have gathered, and ordinary citizens can leave their own perceptions and questions as comments.
[0231] Step 7:
[0232] The server manages the exchange of comments between users and filters or moderates them as needed, thereby maintaining a healthy environment for discussion.
[0233] (Example 1)
[0234] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0235] In today's information society, there is a need to quickly and efficiently grasp audio information and accurately share necessary information. However, journalists and citizens who are not physically present at the scene face the challenge of not being able to easily acquire audio information and grasp the important points. Furthermore, there is a lack of effective means to promote discussion from diverse perspectives.
[0236] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0237] In this invention, the server includes receiving means for receiving acoustic information, transcription means for converting it into a written representation, and generation means for generating summaries and key points. This makes it possible for users who are not physically present at the site to quickly and accurately grasp the audio information and engage in discussions from diverse perspectives.
[0238] "Acoustic information" refers to all data acquired as sound, and is typically a signal acquired through devices such as microphones.
[0239] A "receiving means" is an element that has the function of receiving acoustic information from an external source and transferring it to the system for processing.
[0240] A "transcription means" is an element that performs the process of converting acquired acoustic information into a string of characters and generating a written representation.
[0241] "Written representation" refers to the representation of acoustic information in text format, visually showing the content of audio data.
[0242] "Generative means" refers to elements that perform processing to create a summary based on written expression and extract important information.
[0243] An "information distribution platform" is a system that provides users with generated summaries and important information, and enables interaction.
[0244] The "opinion function" is a feature that provides an interface for users to add their own opinions and comments to generated information and to communicate with others.
[0245] "User operation means" refers to the means used by users to view information or add opinions on the information distribution platform.
[0246] The "response function" is a feature that allows users to reply to comments in order to facilitate communication between them.
[0247] The "evaluation function" is a feature that allows other users to evaluate the quality of other users' opinions and comments, thereby promoting more active discussion.
[0248] This system achieves efficient processing and distribution of acoustic information by automating multiple processes. The system consists of servers, terminals (user computers and mobile devices), and users, all connected via a network.
[0249] The server first receives acoustic information. This acoustic information is typically audio data sent over the internet. Secure data transfer is ensured using a secure protocol. The received acoustic information is then converted into written form using speech recognition software. Specific software used for this process includes commonly available speech recognition APIs.
[0250] Next, the server uses a generative AI model to generate a summary and key points from the written text. The AI models used here include models widely known in the field of natural language processing. Specifically, the generative AI model is run by inputting a prompt such as "Please summarize this text and list the five key points."
[0251] The generated summaries and key information are immediately published through the information distribution platform. This platform is designed to update information in real time over the internet and enables two-way communication between users and servers.
[0252] Users of the device can view published summaries and key points, and express their own opinions on the platform through the opinion function. This enables discussions from diverse perspectives and stimulates information exchange among users.
[0253] As a concrete example, this system can receive audio information from a public press conference, input written statements in real time using pre-prepared prompts, generate summaries, and efficiently process the entire process up to publication. Thus, the ability to quickly access high-quality information regardless of location is a key feature of this system.
[0254] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0255] Step 1:
[0256] The server receives audio data over the network. It takes audio data as input and transfers it to the server using a secure protocol, ensuring data integrity and confidentiality. The output is the storage of audio data, which forms the basis for subsequent processing.
[0257] Step 2:
[0258] The server transcribes the received audio data using speech recognition software. The input is audio data, and the speech recognition engine analyzes the audio signal and converts it into text data. Specific operations include the extraction of acoustic features and language analysis using a probabilistic model. The output is text data representing the content of the audio in written form.
[0259] Step 3:
[0260] The server generates a summary and key points from transcribed text data using a generative AI model. The input is text data, and the generative model is given the prompt "Summarize this text and list five key points" for processing. Specifically, it uses natural language processing techniques to understand the overall structure of the text and extract important context. The output is a summary and a list of key points.
[0261] Step 4:
[0262] The server publishes the generated summary and key points to the information distribution platform. The input consists of the summary and key points, which are formatted into an appropriate format and uploaded to the information distribution platform. Specifically, this involves format conversion and data transmission to the web server, ensuring real-time publication. The output is text information made public to users.
[0263] Step 5:
[0264] Users using the terminal can access publicly available information and add their opinions. Input consists of information on the distribution platform and user comments. Users enter comments, which the system saves to a database and displays to other users. Output is updated content, enabling the exchange of opinions from diverse perspectives.
[0265] (Application Example 1)
[0266] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0267] Traditional information distribution services face challenges in the efficiency of their information gathering and distribution processes, particularly regarding press conferences and other events. They lack features for quickly summarizing important information and facilitating the exchange of opinions from multiple perspectives. To address these challenges, there is a need to provide a system that facilitates real-time information conversion and sharing, as well as smooth dialogue among users.
[0268] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0269] In this invention, the server includes an acquisition means for acquiring information, a conversion means for converting the information into a string, and a generation means for inputting the string and generating a summary and key points. This enables efficient information collection and distribution, as well as smooth exchange of opinions among users.
[0270] "Acquisition means" refers to a mechanism for receiving information from an external source and incorporating it into the system.
[0271] A "conversion method" is a process for analyzing acquired information and converting it into a string in a useful format.
[0272] The "generation method" refers to a function that uses the converted string to summarize information and automatically extract key perspectives.
[0273] A "distribution method" is a mechanism for publishing the generated summaries and perspectives on an information distribution platform and sharing them with a large number of users.
[0274] A "user interface means" is an interface that allows users to add opinions through the system and exchange opinions with others.
[0275] The "opinion exchange function" is a feature that allows users to post their own opinions and engage in dialogue with the opinions of other users.
[0276] The "response function" is a feature that allows multiple users to reply to or comment on opinions posted by other users.
[0277] The "rating function" is a function that provides a means to evaluate the quality and usefulness of opinions posted by users.
[0278] This invention is an information processing system composed of an acquisition means, a conversion means, a generation means, a distribution means, and a user interface means.
[0279] The server receives voice data such as press conferences and events using the information acquisition means. The received voice data is converted into text data by the conversion means. In this process, speech recognition software such as Google Cloud Speech-to-Text is used.
[0280] The text data is summarized by the generation means using a generation AI model (for example, the OpenAI GPT model), and important viewpoints are extracted. This summary data is publicly released to the information distribution infrastructure in real time by the distribution means.
[0281] Users can access the user interface means using a terminal (smartphone or PC) and add comments or exchange opinions with other users through the provided opinion exchange function. For this, a real-time database such as Firebase is used.
[0282] As a specific example, when a user opens a news app and checks the flash report of a politician's press conference, they can immediately view the summarized information and the comments of other users and also share their own opinions. Also, as an example of a prompt sentence for the generation AI model, an instruction such as "Summarize the following meeting text and list three important viewpoints." can be considered.
[0283] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0284] Step 1:
[0285] The server receives voice data from the information source. The input here is real-time or recorded voice data. The server acquires this data via a network connection and sends it to the conversion means.
[0286] Step 2:
[0287] The server uses conversion means to convert the received voice data into text data. The input is voice data, and string data is output using speech recognition software such as Google Cloud Speech-to-Text. Feature extraction of the voice data and character transcription from the voice by deep learning are performed.
[0288] Step 3:
[0289] The server utilizes generation means to summarize the text data and extract important viewpoints. The input is the converted text data, and the output is summary and viewpoint information. Generation AI models such as the OpenAI GPT model are used. Summarization and viewpoint generation are performed through text analysis and natural language processing techniques.
[0290] Step 4:
[0291] The server publishes the generated summary and viewpoints to the information distribution infrastructure via distribution means. The input is the generated summary data, and the output is online distribution content. Cloud storage or a database is used for publication.
[0292] Step 5:
[0293] The user uses the terminal to access the information distribution infrastructure and uses the provided opinion exchange function. The input is the distributed information and the user's comments, and the output is the user's opinion. Interactive opinion sharing is realized through user interface means.
[0294] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion specific model 59 and perform specific processing using the user's emotion.
[0295] The system according to the present invention provides specific technologies for considering user emotions in information provision and achieving more effective communication.
[0296] In one embodiment of the invention, this system consists of a server, a terminal (the user's PC or smartphone), and the user. In addition to conventional voice data reception and transcription functions, the server has an emotion engine. This emotion engine analyzes the user's input comment text and recognizes the user's underlying emotions. The recognized emotions are then used in subsequent interactions with the user.
[0297] As a concrete example, this sentiment engine is activated when a user reads an article on an information distribution platform and posts a comment. If a user posts a stressful comment, the server can analyze it and generate an interface that soothes the user's emotions by providing additional information and diagrams to help them relax. The server also uses a sentiment-based feedback system to encourage constructive discussion and facilitate the exchange of opinions among multiple users.
[0298] Furthermore, the device displays pre-adjusted content sent from the server, allowing users to receive information tailored to their emotions. In this way, the emotion engine understands the user's feelings and adjusts the content accordingly, enabling users to enjoy a more satisfying experience. This system aims to create high-quality information delivery and communication by taking emotions into consideration.
[0299] The following describes the processing flow.
[0300] Step 1:
[0301] The server receives audio data from the press conference from an external source. The audio files are stored for processing according to the format.
[0302] Step 2:
[0303] The server uses an audio recognition engine to convert the received audio data into text data. Through this conversion, the content is transcribed.
[0304] Step 3:
[0305] The server passes the transcribed text data to a generative AI model to generate a summary and key points. The generated information is used for subsequent processing.
[0306] Step 4:
[0307] The server publishes the information so that the user can view it on the platform. Also, the user's interactions are monitored by an emotion engine.
[0308] Step 5:
[0309] The terminal displays the published summary and key points. An interface is provided for the user to view the information and enter comments.
[0310] Step 6:
[0311] When the user enters and sends a comment, the server analyzes the comment with an emotion engine. Through this analysis, the emotion contained in the comment is identified.
[0312] Step 7:
[0313] Based on the emotion analysis results, the server adjusts the information presentation. Additional information and feedback adapted to the user's emotion are generated and provided to the user via the user interface.
[0314] Step 8:
[0315] Users receive responses from the system based on their own reactions, allowing them to provide context-appropriate feedback and additional interactions. This process is repeated to continuously enhance the user experience.
[0316] (Example 2)
[0317] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0318] Conventional information distribution systems provide information uniformly without considering users' emotions, resulting in a lack of emotionally-driven interaction and challenges in information acceptance and satisfaction. Furthermore, they lacked appropriate functions to facilitate constructive exchange of opinions among users, leading to insufficient communication quality.
[0319] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0320] In this invention, the server includes receiving means for receiving voice information, conversion means for converting it into text information, and emotion recognition means for analyzing the text information to identify emotions. This makes it possible to design interactions that respond to the user's emotions, and further enables the provision of emotion-based, tailored information and the promotion of constructive exchange of opinions among users.
[0321] "Audio information" refers to data that represents spoken language and other audio signals in digital or analog format.
[0322] "Receiving means" refers to a device or program for receiving information from an external source, and in this case, it refers to the function of receiving audio information.
[0323] "Textual information" refers to data obtained by converting audio information into text format, which is interpreted as sentences or words.
[0324] "Conversion means" refers to a device or program for converting data in one format to another, and in this case, it refers to the function of converting audio information into text information.
[0325] "Emotion recognition means" refers to a technology or program that analyzes textual information and identifies the user's emotions from it.
[0326] "Interaction design" is the process of dynamically adjusting the content and user interface provided by a system in response to the user's emotions and needs.
[0327] An "information sharing medium" is a platform or network for providing and sharing information with a wide range of users.
[0328] "Interactive features" refer to functions that enable users to input information into the system and exchange information with other users.
[0329] A "response function" is a function that provides appropriate responses or feedback to received information.
[0330] A "feedback function" is a feature that incorporates user input and comments into the system or service.
[0331] This system is comprised of a server, terminals (e.g., the user's personal computer or mobile device), and the user themselves. The server, at the core of the invention, is equipped with a speech-to-text engine for receiving speech information and converting it into text information. Specifically, speech information is transmitted to the server via a network and converted into text data using the conversion engine. General-purpose speech recognition software is used for this engine. The core operation of this system is described below.
[0332] After the server converts the audio information to text, it uses an emotion recognition engine to identify the emotions. This engine analyzes the text using a natural language processing library to identify the emotions contained within it. The results of the emotion analysis are then used to design subsequent interactions. For example, the server adjusts the information and content it provides according to specific emotions (e.g., stress or joy). To this end, a content management system works in conjunction to dynamically generate and provide content.
[0333] The terminal is responsible for displaying the adjusted content sent from the server to the user. The displayed content is visually formatted using HTML and CSS. This allows users to receive information that aligns with their emotions, enabling them to enjoy a more satisfying user experience.
[0334] As a concrete example, consider a scenario where a user reads a news article and enters a comment on it. In this case, the entered comment is analyzed by the server, and if the emotion is determined to be "anger," the server provides content that addresses that emotion. This may include information on relaxation techniques or detailed illustrations on a specific topic.
[0335] An example of a prompt for a generative AI model is, "Please show a method for identifying emotions from user comments and providing appropriate feedback." This prompt allows the AI model to provide insights into emotion analysis and appropriate content generation.
[0336] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0337] Step 1:
[0338] The server receives voice information from the user. A voice data file is passed as input, and the server uses a speech recognition engine to analyze the data. This converts the voice into text information, generating text data. The output is text information.
[0339] Step 2:
[0340] The server inputs the generated text information into the sentiment recognition engine. Here, natural language processing algorithms are used to analyze keywords and context within the text and identify the user's emotions. During this sentiment analysis process, the text is assigned a sentiment score, such as positive, negative, or neutral. The output is the analysis result, including the sentiment score.
[0341] Step 3:
[0342] The server adjusts the content it provides based on the results of sentiment analysis. Sentiment scores are used as input, and a content management system is utilized to select information corresponding to specific emotions. For example, if the user is identified as stressed, relaxation information and specific visual content will be selected. The adjusted content is then generated as output.
[0343] Step 4:
[0344] The device receives the formatted content from the server and displays it to the user. The input is HTML and CSS code sent from the server, which is then rendered in a web browser. The device provides a visual interface through the browser, allowing the user to view the content and send feedback as needed. The output is the provision of visual information to the user.
[0345] Step 5:
[0346] Users enter feedback and new comments through the displayed interface. The input is sent to the server as text data, which then re-analyzes the data to initiate a new analysis cycle. This facilitates constructive exchange of ideas among users and deepens the discussion. Further refined content is provided as output from the re-analysis.
[0347] (Application Example 2)
[0348] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0349] In today's information delivery landscape, users offer comments and feedback with a wide range of emotions, but the lack of information delivery that adequately considers these emotions results in insufficient improvement of the user experience. To address this issue, there is a need for a means to accurately capture users' emotions and provide appropriate information that responds to those emotions.
[0350] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0351] In this invention, the server includes an acquisition means for receiving audio data, a conversion means for converting the audio data into text information, and an emotion recognition means for inputting the text information and analyzing emotions. This enables the provision of information and recommendation of appropriate content based on the user's emotions.
[0352] "Audio data" refers to data recorded in digital format, which is fundamental information used to analyze or process audio.
[0353] "Acquisition means" refers to functions or devices for receiving or collecting audio data.
[0354] A "conversion means" is a mechanism for processing received audio data into text information.
[0355] "Textual information" refers to information that is expressed as text based on audio data, and is written in text format.
[0356] "Emotion recognition means" refers to algorithms and technologies for analyzing and identifying a user's emotions from text information.
[0357] A "generation means" is a device or process that has the function of generating related information or content based on analyzed emotional information.
[0358] "Information provision means" refers to the means of presenting generated information to the user, and includes user interfaces and platforms.
[0359] "User interface means" refers to an interface that allows users to receive and manipulate information, such as a device screen or application interface.
[0360] To realize this invention, a server plays a central role. The server first has a receiving means for acquiring audio data and utilizes a conversion means to convert this audio data into text information. By using a widely used solution as speech recognition software, such as the Google Cloud Speech-to-Text API, it is possible to convert speech to text with high accuracy.
[0361] Next, the server applies emotion recognition tools to analyze the emotions based on the converted text information. In this process, the Google Cloud Natural Language API is used as the software tool for emotion analysis. This allows the server to determine positive, negative, or neutral emotions from user comments and feedback.
[0362] Based on the analyzed sentiment information, the server generates relevant information and content tailored to the user's emotions through a generation mechanism. This generation includes summarizing the transcribed content and algorithms to provide information best suited to the user's current emotions. The generated information is presented to the user through an information delivery mechanism. The user interface operates on devices such as smartphones and tablets, displaying the information in a format that the user can view and interact with.
[0363] As a concrete example, suppose a user watches a movie review video and comments, "This ending is confusing." In this case, the server's emotion engine identifies the emotion of "confusion," generates content recommending similar movie commentary articles or videos, and displays it in the user interface.
[0364] An example of a prompt to input into a generative AI model is as follows: "A user posted the comment 'I didn't understand the ending of this movie.' Analyze the sentiment of this comment and suggest appropriate content."
[0365] This allows users to receive information that takes their feelings into consideration, enabling them to enjoy a better experience.
[0366] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0367] Step 1:
[0368] The server receives audio data. Using the received audio data as input, it performs speech recognition using a conversion method and converts it into text information. By utilizing the Google Cloud Speech-to-Text API, it outputs highly accurate transcription results. This is for use as basic data for information provision.
[0369] Step 2:
[0370] The server passes text information as input to an emotion recognition system to analyze the user's emotions. This process utilizes the Google Cloud Natural Language API to extract emotions from various parts of the text. Specifically, it generates emotional labels such as positive, negative, and neutral, and outputs the analysis results. This information is used to generate content optimized for each individual user.
[0371] Step 3:
[0372] The server generates relevant content using a generation mechanism based on the analyzed emotional information. The generated content must be in harmony with the user's emotions. For example, if negative emotions are analyzed, relaxing information or explanations will be selected. This allows for the output of customized information tailored to the user.
[0373] Step 4:
[0374] The server transmits the generated information to the terminal through an information delivery system. The terminal displays this information via a user interface. Users can receive and view information that corresponds to their own emotions. This can improve the user experience.
[0375] Step 5:
[0376] After receiving information that resonates with their emotions, users can add further feedback and comments as needed. This allows for even more optimized content suggestions in the future.
[0377] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0378] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0379] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0380] [Third Embodiment]
[0381] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0382] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0383] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0384] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0385] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0386] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0387] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0388] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0389] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0390] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0391] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0392] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0393] The system according to the present invention provides a specific implementation method for efficiently processing audio data and distributing information from press conferences and the like. A specific embodiment is described below.
[0394] This system consists of a server, terminals (users' PCs and smartphones), and users, all connected via a network. The server has the function of receiving audio data from an external source and transcribing it. The received audio data is automatically converted into text data using a speech recognition algorithm. This makes it easy for journalists and citizens who are not physically present at the press conference to understand the content.
[0395] A server equipped with generation AI takes text data as input and generates summaries and key points. The generated summaries are automatically published on an information distribution platform. The platform comes standard with a comment function, through which users (journalists and the general public) can express their opinions and reply to comments from other users. This function allows journalists to organize information efficiently and citizens to be exposed to diverse perspectives.
[0396] As a concrete example, if a political press conference is held, the server immediately receives the audio data and simultaneously begins transcribing it into text. The generation AI creates a summary from the text in real time, and that summary is quickly published. Users can comment on the summary, thereby promoting multifaceted discussion.
[0397] This system provides a platform for information sharing and discussion that transcends the boundaries between mass media and online media, enabling high-quality discourse.
[0398] The following describes the processing flow.
[0399] Step 1:
[0400] The server receives audio data from press conferences from an external source. Even when the audio arrives in real time, the buffering function is used to ensure stable data reception.
[0401] Step 2:
[0402] The server uses a speech recognition engine to convert the received audio data into text data. This records the content of the audio as text information.
[0403] Step 3:
[0404] The server inputs the transcribed text into the generating AI. The generating AI then summarizes the text and automatically generates key points based on a specified algorithm.
[0405] Step 4:
[0406] The server formats the generated summaries and perspectives and publishes them on the information distribution platform. The platform makes them accessible to users online.
[0407] Step 5:
[0408] The user's device (PC or smartphone) accesses the platform and views the published summaries and perspectives. The user accesses the information through their device.
[0409] Step 6:
[0410] Users can express their opinions through the comment function. Journalists can provide supplementary information based on the information they have gathered, and ordinary citizens can leave their own perceptions and questions as comments.
[0411] Step 7:
[0412] The server manages the exchange of comments between users and filters or moderates them as needed, thereby maintaining a healthy environment for discussion.
[0413] (Example 1)
[0414] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0415] In today's information society, there is a need to quickly and efficiently grasp audio information and accurately share necessary information. However, journalists and citizens who are not physically present at the scene face the challenge of not being able to easily acquire audio information and grasp the important points. Furthermore, there is a lack of effective means to promote discussion from diverse perspectives.
[0416] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0417] In this invention, the server includes receiving means for receiving acoustic information, transcription means for converting it into a written representation, and generation means for generating summaries and key points. This makes it possible for users who are not physically present at the site to quickly and accurately grasp the audio information and engage in discussions from diverse perspectives.
[0418] "Acoustic information" refers to all data acquired as sound, and is typically a signal acquired through devices such as microphones.
[0419] A "receiving means" is an element that has the function of receiving acoustic information from an external source and transferring it to the system for processing.
[0420] A "transcription means" is an element that performs the process of converting acquired acoustic information into a string of characters and generating a written representation.
[0421] "Written representation" refers to the representation of acoustic information in text format, visually showing the content of audio data.
[0422] "Generative means" refers to elements that perform processing to create a summary based on written expression and extract important information.
[0423] An "information distribution platform" is a system that provides users with generated summaries and important information, and enables interaction.
[0424] The "opinion function" is a feature that provides an interface for users to add their own opinions and comments to generated information and to communicate with others.
[0425] "User operation means" refers to the means used by users to view information or add opinions on the information distribution platform.
[0426] The "response function" is a feature that allows users to reply to comments in order to facilitate communication between them.
[0427] The "evaluation function" is a feature that allows other users to evaluate the quality of other users' opinions and comments, thereby promoting more active discussion.
[0428] This system achieves efficient processing and distribution of acoustic information by automating multiple processes. The system consists of servers, terminals (user computers and mobile devices), and users, all connected via a network.
[0429] The server first receives acoustic information. This acoustic information is typically audio data sent over the internet. Secure data transfer is ensured using a secure protocol. The received acoustic information is then converted into written form using speech recognition software. Specific software used for this process includes commonly available speech recognition APIs.
[0430] Next, the server uses a generative AI model to generate a summary and key points from the written text. The AI models used here include models widely known in the field of natural language processing. Specifically, the generative AI model is run by inputting a prompt such as "Please summarize this text and list the five key points."
[0431] The generated summaries and key information are immediately published through the information distribution platform. This platform is designed to update information in real time over the internet and enables two-way communication between users and servers.
[0432] Users of the device can view published summaries and key points, and express their own opinions on the platform through the opinion function. This enables discussions from diverse perspectives and stimulates information exchange among users.
[0433] As a concrete example, this system can receive audio information from a public press conference, input written statements in real time using pre-prepared prompts, generate summaries, and efficiently process the entire process up to publication. Thus, the ability to quickly access high-quality information regardless of location is a key feature of this system.
[0434] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0435] Step 1:
[0436] The server receives audio data over the network. It takes audio data as input and transfers it to the server using a secure protocol, ensuring data integrity and confidentiality. The output is the storage of audio data, which forms the basis for subsequent processing.
[0437] Step 2:
[0438] The server transcribes the received audio data using speech recognition software. The input is audio data, and the speech recognition engine analyzes the audio signal and converts it into text data. Specific operations include the extraction of acoustic features and language analysis using a probabilistic model. The output is text data representing the content of the audio in written form.
[0439] Step 3:
[0440] The server generates a summary and key points from transcribed text data using a generative AI model. The input is text data, and the generative model is given the prompt "Summarize this text and list five key points" for processing. Specifically, it uses natural language processing techniques to understand the overall structure of the text and extract important context. The output is a summary and a list of key points.
[0441] Step 4:
[0442] The server publishes the generated summary and key points to the information distribution platform. The input consists of the summary and key points, which are formatted into an appropriate format and uploaded to the information distribution platform. Specifically, this involves format conversion and data transmission to the web server, ensuring real-time publication. The output is text information made public to users.
[0443] Step 5:
[0444] Users using the terminal can access publicly available information and add their opinions. Input consists of information on the distribution platform and user comments. Users enter comments, which the system saves to a database and displays to other users. Output is updated content, enabling the exchange of opinions from diverse perspectives.
[0445] (Application Example 1)
[0446] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0447] Traditional information distribution services face challenges in the efficiency of their information gathering and distribution processes, particularly regarding press conferences and other events. They lack features for quickly summarizing important information and facilitating the exchange of opinions from multiple perspectives. To address these challenges, there is a need to provide a system that facilitates real-time information conversion and sharing, as well as smooth dialogue among users.
[0448] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0449] In this invention, the server includes an acquisition means for acquiring information, a conversion means for converting the information into a string, and a generation means for inputting the string and generating a summary and key points. This enables efficient information collection and distribution, as well as smooth exchange of opinions among users.
[0450] "Acquisition means" refers to a mechanism for receiving information from an external source and incorporating it into the system.
[0451] A "conversion method" is a process for analyzing acquired information and converting it into a string in a useful format.
[0452] The "generation method" refers to a function that uses the converted string to summarize information and automatically extract key perspectives.
[0453] A "distribution method" is a mechanism for publishing the generated summaries and perspectives on an information distribution platform and sharing them with a large number of users.
[0454] A "user interface means" is an interface that allows users to add opinions through the system and exchange opinions with others.
[0455] The "opinion exchange function" is a feature that allows users to post their own opinions and engage in dialogue with the opinions of other users.
[0456] The "response function" is a feature that allows multiple users to reply to or comment on opinions posted by other users.
[0457] The "rating function" is a function that provides a means to evaluate the quality and usefulness of opinions posted by users.
[0458] This invention is an information processing system comprising acquisition means, conversion means, generation means, distribution means, and user interface means.
[0459] The server receives audio data from press conferences, events, and other sources using information acquisition methods. The received audio data is then converted into text data using conversion methods. This process utilizes speech recognition software such as Google Cloud Speech-to-Text.
[0460] Text data is summarized using a generation AI model (e.g., the OpenAI GPT model) by a generation method, and key perspectives are extracted. This summarized data is then published in real time on an information distribution platform by a distribution method.
[0461] Users can access the user interface using their devices (smartphones or PCs) and add comments and exchange opinions with other users through the provided opinion exchange function. This utilizes a real-time database such as Firebase.
[0462] As a concrete example, when a user opens a news app to check the breaking news of a politician's press conference, they can immediately view summarized information, see comments from other users, and share their own opinions. An example of a prompt for the generative AI model could be an instruction such as, "Summarize the following press conference text and identify three key points."
[0463] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0464] Step 1:
[0465] The server receives audio data from the information source. The input here is either real-time or recorded audio data. The server acquires this data via a network connection and sends it to the conversion means.
[0466] Step 2:
[0467] The server uses a conversion mechanism to convert received audio data into text data. The input is audio data, and speech recognition software such as Google Cloud Speech-to-Text is used to output text data. Feature extraction from the audio data and transcription from the audio are performed using deep learning.
[0468] Step 3:
[0469] The server utilizes generation methods to summarize text data and extract key perspectives. The input is transformed text data, and the output is a summary and perspective information. Generative AI models such as the OpenAI GPT model are used. Text analysis and natural language processing techniques are employed to generate the summary and perspectives.
[0470] Step 4:
[0471] The server publishes the generated summaries and perspectives to the information distribution platform via a distribution method. The input is the generated summary data, and the output is the online distribution content. Cloud storage or databases are used for publication.
[0472] Step 5:
[0473] Users access the information distribution platform using their terminals and utilize the provided opinion exchange function. Input consists of distributed information and user comments, while output is the user's opinion. Interactive opinion sharing is achieved through the user interface.
[0474] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0475] The system according to the present invention provides specific technologies for considering user emotions in information provision and achieving more effective communication.
[0476] In one embodiment of the invention, this system consists of a server, a terminal (the user's PC or smartphone), and the user. In addition to conventional voice data reception and transcription functions, the server has an emotion engine. This emotion engine analyzes the user's input comment text and recognizes the user's underlying emotions. The recognized emotions are then used in subsequent interactions with the user.
[0477] As a concrete example, this sentiment engine is activated when a user reads an article on an information distribution platform and posts a comment. If a user posts a stressful comment, the server can analyze it and generate an interface that soothes the user's emotions by providing additional information and diagrams to help them relax. The server also uses a sentiment-based feedback system to encourage constructive discussion and facilitate the exchange of opinions among multiple users.
[0478] Furthermore, the device displays pre-adjusted content sent from the server, allowing users to receive information tailored to their emotions. In this way, the emotion engine understands the user's feelings and adjusts the content accordingly, enabling users to enjoy a more satisfying experience. This system aims to create high-quality information delivery and communication by taking emotions into consideration.
[0479] The following describes the processing flow.
[0480] Step 1:
[0481] The server receives audio data from the press conference from an external source. The audio files are stored for processing according to the format.
[0482] Step 2:
[0483] The server uses a speech recognition engine to convert the received audio data into text data. This conversion transcribes the content.
[0484] Step 3:
[0485] The server passes the transcribed text data to a generation AI model, which generates a summary and key points. The generated information is then used for subsequent processing.
[0486] Step 4:
[0487] The server publishes information so that users can view it on the platform. Additionally, a sentiment engine monitors user interactions.
[0488] Step 5:
[0489] The terminal displays published summaries and perspectives. An interface is provided for users to view the information and enter comments.
[0490] Step 6:
[0491] When a user enters and submits a comment, the server analyzes the comment using a sentiment engine. This analysis identifies the sentiment contained in the comment.
[0492] Step 7:
[0493] The server adjusts the information presentation based on the sentiment analysis results. Additional information and feedback adapted to the user's emotions are generated and provided to the user through the user interface.
[0494] Step 8:
[0495] Users receive responses from the system based on their own reactions, allowing them to provide context-appropriate feedback and additional interactions. This process is repeated to continuously enhance the user experience.
[0496] (Example 2)
[0497] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0498] Conventional information distribution systems provide information uniformly without considering users' emotions, resulting in a lack of emotionally-driven interaction and challenges in information acceptance and satisfaction. Furthermore, they lacked appropriate functions to facilitate constructive exchange of opinions among users, leading to insufficient communication quality.
[0499] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0500] In this invention, the server includes receiving means for receiving voice information, conversion means for converting it into text information, and emotion recognition means for analyzing the text information to identify emotions. This makes it possible to design interactions that respond to the user's emotions, and further enables the provision of emotion-based, tailored information and the promotion of constructive exchange of opinions among users.
[0501] "Audio information" refers to data that represents spoken language and other audio signals in digital or analog format.
[0502] "Receiving means" refers to a device or program for receiving information from an external source, and in this case, it refers to the function of receiving audio information.
[0503] "Textual information" refers to data obtained by converting audio information into text format, which is interpreted as sentences or words.
[0504] "Conversion means" refers to a device or program for converting data in one format to another, and in this case, it refers to the function of converting audio information into text information.
[0505] "Emotion recognition means" refers to a technology or program that analyzes textual information and identifies the user's emotions from it.
[0506] "Interaction design" is the process of dynamically adjusting the content and user interface provided by a system in response to the user's emotions and needs.
[0507] An "information sharing medium" is a platform or network for providing and sharing information with a wide range of users.
[0508] "Interactive features" refer to functions that enable users to input information into the system and exchange information with other users.
[0509] A "response function" is a function that provides appropriate responses or feedback to received information.
[0510] A "feedback function" is a feature that incorporates user input and comments into the system or service.
[0511] This system is comprised of a server, terminals (e.g., the user's personal computer or mobile device), and the user themselves. The server, at the core of the invention, is equipped with a speech-to-text engine for receiving speech information and converting it into text information. Specifically, speech information is transmitted to the server via a network and converted into text data using the conversion engine. General-purpose speech recognition software is used for this engine. The core operation of this system is described below.
[0512] After the server converts the audio information to text, it uses an emotion recognition engine to identify the emotions. This engine analyzes the text using a natural language processing library to identify the emotions contained within it. The results of the emotion analysis are then used to design subsequent interactions. For example, the server adjusts the information and content it provides according to specific emotions (e.g., stress or joy). To this end, a content management system works in conjunction to dynamically generate and provide content.
[0513] The terminal is responsible for displaying the adjusted content sent from the server to the user. The displayed content is visually formatted using HTML and CSS. This allows users to receive information that aligns with their emotions, enabling them to enjoy a more satisfying user experience.
[0514] As a concrete example, consider a scenario where a user reads a news article and enters a comment on it. In this case, the entered comment is analyzed by the server, and if the emotion is determined to be "anger," the server provides content that addresses that emotion. This may include information on relaxation techniques or detailed illustrations on a specific topic.
[0515] An example of a prompt for a generative AI model is, "Please show a method for identifying emotions from user comments and providing appropriate feedback." This prompt allows the AI model to provide insights into emotion analysis and appropriate content generation.
[0516] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0517] Step 1:
[0518] The server receives voice information from the user. A voice data file is passed as input, and the server uses a speech recognition engine to analyze the data. This converts the voice into text information, generating text data. The output is text information.
[0519] Step 2:
[0520] The server inputs the generated text information into the sentiment recognition engine. Here, natural language processing algorithms are used to analyze keywords and context within the text and identify the user's emotions. During this sentiment analysis process, the text is assigned a sentiment score, such as positive, negative, or neutral. The output is the analysis result, including the sentiment score.
[0521] Step 3:
[0522] The server adjusts the content it provides based on the results of sentiment analysis. Sentiment scores are used as input, and a content management system is utilized to select information corresponding to specific emotions. For example, if the user is identified as stressed, relaxation information and specific visual content will be selected. The adjusted content is then generated as output.
[0523] Step 4:
[0524] The device receives the formatted content from the server and displays it to the user. The input is HTML and CSS code sent from the server, which is then rendered in a web browser. The device provides a visual interface through the browser, allowing the user to view the content and send feedback as needed. The output is the provision of visual information to the user.
[0525] Step 5:
[0526] Users enter feedback and new comments through the displayed interface. The input is sent to the server as text data, which then re-analyzes the data to initiate a new analysis cycle. This facilitates constructive exchange of ideas among users and deepens the discussion. Further refined content is provided as output from the re-analysis.
[0527] (Application Example 2)
[0528] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0529] In today's information delivery landscape, users offer comments and feedback with a wide range of emotions, but the lack of information delivery that adequately considers these emotions results in insufficient improvement of the user experience. To address this issue, there is a need for a means to accurately capture users' emotions and provide appropriate information that responds to those emotions.
[0530] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0531] In this invention, the server includes an acquisition means for receiving audio data, a conversion means for converting the audio data into text information, and an emotion recognition means for inputting the text information and analyzing emotions. This enables the provision of information and recommendation of appropriate content based on the user's emotions.
[0532] "Audio data" refers to data recorded in digital format, which is fundamental information used to analyze or process audio.
[0533] "Acquisition means" refers to functions or devices for receiving or collecting audio data.
[0534] A "conversion means" is a mechanism for processing received audio data into text information.
[0535] "Textual information" refers to information that is expressed as text based on audio data, and is written in text format.
[0536] "Emotion recognition means" refers to algorithms and technologies for analyzing and identifying a user's emotions from text information.
[0537] A "generation means" is a device or process that has the function of generating related information or content based on analyzed emotional information.
[0538] "Information provision means" refers to the means of presenting generated information to the user, and includes user interfaces and platforms.
[0539] "User interface means" refers to an interface that allows users to receive and manipulate information, such as a device screen or application interface.
[0540] To realize this invention, a server plays a central role. The server first has a receiving means for acquiring audio data and utilizes a conversion means to convert this audio data into text information. By using a widely used solution as speech recognition software, such as the Google Cloud Speech-to-Text API, it is possible to convert speech to text with high accuracy.
[0541] Next, the server applies emotion recognition tools to analyze the emotions based on the converted text information. In this process, the Google Cloud Natural Language API is used as the software tool for emotion analysis. This allows the server to determine positive, negative, or neutral emotions from user comments and feedback.
[0542] Based on the analyzed sentiment information, the server generates relevant information and content tailored to the user's emotions through a generation mechanism. This generation includes summarizing the transcribed content and algorithms to provide information best suited to the user's current emotions. The generated information is presented to the user through an information delivery mechanism. The user interface operates on devices such as smartphones and tablets, displaying the information in a format that the user can view and interact with.
[0543] As a concrete example, suppose a user watches a movie review video and comments, "This ending is confusing." In this case, the server's emotion engine identifies the emotion of "confusion," generates content recommending similar movie commentary articles or videos, and displays it in the user interface.
[0544] An example of a prompt to input into a generative AI model is as follows: "A user posted the comment 'I didn't understand the ending of this movie.' Analyze the sentiment of this comment and suggest appropriate content."
[0545] This allows users to receive information that takes their feelings into consideration, enabling them to enjoy a better experience.
[0546] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0547] Step 1:
[0548] The server receives audio data. Using the received audio data as input, it performs speech recognition using a conversion method and converts it into text information. By utilizing the Google Cloud Speech-to-Text API, it outputs highly accurate transcription results. This is for use as basic data for information provision.
[0549] Step 2:
[0550] The server passes text information as input to an emotion recognition system to analyze the user's emotions. This process utilizes the Google Cloud Natural Language API to extract emotions from various parts of the text. Specifically, it generates emotional labels such as positive, negative, and neutral, and outputs the analysis results. This information is used to generate content optimized for each individual user.
[0551] Step 3:
[0552] The server generates relevant content using a generation mechanism based on the analyzed emotional information. The generated content must be in harmony with the user's emotions. For example, if negative emotions are analyzed, relaxing information or explanations will be selected. This allows for the output of customized information tailored to the user.
[0553] Step 4:
[0554] The server transmits the generated information to the terminal through an information delivery system. The terminal displays this information via a user interface. Users can receive and view information that corresponds to their own emotions. This can improve the user experience.
[0555] Step 5:
[0556] After receiving information that resonates with their emotions, users can add further feedback and comments as needed. This allows for even more optimized content suggestions in the future.
[0557] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0558] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0559] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0560] [Fourth Embodiment]
[0561] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0562] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0563] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0564] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0565] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0566] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0567] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0568] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0569] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0570] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0571] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0572] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0573] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0574] The system according to the present invention provides a specific implementation method for efficiently processing audio data and distributing information from press conferences and the like. A specific embodiment is described below.
[0575] This system consists of a server, terminals (users' PCs and smartphones), and users, all connected via a network. The server has the function of receiving audio data from an external source and transcribing it. The received audio data is automatically converted into text data using a speech recognition algorithm. This makes it easy for journalists and citizens who are not physically present at the press conference to understand the content.
[0576] A server equipped with generation AI takes text data as input and generates summaries and key points. The generated summaries are automatically published on an information distribution platform. The platform comes standard with a comment function, through which users (journalists and the general public) can express their opinions and reply to comments from other users. This function allows journalists to organize information efficiently and citizens to be exposed to diverse perspectives.
[0577] As a concrete example, if a political press conference is held, the server immediately receives the audio data and simultaneously begins transcribing it into text. The generation AI creates a summary from the text in real time, and that summary is quickly published. Users can comment on the summary, thereby promoting multifaceted discussion.
[0578] This system provides a platform for information sharing and discussion that transcends the boundaries between mass media and online media, enabling high-quality discourse.
[0579] The following describes the processing flow.
[0580] Step 1:
[0581] The server receives audio data from press conferences from an external source. Even when the audio arrives in real time, the buffering function is used to ensure stable data reception.
[0582] Step 2:
[0583] The server uses a speech recognition engine to convert the received audio data into text data. This records the content of the audio as text information.
[0584] Step 3:
[0585] The server inputs the transcribed text into the generating AI. The generating AI then summarizes the text and automatically generates key points based on a specified algorithm.
[0586] Step 4:
[0587] The server formats the generated summaries and perspectives and publishes them on the information distribution platform. The platform makes them accessible to users online.
[0588] Step 5:
[0589] The user's device (PC or smartphone) accesses the platform and views the published summaries and perspectives. The user accesses the information through their device.
[0590] Step 6:
[0591] Users can express their opinions through the comment function. Journalists can provide supplementary information based on the information they have gathered, and ordinary citizens can leave their own perceptions and questions as comments.
[0592] Step 7:
[0593] The server manages the exchange of comments between users and filters or moderates them as needed, thereby maintaining a healthy environment for discussion.
[0594] (Example 1)
[0595] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0596] In today's information society, there is a need to quickly and efficiently grasp audio information and accurately share necessary information. However, journalists and citizens who are not physically present at the scene face the challenge of not being able to easily acquire audio information and grasp the important points. Furthermore, there is a lack of effective means to promote discussion from diverse perspectives.
[0597] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0598] In this invention, the server includes receiving means for receiving acoustic information, transcription means for converting it into a written representation, and generation means for generating summaries and key points. This makes it possible for users who are not physically present at the site to quickly and accurately grasp the audio information and engage in discussions from diverse perspectives.
[0599] "Acoustic information" refers to all data acquired as sound, and is typically a signal acquired through devices such as microphones.
[0600] A "receiving means" is an element that has the function of receiving acoustic information from an external source and transferring it to the system for processing.
[0601] A "transcription means" is an element that performs the process of converting acquired acoustic information into a string of characters and generating a written representation.
[0602] "Written representation" refers to the representation of acoustic information in text format, visually showing the content of audio data.
[0603] "Generative means" refers to elements that perform processing to create a summary based on written expression and extract important information.
[0604] An "information distribution platform" is a system that provides users with generated summaries and important information, and enables interaction.
[0605] The "opinion function" is a feature that provides an interface for users to add their own opinions and comments to generated information and to communicate with others.
[0606] "User operation means" refers to the means used by users to view information or add opinions on the information distribution platform.
[0607] The "response function" is a feature that allows users to reply to comments in order to facilitate communication between them.
[0608] The "evaluation function" is a feature that allows other users to evaluate the quality of other users' opinions and comments, thereby promoting more active discussion.
[0609] This system achieves efficient processing and distribution of acoustic information by automating multiple processes. The system consists of servers, terminals (user computers and mobile devices), and users, all connected via a network.
[0610] The server first receives acoustic information. This acoustic information is typically audio data sent over the internet. Secure data transfer is ensured using a secure protocol. The received acoustic information is then converted into written form using speech recognition software. Specific software used for this process includes commonly available speech recognition APIs.
[0611] Next, the server uses a generative AI model to generate a summary and key points from the written text. The AI models used here include models widely known in the field of natural language processing. Specifically, the generative AI model is run by inputting a prompt such as "Please summarize this text and list the five key points."
[0612] The generated summaries and key information are immediately published through the information distribution platform. This platform is designed to update information in real time over the internet and enables two-way communication between users and servers.
[0613] Users of the device can view published summaries and key points, and express their own opinions on the platform through the opinion function. This enables discussions from diverse perspectives and stimulates information exchange among users.
[0614] As a concrete example, this system can receive audio information from a public press conference, input written statements in real time using pre-prepared prompts, generate summaries, and efficiently process the entire process up to publication. Thus, the ability to quickly access high-quality information regardless of location is a key feature of this system.
[0615] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0616] Step 1:
[0617] The server receives audio data over the network. It takes audio data as input and transfers it to the server using a secure protocol, ensuring data integrity and confidentiality. The output is the storage of audio data, which forms the basis for subsequent processing.
[0618] Step 2:
[0619] The server transcribes the received audio data using speech recognition software. The input is audio data, and the speech recognition engine analyzes the audio signal and converts it into text data. Specific operations include the extraction of acoustic features and language analysis using a probabilistic model. The output is text data representing the content of the audio in written form.
[0620] Step 3:
[0621] The server generates a summary and key points from transcribed text data using a generative AI model. The input is text data, and the generative model is given the prompt "Summarize this text and list five key points" for processing. Specifically, it uses natural language processing techniques to understand the overall structure of the text and extract important context. The output is a summary and a list of key points.
[0622] Step 4:
[0623] The server publishes the generated summary and key points to the information distribution platform. The input consists of the summary and key points, which are formatted into an appropriate format and uploaded to the information distribution platform. Specifically, this involves format conversion and data transmission to the web server, ensuring real-time publication. The output is text information made public to users.
[0624] Step 5:
[0625] Users using the terminal can access publicly available information and add their opinions. Input consists of information on the distribution platform and user comments. Users enter comments, which the system saves to a database and displays to other users. Output is updated content, enabling the exchange of opinions from diverse perspectives.
[0626] (Application Example 1)
[0627] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0628] Traditional information distribution services face challenges in the efficiency of their information gathering and distribution processes, particularly regarding press conferences and other events. They lack features for quickly summarizing important information and facilitating the exchange of opinions from multiple perspectives. To address these challenges, there is a need to provide a system that facilitates real-time information conversion and sharing, as well as smooth dialogue among users.
[0629] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0630] In this invention, the server includes an acquisition means for acquiring information, a conversion means for converting the information into a string, and a generation means for inputting the string and generating a summary and key points. This enables efficient information collection and distribution, as well as smooth exchange of opinions among users.
[0631] "Acquisition means" refers to a mechanism for receiving information from an external source and incorporating it into the system.
[0632] A "conversion method" is a process for analyzing acquired information and converting it into a string in a useful format.
[0633] The "generation method" refers to a function that uses the converted string to summarize information and automatically extract key perspectives.
[0634] A "distribution method" is a mechanism for publishing the generated summaries and perspectives on an information distribution platform and sharing them with a large number of users.
[0635] A "user interface means" is an interface that allows users to add opinions through the system and exchange opinions with others.
[0636] The "opinion exchange function" is a feature that allows users to post their own opinions and engage in dialogue with the opinions of other users.
[0637] The "response function" is a feature that allows multiple users to reply to or comment on opinions posted by other users.
[0638] The "rating function" is a function that provides a means to evaluate the quality and usefulness of opinions posted by users.
[0639] This invention is an information processing system comprising acquisition means, conversion means, generation means, distribution means, and user interface means.
[0640] The server receives audio data from press conferences, events, and other sources using information acquisition methods. The received audio data is then converted into text data using conversion methods. This process utilizes speech recognition software such as Google Cloud Speech-to-Text.
[0641] Text data is summarized using a generation AI model (e.g., the OpenAI GPT model) by a generation method, and key perspectives are extracted. This summarized data is then published in real time on an information distribution platform by a distribution method.
[0642] Users can access the user interface using their devices (smartphones or PCs) and add comments and exchange opinions with other users through the provided opinion exchange function. This utilizes a real-time database such as Firebase.
[0643] As a concrete example, when a user opens a news app to check the breaking news of a politician's press conference, they can immediately view summarized information, see comments from other users, and share their own opinions. An example of a prompt for the generative AI model could be an instruction such as, "Summarize the following press conference text and identify three key points."
[0644] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0645] Step 1:
[0646] The server receives audio data from the information source. The input here is either real-time or recorded audio data. The server acquires this data via a network connection and sends it to the conversion means.
[0647] Step 2:
[0648] The server uses a conversion mechanism to convert received audio data into text data. The input is audio data, and speech recognition software such as Google Cloud Speech-to-Text is used to output text data. Feature extraction from the audio data and transcription from the audio are performed using deep learning.
[0649] Step 3:
[0650] The server utilizes generation methods to summarize text data and extract key perspectives. The input is transformed text data, and the output is a summary and perspective information. Generative AI models such as the OpenAI GPT model are used. Text analysis and natural language processing techniques are employed to generate the summary and perspectives.
[0651] Step 4:
[0652] The server publishes the generated summaries and perspectives to the information distribution platform via a distribution method. The input is the generated summary data, and the output is the online distribution content. Cloud storage or databases are used for publication.
[0653] Step 5:
[0654] Users access the information distribution platform using their terminals and utilize the provided opinion exchange function. Input consists of distributed information and user comments, while output is the user's opinion. Interactive opinion sharing is achieved through the user interface.
[0655] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0656] The system according to the present invention provides specific technologies for considering user emotions in information provision and achieving more effective communication.
[0657] In one embodiment of the invention, this system consists of a server, a terminal (the user's PC or smartphone), and the user. In addition to conventional voice data reception and transcription functions, the server has an emotion engine. This emotion engine analyzes the user's input comment text and recognizes the user's underlying emotions. The recognized emotions are then used in subsequent interactions with the user.
[0658] As a concrete example, this sentiment engine is activated when a user reads an article on an information distribution platform and posts a comment. If a user posts a stressful comment, the server can analyze it and generate an interface that soothes the user's emotions by providing additional information and diagrams to help them relax. The server also uses a sentiment-based feedback system to encourage constructive discussion and facilitate the exchange of opinions among multiple users.
[0659] Furthermore, the device displays pre-adjusted content sent from the server, allowing users to receive information tailored to their emotions. In this way, the emotion engine understands the user's feelings and adjusts the content accordingly, enabling users to enjoy a more satisfying experience. This system aims to create high-quality information delivery and communication by taking emotions into consideration.
[0660] The following describes the processing flow.
[0661] Step 1:
[0662] The server receives audio data from the press conference from an external source. The audio files are stored for processing according to the format.
[0663] Step 2:
[0664] The server uses a speech recognition engine to convert the received audio data into text data. This conversion transcribes the content.
[0665] Step 3:
[0666] The server passes the transcribed text data to a generation AI model, which generates a summary and key points. The generated information is then used for subsequent processing.
[0667] Step 4:
[0668] The server publishes information so that users can view it on the platform. Additionally, a sentiment engine monitors user interactions.
[0669] Step 5:
[0670] The terminal displays published summaries and perspectives. An interface is provided for users to view the information and enter comments.
[0671] Step 6:
[0672] When a user enters and submits a comment, the server analyzes the comment using a sentiment engine. This analysis identifies the sentiment contained in the comment.
[0673] Step 7:
[0674] The server adjusts the information presentation based on the sentiment analysis results. Additional information and feedback adapted to the user's emotions are generated and provided to the user through the user interface.
[0675] Step 8:
[0676] Users receive responses from the system based on their own reactions, allowing them to provide context-appropriate feedback and additional interactions. This process is repeated to continuously enhance the user experience.
[0677] (Example 2)
[0678] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0679] Conventional information distribution systems provide information uniformly without considering users' emotions, resulting in a lack of emotionally-driven interaction and challenges in information acceptance and satisfaction. Furthermore, they lacked appropriate functions to facilitate constructive exchange of opinions among users, leading to insufficient communication quality.
[0680] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0681] In this invention, the server includes receiving means for receiving voice information, conversion means for converting it into text information, and emotion recognition means for analyzing the text information to identify emotions. This makes it possible to design interactions that respond to the user's emotions, and further enables the provision of emotion-based, tailored information and the promotion of constructive exchange of opinions among users.
[0682] "Audio information" refers to data that represents spoken language and other audio signals in digital or analog format.
[0683] "Receiving means" refers to a device or program for receiving information from an external source, and in this case, it refers to the function of receiving audio information.
[0684] "Textual information" refers to data obtained by converting audio information into text format, which is interpreted as sentences or words.
[0685] "Conversion means" refers to a device or program for converting data in one format to another, and in this case, it refers to the function of converting audio information into text information.
[0686] "Emotion recognition means" refers to a technology or program that analyzes textual information and identifies the user's emotions from it.
[0687] "Interaction design" is the process of dynamically adjusting the content and user interface provided by a system in response to the user's emotions and needs.
[0688] An "information sharing medium" is a platform or network for providing and sharing information with a wide range of users.
[0689] "Interactive features" refer to functions that enable users to input information into the system and exchange information with other users.
[0690] A "response function" is a function that provides appropriate responses or feedback to received information.
[0691] A "feedback function" is a feature that incorporates user input and comments into the system or service.
[0692] This system is comprised of a server, terminals (e.g., the user's personal computer or mobile device), and the user themselves. The server, at the core of the invention, is equipped with a speech-to-text engine for receiving speech information and converting it into text information. Specifically, speech information is transmitted to the server via a network and converted into text data using the conversion engine. General-purpose speech recognition software is used for this engine. The core operation of this system is described below.
[0693] After the server converts the audio information to text, it uses an emotion recognition engine to identify the emotions. This engine analyzes the text using a natural language processing library to identify the emotions contained within it. The results of the emotion analysis are then used to design subsequent interactions. For example, the server adjusts the information and content it provides according to specific emotions (e.g., stress or joy). To this end, a content management system works in conjunction to dynamically generate and provide content.
[0694] The terminal is responsible for displaying the adjusted content sent from the server to the user. The displayed content is visually formatted using HTML and CSS. This allows users to receive information that aligns with their emotions, enabling them to enjoy a more satisfying user experience.
[0695] As a concrete example, consider a scenario where a user reads a news article and enters a comment on it. In this case, the entered comment is analyzed by the server, and if the emotion is determined to be "anger," the server provides content that addresses that emotion. This may include information on relaxation techniques or detailed illustrations on a specific topic.
[0696] An example of a prompt for a generative AI model is, "Please show a method for identifying emotions from user comments and providing appropriate feedback." This prompt allows the AI model to provide insights into emotion analysis and appropriate content generation.
[0697] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0698] Step 1:
[0699] The server receives voice information from the user. A voice data file is passed as input, and the server uses a speech recognition engine to analyze the data. This converts the voice into text information, generating text data. The output is text information.
[0700] Step 2:
[0701] The server inputs the generated text information into the sentiment recognition engine. Here, natural language processing algorithms are used to analyze keywords and context within the text and identify the user's emotions. During this sentiment analysis process, the text is assigned a sentiment score, such as positive, negative, or neutral. The output is the analysis result, including the sentiment score.
[0702] Step 3:
[0703] The server adjusts the content it provides based on the results of sentiment analysis. Sentiment scores are used as input, and a content management system is utilized to select information corresponding to specific emotions. For example, if the user is identified as stressed, relaxation information and specific visual content will be selected. The adjusted content is then generated as output.
[0704] Step 4:
[0705] The device receives the formatted content from the server and displays it to the user. The input is HTML and CSS code sent from the server, which is then rendered in a web browser. The device provides a visual interface through the browser, allowing the user to view the content and send feedback as needed. The output is the provision of visual information to the user.
[0706] Step 5:
[0707] Users enter feedback and new comments through the displayed interface. The input is sent to the server as text data, which then re-analyzes the data to initiate a new analysis cycle. This facilitates constructive exchange of ideas among users and deepens the discussion. Further refined content is provided as output from the re-analysis.
[0708] (Application Example 2)
[0709] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0710] In today's information delivery landscape, users offer comments and feedback with a wide range of emotions, but the lack of information delivery that adequately considers these emotions results in insufficient improvement of the user experience. To address this issue, there is a need for a means to accurately capture users' emotions and provide appropriate information that responds to those emotions.
[0711] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0712] In this invention, the server includes an acquisition means for receiving audio data, a conversion means for converting the audio data into text information, and an emotion recognition means for inputting the text information and analyzing emotions. This enables the provision of information and recommendation of appropriate content based on the user's emotions.
[0713] "Audio data" refers to data recorded in digital format, which is fundamental information used to analyze or process audio.
[0714] "Acquisition means" refers to functions or devices for receiving or collecting audio data.
[0715] A "conversion means" is a mechanism for processing received audio data into text information.
[0716] "Textual information" refers to information that is expressed as text based on audio data, and is written in text format.
[0717] "Emotion recognition means" refers to algorithms and technologies for analyzing and identifying a user's emotions from text information.
[0718] A "generation means" is a device or process that has the function of generating related information or content based on analyzed emotional information.
[0719] "Information provision means" refers to the means of presenting generated information to the user, and includes user interfaces and platforms.
[0720] "User interface means" refers to an interface that allows users to receive and manipulate information, such as a device screen or application interface.
[0721] To realize this invention, a server plays a central role. The server first has a receiving means for acquiring audio data and utilizes a conversion means to convert this audio data into text information. By using a widely used solution as speech recognition software, such as the Google Cloud Speech-to-Text API, it is possible to convert speech to text with high accuracy.
[0722] Next, the server applies emotion recognition tools to analyze the emotions based on the converted text information. In this process, the Google Cloud Natural Language API is used as the software tool for emotion analysis. This allows the server to determine positive, negative, or neutral emotions from user comments and feedback.
[0723] Based on the analyzed sentiment information, the server generates relevant information and content tailored to the user's emotions through a generation mechanism. This generation includes summarizing the transcribed content and algorithms to provide information best suited to the user's current emotions. The generated information is presented to the user through an information delivery mechanism. The user interface operates on devices such as smartphones and tablets, displaying the information in a format that the user can view and interact with.
[0724] As a concrete example, suppose a user watches a movie review video and comments, "This ending is confusing." In this case, the server's emotion engine identifies the emotion of "confusion," generates content recommending similar movie commentary articles or videos, and displays it in the user interface.
[0725] An example of a prompt to input into a generative AI model is as follows: "A user posted the comment 'I didn't understand the ending of this movie.' Analyze the sentiment of this comment and suggest appropriate content."
[0726] This allows users to receive information that takes their feelings into consideration, enabling them to enjoy a better experience.
[0727] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0728] Step 1:
[0729] The server receives audio data. Using the received audio data as input, it performs speech recognition using a conversion method and converts it into text information. By utilizing the Google Cloud Speech-to-Text API, it outputs highly accurate transcription results. This is for use as basic data for information provision.
[0730] Step 2:
[0731] The server passes text information as input to an emotion recognition system to analyze the user's emotions. This process utilizes the Google Cloud Natural Language API to extract emotions from various parts of the text. Specifically, it generates emotional labels such as positive, negative, and neutral, and outputs the analysis results. This information is used to generate content optimized for each individual user.
[0732] Step 3:
[0733] The server generates relevant content using a generation mechanism based on the analyzed emotional information. The generated content must be in harmony with the user's emotions. For example, if negative emotions are analyzed, relaxing information or explanations will be selected. This allows for the output of customized information tailored to the user.
[0734] Step 4:
[0735] The server transmits the generated information to the terminal through an information delivery system. The terminal displays this information via a user interface. Users can receive and view information that corresponds to their own emotions. This can improve the user experience.
[0736] Step 5:
[0737] After receiving information that resonates with their emotions, users can add further feedback and comments as needed. This allows for even more optimized content suggestions in the future.
[0738] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0739] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0740] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0741] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0742] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0743] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0744] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0745] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0746] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0747] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0748] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0749] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0750] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0751] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0752] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0753] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0754] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0755] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0756] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0757] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0758] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0759] The following is further disclosed regarding the embodiments described above.
[0760] (Claim 1)
[0761] A receiving means for receiving audio data,
[0762] A transcription means that converts the aforementioned audio data into text,
[0763] A generation means that inputs the aforementioned text and generates a summary and key points,
[0764] A distribution method for publishing the generated summaries and perspectives to an information distribution platform,
[0765] A user interface means that provides a comment function that allows users to add comments,
[0766] A system that includes this.
[0767] (Claim 2)
[0768] The system according to claim 1, wherein the generation means automatically extracts important points based on the text generated by the transcription means.
[0769] (Claim 3)
[0770] The system according to claim 1, wherein the comment function includes a reply function and a feedback function for facilitating discussions among multiple users.
[0771] "Example 1"
[0772] (Claim 1)
[0773] A receiving means for receiving acoustic information,
[0774] A transcription means that converts the aforementioned acoustic information into a written representation,
[0775] A generation means that inputs the aforementioned written expression and generates a summary and important points,
[0776] A distribution means for publishing the generated summary and key information to the information distribution platform,
[0777] A user operation means that provides an opinion function that allows users to add opinions,
[0778] A system that includes this.
[0779] (Claim 2)
[0780] The system according to claim 1, wherein the generation means automatically extracts important information based on the written expression generated by the transcription means.
[0781] (Claim 3)
[0782] The system according to claim 1, wherein the opinion function is equipped with a response function and an evaluation function for facilitating discussion among a large number of users.
[0783] "Application Example 1"
[0784] (Claim 1)
[0785] Means of acquiring information,
[0786] A conversion means that converts the aforementioned information into a string,
[0787] A generation means that inputs the aforementioned string and generates a summary and key points,
[0788] A distribution method for publishing the generated summaries and perspectives to an information distribution platform,
[0789] A user interface means that provides an opinion exchange function in which users can add their opinions,
[0790] A system that includes this.
[0791] (Claim 2)
[0792] The system according to claim 1, wherein the generation means automatically extracts important points based on the string generated by the conversion means.
[0793] (Claim 3)
[0794] The system according to claim 1, wherein the opinion exchange function is equipped with a response function and an evaluation function for facilitating dialogue among multiple users.
[0795] "Example 2 of combining an emotion engine"
[0796] (Claim 1)
[0797] A receiving means for receiving audio information,
[0798] A conversion means that converts the aforementioned audio information into text information,
[0799] An emotion recognition means that analyzes the aforementioned textual information and identifies emotions,
[0800] An adjustment mechanism that adjusts the content presented based on identified emotions,
[0801] A distribution method for publishing the adjusted content to an information sharing medium,
[0802] A display means that provides a dialogue function that allows users to add their opinions,
[0803] A system that includes this.
[0804] (Claim 2)
[0805] The system according to claim 1, wherein the emotion recognition means identifies an emotion based on the character information generated by the conversion means and designs an interaction.
[0806] (Claim 3)
[0807] The system according to claim 1, wherein the dialogue function includes a response function and an opinion reflection function for facilitating information exchange among multiple users.
[0808] "Application example 2 when combining with an emotional engine"
[0809] (Claim 1)
[0810] A means for receiving audio data,
[0811] A conversion means that converts the aforementioned audio data into text information,
[0812] An emotion recognition means that takes the aforementioned textual information as input and analyzes emotions,
[0813] A generation means that generates highly adjusted information based on analyzed emotions,
[0814] An information provision means equipped with a feedback function to facilitate the exchange of opinions among users regarding the generated information,
[0815] A user interface means that allows users to receive additional information adjusted based on their emotions,
[0816] A system that includes this.
[0817] (Claim 2)
[0818] The system according to claim 1, wherein the generation means automatically evaluates emotions based on the character information generated by the conversion means and selects relevant information.
[0819] (Claim 3)
[0820] The system according to claim 1, wherein the information provision means improves user satisfaction by suggesting optimal content to the user in accordance with the analyzed emotions. [Explanation of Symbols]
[0821] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A receiving means for receiving audio data, A transcription means that converts the aforementioned audio data into text, A generation means that inputs the aforementioned text and generates a summary and key points, A distribution method for publishing the generated summaries and perspectives to an information distribution platform, A user interface means that provides a comment function that allows users to add comments, A system that includes this.
2. The system according to claim 1, wherein the generation means automatically extracts important points based on the text generated by the transcription means.
3. The system according to claim 1, wherein the comment function is further equipped with a reply function and a feedback function for facilitating discussions among multiple users.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A