System
A generative AI-based scoring system enhances transparency and participation in Diet activities by summarizing speeches, generating report cards, and enabling direct voter feedback, addressing the challenges of understanding and interacting with Diet members.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-19
- Publication Date
- 2026-03-04
AI Technical Summary
Ordinary voters face difficulties in understanding the activities and statements of Diet members, leading to a lack of transparency and voter participation, with limited communication channels and trust issues due to inconsistent tracking of election pledges and Diet deliberations.
A scoring system utilizing generative artificial intelligence to summarize speeches and questions, evaluate member activity and performance, generate report cards, and facilitate direct voter feedback through a dashboard and social media sharing.
Increases transparency and promotes political participation by enabling voters to easily understand and interact with Diet members, enhancing communication and trust through detailed evaluations and feedback mechanisms.
Smart Images

Figure 2026035289000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] It is difficult for ordinary voters to understand the activities and statements of Diet members and ensure transparency. This has led to a lack of voter participation and understanding. In particular, it is difficult for the media and journalists to cover all of the deliberations in the Diet, and the current situation, in which it is not possible to track the consistency between election pledges and Diet deliberations, is causing a loss of voter trust. In addition, there is a lack of a system for voters to send direct feedback to Diet members, limiting communication between the two. There is a need to solve this issue and realize a more transparent and participatory democracy. [Means for solving the problem]
[0005] The present invention provides a scoring system that uses generative artificial intelligence to summarize the speeches and questions and answers of Diet members, and evaluates the activity level and achievements of Diet members based on the summaries. This system includes the following means.
[0006] 1. A method of converting voice data into text data using generative artificial intelligence.
[0007] 2. A means of automatically summarizing text data.
[0008] 3. A method for calculating scores based on summarized data and analyzing multiple metrics to assess legislators' activity and performance.
[0009] 4. A means of generating a report card evaluation of each member of parliament based on the calculated scores.
[0010] 5. A means of translating the generated report cards and summary data into multiple languages.
[0011] 6. A means for voters to type and send messages to their representatives through a dashboard.
[0012] 7. A means of sharing report cards and summary data on social media.
[0013] By integrating these methods, it will be possible to increase the transparency of the activities and statements of Diet members and promote political participation among ordinary voters. This will allow voters to easily understand the activities of their Diet members and send direct feedback, promoting communication between them and making it easier to build trust.
[0014] "Generative artificial intelligence" refers to an artificial intelligence system that has the ability to learn from sufficient data sets and generate text and other artifacts in response to new information.
[0015] "Audio data" refers to audio files containing recordings of statements by members of parliament and questions and answers.
[0016] "Text data" refers to data that has been analyzed and converted into text information.
[0017] A "summary" is data that extracts important points and information from text data and summarizes them in a concise form.
[0018] "Metrics" are a set of indicators or standards used to evaluate the performance and success of legislators.
[0019] A "score" refers to an evaluation value calculated based on metrics, and is a numerical representation of a member of parliament's activity and achievements.
[0020] A "report card" is a document that provides an easy-to-read summary of the results of a comprehensive evaluation of each member of parliament's activities and achievements.
[0021] "Multilingual translation" refers to the process of converting a given text into different languages.
[0022] "Dashboard" means the interface through which a User can access the System and enter and view information.
[0023] "Message" refers to text data of feedback or opinions sent by users to legislators.
[0024] "Social media" refers to platforms for sharing information online and communicating with other users. [Brief explanation of the drawings]
[0025] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0026] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0027] First, the terms used in the following description will be explained.
[0028] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0029] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0030] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0031] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0032] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0033] [First embodiment]
[0034] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0035] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0036] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0037] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0038] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0039] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0040] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0041] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0042] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0043] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0044] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0045] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0046] This invention is a system that uses generative artificial intelligence to summarize the speeches and Q&A sessions of Diet members, and evaluates their activity and performance based on the summaries. This system provides integrated functions for voice data conversion, summary generation, scoring, multilingual translation, message sending, and social media sharing.
[0047] Overall system overview
[0048] The system mainly consists of the following components:
[0049] 1. Voice data collection terminal
[0050] 2. Speech Recognition Server
[0051] 3. Abstract Generation Server
[0052] 4. Scoring Server
[0053] 5. Multilingual Translation Server
[0054] 6. Message sending function
[0055] 7. Dashboard
[0056] 8. Social Media Integration
[0057] Audio data collection and conversion
[0058] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal.
[0059] The device sends the uploaded voice data to a speech recognition server, which uses generative artificial intelligence to convert the voice data into text data. The text data returned from the speech recognition server is sent to a summary generation server.
[0060] Summary Generation
[0061] The server receives the text data and generates a summary of the text data using generative artificial intelligence (generative artificial intelligence), using a BERT-based natural language processing model, and then sends the summarized text data to a scoring server.
[0062] Scoring and report card generation
[0063] The server analyzes the summary data and calculates scores based on multiple metrics (e.g., number of comments, number of proposals, number of votes) that evaluate the activity and performance of each member of parliament. A score is generated based on each metric, and an overall evaluation score is calculated. A document visually summarizing these evaluation results is generated as a report card and saved in a database.
[0064] Multilingual Translation
[0065] The server translates report cards and summary data into multiple languages using generative artificial intelligence models or cloud-based translation services (e.g., Google Translate API). The translation results are displayed on a dashboard for users to view.
[0066] Sending a message
[0067] Users input messages to their legislators on the dashboard. The device then sends the input messages to the generative AI, which analyzes the contents of the messages. Based on the analysis results, the messages are sent to the legislators in an appropriate format.
[0068] Share on social media
[0069] To share a report card or score on social media, a user clicks the share button from the dashboard. The device generates a link and a share message and posts it using the social media API.
[0070] Specific examples
[0071] For example, a speech and question-and-answer session by a member of the National Diet is recorded and uploaded to the system as audio data. The speech is converted into text by a speech recognition server, and then summarized by a summary generation server as "a speech in the Diet regarding an increase in the education budget." The scoring server then evaluates the impact of the speech and adds the summarized data to metrics related to education policy. As a result, the member's score for education policy is calculated as 85 points, which is reflected in his / her report card.
[0072] In this way, by using the system of the present invention, it is possible to increase the transparency of the activities and statements of Diet members and promote political participation by ordinary voters.
[0073] The processing flow will be explained below.
[0074] Step 1:
[0075] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal by accessing the system's dedicated upload page, selecting the audio file, and clicking the upload button.
[0076] Step 2:
[0077] The device sends the uploaded voice data to the voice recognition server. At the same time as receiving the voice data, the device sends the voice file using the voice recognition API. This API converts the voice data into text.
[0078] Step 3:
[0079] The server analyzes the voice data received through the voice recognition API and converts it into text data. The voice recognition server then converts the spoken content into text using a matching algorithm, and sends the converted text data to the summary generation server.
[0080] Step 4:
[0081] The server receives the text data at the summary generation server and generates a summary of the text data using generative artificial intelligence. The summary generation server activates the generative artificial intelligence model, extracts important points and keywords, and generates a summary sentence.
[0082] Step 5:
[0083] The server sends the summarized text data to a scoring server, which analyzes metrics that evaluate the legislator's activity and performance. The scoring server analyzes the content, frequency, and activity of each statement, and calculates a score for each metric. The calculated scores are stored in a database.
[0084] Step 6:
[0085] The server generates a report card with the evaluation of each legislator. The scoring server calculates an overall evaluation score for each legislator based on the saved scores. The evaluation scores are compiled into a visually easy-to-understand report card and recorded in a database.
[0086] Step 7:
[0087] The server translates the generated report cards and summary data into multiple languages. The multilingual translation server uses generative artificial intelligence or a cloud translation API to convert data into different languages, and stores the translation results in a database. The results are then reflected on the dashboard.
[0088] Step 8:
[0089] Users access the dashboard and enter messages to their legislators by entering text directly into the message entry form on the dashboard and clicking the send button.
[0090] Step 9:
[0091] The terminal sends the input message to the generative AI, which analyzes the message content. The generative AI analyzes the message and sends the content to the legislator in an appropriate format.
[0092] Step 10:
[0093] Users can share their report cards and scores on social media by clicking the share button on the dashboard. When a user clicks the share button, the device generates a link and a share message and posts it via the corresponding social media API.
[0094] Through the above processing steps, a system will be realized in which the activities and statements of Diet members are managed transparently, leading to a deeper understanding among ordinary voters.
[0095] Example 1
[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0097] In today's political environment, there is a need to increase transparency in the statements and activities of Diet members and provide reliable information that allows voters to evaluate them. However, the current situation makes it difficult to obtain detailed information about Diet member statements and question and answer sessions, which hinders fair evaluation of Diet members' activities and achievements. Furthermore, opportunities for political participation are limited by a lack of information provision in multiple languages, voter feedback, and the means to share that information on social media.
[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0099] In this invention, the server includes means for converting collected voice data into text data, means for automatically summarizing the text data, means for calculating scores by analyzing multiple indicators for evaluating the activities and achievements of Diet members based on the summarized data, means for generating a report card evaluation of each Diet member based on the calculated score, means for translating the generated report card and summary data into multiple languages, means for voters to input and send messages to Diet members via a dashboard, means for sharing the report card and summary data on social media, means for translating data into multiple languages using a cloud-based translation service, means for posting data using a social media API, and means for providing dedicated terminals for uploading voice data. This increases the transparency of Diet members' statements and activities, enabling voters to properly evaluate Diet members, obtain information in multiple languages, send feedback, and further share that information widely.
[0100] "Generative AI" is an AI system that uses natural language processing technology to analyze and generate data.
[0101] "Audio data" refers to data recorded in digital format containing audio, including statements by members of parliament and question and answer sessions.
[0102] "Text data" refers to data obtained by converting voice data into character information.
[0103] A "summary" is a document that extracts key information from long text data and summarizes it in a short form.
[0104] "Multiple indicators" are criteria for evaluating a member of parliament's activity and achievements, such as the number of times they speak, the number of proposals they make, and the number of votes they pass.
[0105] The "score" is a numerical evaluation of a member of parliament's activity and achievements based on multiple indicators.
[0106] A "report card" is a document that visually summarizes a legislator's evaluation, including calculated scores.
[0107] "Multilingual translation" is the process of converting data or documents into multiple languages.
[0108] A "dashboard" is an interface that allows voters to view and interact with information.
[0109] "Social media" is an online platform for sharing information and interacting.
[0110] A "voice recognition model" is an algorithm for converting voice data into text data.
[0111] A "cloud-based translation service" is a service that uses translation functions provided via the Internet.
[0112] A "social media API" is an interface for accessing and operating social media functions from outside.
[0113] A "dedicated terminal" is a specific hardware device used to upload audio data.
[0114] This invention is a system that uses generative artificial intelligence to summarize the speeches and questions and answers of Diet members, and evaluates their activity and performance based on the summaries. This system integrates the following main components and functions:
[0115] Audio data collection and uploading
[0116] Users collect audio data of Diet deliberations using a recording device and upload it to a dedicated terminal. This dedicated terminal can be a general digital device such as a smartphone or PC. Users use these devices to record audio data and upload it to a voice recognition server on the cloud.
[0117] Converting audio data to text
[0118] The device sends the uploaded voice data to a voice recognition server. The server converts the voice data into text data using a voice recognition model (e.g., a cloud-based voice recognition service). Specifically, Google Cloud Speech-to-Text API is used. At this stage, the voice data is converted into text format.
[0119] Summarizing text data
[0120] The server sends the text data obtained from the speech recognition server to the summary generation server, which automatically summarizes the text data using a BERT-based natural language processing model (e.g., Hugging Face's transformers library). In this step, redundant information is removed, leaving the key information in a compact form.
[0121] Scoring summary data
[0122] The server receives the summary data and sends it to a scoring server. The scoring server then analyzes multiple indicators (e.g., number of statements, number of proposals, number of votes) that evaluate the activity and performance of each legislator based on the summary data, and calculates a score. Data analysis tools such as the Python pandas library are used for scoring.
[0123] Multilingual translation of scores and summary data
[0124] The server translates the calculated scores and summary data into multiple languages using a cloud-based translation service (e.g., Google Translate API). The translated data is then reflected on a dashboard, allowing users to view their assessment results in multiple languages.
[0125] Send a message to your legislators
[0126] Users can input feedback messages to legislators on the dashboard. The input messages are sent to the generative AI via the terminal, where the content is analyzed. The generative AI analyzes the content of the messages and sends them to legislators in an appropriate format.
[0127] Share on social media
[0128] Users can share their assessment results and report cards on social media from the dashboard. When a user clicks the share button, the device generates a link and shareable text and posts it using a social media API (e.g., Twitter API).
[0129] Specific examples
[0130] For example, if you upload audio data containing a congressman's speech about the education budget, the system processes it as follows: First, it converts the audio data into text using a cloud-based speech recognition service, and then summarizes the text using a BERT-based natural language processing model. The scoring server then analyzes the summarized data and calculates a score (e.g., 85 points) for the congressman's education policy. These results are then translated into multiple languages and displayed on a dashboard.
[0131] Prompt Sentence Examples
[0132] "Use this speech recognition system to analyze recent questions and answers from members of Congress regarding the education budget. Generate a summary of the results and an evaluation of the members' performance, display it in multiple languages, and make the results shareable on social media."
[0133] In this way, the system of the present invention increases the transparency of the activities and statements of Diet members and makes it possible to promote political participation among ordinary voters.
[0134] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0135] Step 1:
[0136] Users collect audio data of Diet deliberations and upload it to a dedicated device. Specifically, users use smartphones or digital recorders to record Diet members' remarks and Q&A sessions, and then upload the audio files to the device. The input is the audio data (e.g., a Diet member's statement, "We need to increase the education budget"), and the output is the saving of the audio file on the device.
[0137] Step 2:
[0138] The device sends the uploaded voice data to a voice recognition server on the cloud. The input is the voice data stored on the device, and the output is the data transfer to the voice recognition server. Specifically, the device uploads the voice data to the server via an internet connection.
[0139] Step 3:
[0140] The server uses a speech recognition model (e.g., a cloud-based speech recognition service) to convert the voice data into text data. The input is voice data, and the output is text data (e.g., "We need to increase the education budget"). Specifically, it uses services such as the Google Cloud Speech-to-Text API to perform highly accurate voice analysis.
[0141] Step 4:
[0142] The server transmits the text data acquired from the speech recognition server to the summary generation server. The input is text data, and the output is data transfer to the summary generation server. Specifically, data is transferred securely between the servers.
[0143] Step 5:
[0144] The server uses a summary generation server to automatically summarize text data using a BERT-based natural language processing model. The input is text data, and the output is summarized text data (e.g., "An increase in the education budget was discussed"). Specifically, natural language processing is performed using the Hugging Face transformers library.
[0145] Step 6:
[0146] The server sends the summary data to the scoring server, which analyzes multiple indicators and calculates a score. The input is the summary data, and the output is a score based on each indicator (e.g., 85 points). Specifically, analysis is performed using Python's pandas library based on data such as the number of comments, number of proposals, and number of votes.
[0147] Step 7:
[0148] The server translates the generated scores and summary data into multiple languages. The input is the scores and summary data, and the output is the translated data. Specifically, it uses a cloud-based translation service such as Google Translate API. The translated data is reflected in a dashboard and can be viewed by users.
[0149] Step 8:
[0150] The user inputs a feedback message to the legislator on the dashboard. The input is the message entered by the user on the dashboard (e.g., "I agree with the opinion on the education budget"), and the output is the message sent to the generative AI model. Specifically, the generative AI model analyzes the content of the message and sends it to the legislator in an appropriate format.
[0151] Step 9:
[0152] Users can share their assessment results and report cards on social media from the dashboard. The input is a click on the share button, and the output is a link and text for sharing. Specifically, the device generates the link and text for sharing and posts it to social media using the Twitter API or similar.
[0153] This series of processes will increase transparency in the statements and activities of Diet members, enable voters to properly evaluate their members, obtain information in multiple languages, send feedback, and share that information widely.
[0154] (Application example 1)
[0155] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0156] Since there is no means to evaluate the driving performance of autonomous vehicles, it is difficult to analyze and evaluate the appropriateness of decisions and actions while driving. Conventional systems do not adequately collect and analyze voice data and behavioral data while driving, making it difficult to improve autonomous driving technology and evaluate driver performance.
[0157] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0158] In this invention, the server includes means for converting collected voice data into text data using generative artificial intelligence, means for automatically summarizing the text data, means for calculating a score by analyzing multiple measurement indicators for evaluating the activity level and achievements of the subject based on the summarized data, means for generating a visualized document of the subject's evaluation based on the calculated score, means for translating the generated evaluation document and summary data into multiple languages, means for a user to input and send a message to the subject through a dashboard, means for sharing the evaluation document and summary data on social media, means for collecting driving data, analyzing and evaluating the data, and means for providing a management screen for visualizing and displaying the evaluation results. This makes it possible to effectively evaluate the driving performance of autonomous vehicles and provide transparent information to stakeholders.
[0159] "Generative AI" is a type of AI that has the ability to learn patterns based on large amounts of data and generate or predict new data.
[0160] "Text data" refers to character string data that has been converted into a natural language format from voice data or other input data.
[0161] "Summarization" refers to extracting important information from the original text data and presenting it in a short, concise form.
[0162] "Metrics" are the standards or measures used to evaluate the activities and outcomes of an evaluation.
[0163] A "score" is the result of quantifying the performance of an evaluation target based on measurement indicators.
[0164] A "visualized document" is a document that visually displays information such as evaluation results using graphs and charts.
[0165] "Multilingual translation" is the process of converting text data or evaluation documents into multiple languages.
[0166] A "dashboard" is an interface that allows users to enter information and view results.
[0167] A "message" is a text-based communication sent by a user to a target of evaluation.
[0168] "Social media" refers to online platforms for sharing information and engaging in two-way communication with the community.
[0169] "Driving data" refers to data about the vehicle's movements and surrounding conditions that is collected by an autonomous vehicle while it is driving.
[0170] The "management screen" is an interface for visually displaying and managing information such as evaluation results.
[0171] This invention relates to a system for evaluating the driving performance of autonomous vehicles. This system analyzes speech data and driving data, and provides the function of summarizing and evaluating the results. The entire system consists of the following components:
[0172] Hardware Configuration
[0173] 1. Voice data collection terminal
[0174] The device has a built-in microphone that collects conversations and instructions that occur inside the self-driving vehicle.
[0175] 2. Speech Recognition Server
[0176] The collected voice data is sent to this server and converted into text data using voice recognition technology.
[0177] 3. Abstract Generation Server
[0178] Text data received from a speech recognition server is summarized using generative artificial intelligence.
[0179] 4. Scoring Server
[0180] Based on the summarized data, multiple measurement indicators are analyzed to evaluate driving performance and a score is calculated.
[0181] 5. Multilingual Translation Server
[0182] Translate the generated summary and evaluation documents into multiple languages.
[0183] 6. Message sending function
[0184] Users enter messages through a dashboard, and generative artificial intelligence converts them into the appropriate format and sends them.
[0185] 7. Dashboard
[0186] This interface is used to visually display and manage assessment results and summaries.
[0187] 8. Social Media Integration
[0188] Users can easily share assessment results and summaries on social media from the dashboard.
[0189] Software Configuration
[0190] 1. Generative AI Model
[0191] It uses Hugging Face's BART model and a BERT-based natural language processing model.
[0192] 2. Voice Recognition Software
[0193] Convert audio data into text data using the Google Speech Recognition API or similar.
[0194] 3. Translation Services
[0195] Use a cloud-based translation service such as the Google Translate API.
[0196] 4. Data Visualization Tools
[0197] Display the evaluation results as a graph using seaborn and matplotlib.
[0198] Specific examples
[0199] For example, a voice command for an autonomous vehicle to turn left at an intersection is collected. This voice data is converted into text data by a speech recognition server, and summarized as a "left turn command" by a summary generation server. The scoring server then evaluates the appropriateness of this command and calculates an "accuracy score of 85 points." The evaluation results are translated into other languages by a multilingual translation server and displayed on a dashboard.
[0200] Prompt Sentence Examples
[0201] For example, use the following prompt:
[0202] "Collect audio data while driving, summarize and evaluate the data, translate the evaluation results into other languages, send them by email, and display the evaluation results as graphs."
[0203] In this way, by using the system of the present invention, it is possible to effectively evaluate the driving performance of an autonomous vehicle and provide transparent information.
[0204] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0205] Step 1:
[0206] Audio data collection
[0207] The device collects audio data using a microphone installed inside the autonomous vehicle.
[0208] Input: Voice data of conversations and instructions inside the vehicle
[0209] Output: Collected audio data file
[0210] Step 2:
[0211] Audio data conversion
[0212] The server converts the collected voice data files into text data using voice recognition software (e.g., Google Speech Recognition API).
[0213] Input: Audio data file
[0214] Output: Converted text data
[0215] What it does: Speech recognition software analyzes the audio waveform and outputs the corresponding text.
[0216] Step 3:
[0217] Summarizing text data
[0218] The server uses a generative artificial intelligence model (e.g., the BART model of Hugging Face) to summarize the text data.
[0219] Input: Converted text data
[0220] Output: Summary text
[0221] How it works: A generative artificial intelligence model extracts important parts of text data and generates a concise summary.
[0222] Step 4:
[0223] Driving performance scoring
[0224] The server evaluates driving performance using a BERT-based model based on the summarized text data and calculates a score.
[0225] Input: Summary text
[0226] Output: Performance score
[0227] How it works: A BERT-based model analyzes the summarized text and generates a score based on the evaluation metrics.
[0228] Step 5:
[0229] Evaluation document generation and multilingual translation
[0230] The server generates an evaluation document based on the calculated score and translates it into multiple languages using the Google Translate API or similar.
[0231] Input: Performance score, summary text
[0232] Output: Evaluation documents translated into multiple languages
[0233] Specific operation: Automatically generate evaluation documents and convert them into other languages using the translation API.
[0234] Step 6:
[0235] Sending a message
[0236] Users input messages to the subject of evaluation through the dashboard, and the generative artificial intelligence analyzes them and sends the messages in an appropriate format.
[0237] Input: The message entered by the user
[0238] Output: Message sent
[0239] Specific operation: Generative AI analyzes the content of the message, converts it into a format appropriate for the recipient, and then sends it.
[0240] Step 7:
[0241] Displaying evaluation results on the dashboard
[0242] The server visually displays the evaluation results on a dashboard.
[0243] Input: Performance score and summary text
[0244] Output: Visualized evaluation results (graphs and charts)
[0245] Specific operation: Use data visualization tools such as seaborn and matplotlib to display the evaluation results as graphs and charts.
[0246] Step 8:
[0247] Social media integration
[0248] Users can share evaluation documents and summary text on social media from the dashboard, and the device posts using social media APIs.
[0249] Input: Evaluation document and summary text
[0250] Output: Social media posts
[0251] Specific operation: Evaluation results are automatically posted via social media API.
[0252] The above processing steps make it possible to effectively evaluate the driving performance of an autonomous vehicle and provide transparent information.
[0253] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0254] This invention is a system that uses generative artificial intelligence to summarize the speeches and Q&A sessions of Diet members, and evaluates their activity levels and achievements based on those summaries. In particular, by combining it with an emotion engine that recognizes and analyzes the user's emotions, this invention realizes message exchange that reflects the user's emotional expressions.
[0255] Overall system overview
[0256] The system mainly consists of the following components:
[0257] 1. Voice data collection terminal
[0258] 2. Speech Recognition Server
[0259] 3. Abstract Generation Server
[0260] 4. Scoring Server
[0261] 5. Multilingual Translation Server
[0262] 6. Messaging and emotion engine
[0263] 7. Dashboard
[0264] 8. Social Media Integration
[0265] Audio data collection and conversion
[0266] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal by accessing the system's dedicated upload page, selecting the audio file, and clicking the upload button.
[0267] The terminal sends the uploaded voice data to a voice recognition server, which converts the voice data into text data and sends it to a summary generation server.
[0268] Summary Generation
[0269] The server receives the text data and generates a summary of the text data using generative artificial intelligence. The summary is generated using a BERT-based natural language processing model. The generated summary data is sent to a scoring server.
[0270] Scoring and report card generation
[0271] The server analyzes the summary data and calculates a score based on multiple metrics (e.g., number of comments, number of proposals, number of votes) that evaluate the activity and performance of each member of parliament. Based on the analyzed scores, a report card is generated and stored in a database.
[0272] Multilingual Translation
[0273] The server translates report cards and summary data into multiple languages using generative artificial intelligence models or cloud-based translation services (e.g., Google Translate API). The translation results are stored in a database and displayed on a dashboard.
[0274] Messaging and Sentiment Analysis
[0275] Users can input messages to lawmakers on the dashboard. As they input their messages, the emotion engine analyzes their emotions and reflects the results in the message content.
[0276] The device sends the input message and analyzed emotion data to the generative AI, which analyzes the message and sends it to the legislator in an appropriate format. The emotion data is used by the legislator to understand the user's emotions and suggest an appropriate reply.
[0277] Message examples
[0278] For example, when a user types, "I strongly desire an increase in the education budget," the emotion engine recognizes from the phrase "strongly desire" that the user has a high positive emotion. This emotion data is parsed as part of the message content and sent to the legislator. The legislator receives feedback with a high positive emotion and generates an appropriate reply based on that emotion.
[0279] Share on social media
[0280] Users select a report card or score from the dashboard and click the share button, and the device generates a link and share message that is posted via social media APIs.
[0281] As described above, the present invention is a system that transparently manages the activities and statements of Diet members, and further promotes richer communication by recognizing the user's emotions.
[0282] The processing flow will be explained below.
[0283] Step 1:
[0284] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal by accessing the system's dedicated upload page, selecting the audio file, and clicking the upload button.
[0285] Step 2:
[0286] The device sends the uploaded voice data to the voice recognition server. At the same time, the device receives the voice data and sends the voice file using the voice recognition API. This API converts the voice data into text.
[0287] Step 3:
[0288] The server analyzes the voice data received through the voice recognition API and converts it into text data. The voice recognition server converts the voice data into text data and sends the converted text data to the summary generation server.
[0289] Step 4:
[0290] The server receives text data at the summary generation server and generates a summary of the text data using generative artificial intelligence. The summary generation server activates the generative artificial intelligence model, extracts important points and keywords, and generates a summary. The generated summary data is sent to the scoring server.
[0291] Step 5:
[0292] The server sends the summarized text data to a scoring server, which analyzes metrics that evaluate the legislator's activity and performance. The scoring server analyzes the content, frequency, and activity of each statement, and calculates a score for each metric. The calculated scores are stored in a database.
[0293] Step 6:
[0294] The server generates a report card with the evaluation of each legislator. The scoring server calculates an overall evaluation score for each legislator based on the saved scores. The evaluation scores are compiled into a visually easy-to-understand report card and recorded in a database.
[0295] Step 7:
[0296] The server translates the generated report cards and summary data into multiple languages. The multilingual translation server uses generative artificial intelligence or a cloud translation API to convert data into different languages. The translation results are stored in a database and then reflected on the dashboard.
[0297] Step 8:
[0298] Users access the dashboard and input messages to legislators. When inputting a message, the emotion engine analyzes the user's emotions and reflects the results in the message content. Users enter text in the message input form on the dashboard and click the send button.
[0299] Step 9:
[0300] The device sends the input message and analyzed emotion data to the generative AI, which analyzes the message and sends the content to the legislator in an appropriate format. The emotion data is used by the legislator to understand the user's emotion and generate a reply accordingly.
[0301] Step 10:
[0302] Users can share their report cards or scores on social media by clicking the share button on the dashboard. The device generates a link and a share message and posts it via the social media API.
[0303] Through this series of processing steps, a system will be realized that transparently manages the activities and statements of Diet members, deepening the understanding of ordinary voters. In addition, the addition of an emotion engine will allow users' emotions to be reflected in messages, enabling richer communication.
[0304] Example 2
[0305] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0306] In conventional systems for evaluating the activities of legislators, the collection and analysis of voice data is performed manually, making it difficult to efficiently process large amounts of data. Furthermore, there are few ways for voters to provide direct, emotional feedback to legislators, preventing improvements in transparency and deeper communication. This has led to challenges in the accuracy of legislator activity evaluations and two-way communication with voters.
[0307] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0308] In this invention, the server includes: means for converting collected voice data into text data using generative artificial intelligence; means for automatically summarizing the text data; means for calculating scores by analyzing multiple metrics that evaluate the activity and performance of legislators based on the summarized data; means for translating the generated report card and summary data into multiple languages; means for analyzing messages from users using an emotion analysis engine and reflecting the emotion data in messages to legislators; means for voters to input and send messages to legislators via a dashboard; means for the generative artificial intelligence to analyze the input message and emotion data and send them to legislators in an appropriate format; and means for sharing the report card and summary data on social media. This enables efficient analysis and summarization of voice data, automated evaluation of legislator activity, and two-way communication with voters that reflects their emotions.
[0309] "Generative AI" refers to AI that has the ability to generate new information and data based on large amounts of data.
[0310] "Audio data" refers to recorded sounds or speech stored in digital format.
[0311] "Text data" refers to digital data that has been converted from audio data into text information.
[0312] A "summary" refers to information that has been shortened by extracting important information from long text data.
[0313] "Metrics" refers to the indicators and standards used to evaluate the activities of lawmakers.
[0314] "Sentiment analysis engine" refers to software or technology for analyzing and recognizing emotions from input text or voice.
[0315] "Dashboard" refers to the interface that allows system users to access and operate data and functions.
[0316] A "report card" is a report that evaluates a member of parliament's activities and achievements and compiles information such as scores.
[0317] "Multilingual translation" refers to the process or function of converting text from one language into another.
[0318] "Social media" refers to a platform on the Internet that allows users to share and interact with each other.
[0319] This invention is a system that uses generative artificial intelligence to summarize the speeches and Q&A sessions of legislators, and evaluates their activity levels and achievements based on the summaries. In particular, by combining it with an emotion engine that recognizes and analyzes the user's emotions, the system realizes message exchanges that reflect the user's emotional expressions.
[0320] System configuration
[0321] The system consists of the following components:
[0322] 1. Voice data collection terminal
[0323] 2. Speech Recognition Server
[0324] 3. Abstract Generation Server
[0325] 4. Scoring Server
[0326] 5. Multilingual Translation Server
[0327] 6. Messaging and emotion engine
[0328] 7. Dashboard
[0329] 8. Social Media Integration
[0330] Hardware and software used
[0331] The Google Cloud Speech-to-Text API is used as the speech recognition server.
[0332] The summary generation server uses a BERT-based natural language processing model.
[0333] The multilingual translation server uses the Google Translate API.
[0334] The Sentiment Analysis tool is used for sentiment analysis.
[0335] Audio data collection
[0336] Users record audio data of Diet deliberations and upload it to an audio data collection terminal. This process is performed using a dedicated upload page. Users access the upload page, select the audio file, and click the "Upload" button.
[0337] Converting audio data to text
[0338] The device sends the uploaded voice data to the speech recognition server, which receives the voice data and converts it into text data using the Google Cloud Speech-to-Text API, which then sends the converted text data to the summary generation server.
[0339] Summary Generation
[0340] The server receives the text data and generates a summary of the text data using generative artificial intelligence (e.g., a BERT-based natural language processing model). The summary generation process extracts important statements and keywords and shortens them while preserving the context. The generated summary data is sent to the scoring server.
[0341] Activity evaluation and scoring
[0342] The server analyzes the summary data to evaluate the activity and performance of each member of parliament (e.g., number of comments, number of proposals, number of votes). It calculates a score based on the analyzed data and generates a report card based on that score. The calculated report card is stored in a database.
[0343] Multilingual Translation
[0344] The server translates the generated report cards and summary data into multiple languages using the Google Translate API. The translated data is stored in a database and displayed on the dashboard.
[0345] Sentiment analysis and messaging
[0346] The user inputs a message to the legislator on the dashboard. At this time, the emotion engine analyzes the input message and recognizes the user's emotions. For example, Sentiment Analysis can be used to identify "highly positive emotions" and "negative emotions." The device then sends the analyzed emotion data to the generative AI, which then analyzes the message and sends it to the legislator in an appropriate format.
[0347] Message examples
[0348] For example, if a user types, "I strongly desire an increase in the education budget," the emotion engine recognizes from the phrase "strongly desire" that the user has a high positive emotion. This emotion data is parsed as part of the message content and sent to the legislator. The legislator receives feedback with a high positive emotion and generates an appropriate reply based on that emotion.
[0349] Share on social media
[0350] Users select a report card or score from the dashboard and click the share button, and the device generates a link and a share message, which is posted via social media APIs (such as Twitter API or Facebook API).
[0351] Example prompts for generative AI models
[0352] The user can provide prompts to the generative AI model, such as:
[0353] "We will create a report card to evaluate the activities of politicians. We will convert audio data into text, summarize the text, and analyze the number of statements and proposals made by politicians. As an example, please summarize and evaluate the following statements:
[0354] "There have been many problems with education in Japan over the past few years. I strongly hope that the education budget will be increased."
[0355] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0356] Step 1:
[0357] Users record audio data of Diet deliberations and upload it to the audio data collection terminal by accessing the upload page, selecting the audio file, and clicking the "Upload" button. This inputs the audio data into the system.
[0358] Step 2:
[0359] The device sends the uploaded voice data to the speech recognition server, which receives the voice data and converts it into text data using the Google Cloud Speech-to-Text API. Through this process, the voice data is converted into text data.
[0360] Step 3:
[0361] The server receives the text data and generates a summary of the text data using a BERT-based natural language processing model. The summary generation process extracts important information from the text data and generates a compact summary. The generated summary data is then sent to the scoring server.
[0362] Step 4:
[0363] The server analyzes the summary data to evaluate the activity and performance of each member of parliament using multiple metrics (e.g., number of comments, number of proposals, number of votes). It calculates a score based on the analyzed data and generates a report card based on that score. This report card is stored in a database.
[0364] Step 5:
[0365] The server translates the generated report cards and summary data into multiple languages. The generated data is converted into other languages using the Google Translate API, etc. The translated data is stored in a database and displayed on a dashboard.
[0366] Step 6:
[0367] Users input messages to their legislators on the dashboard. The input message is analyzed by a sentiment analysis engine to identify the user's sentiment. For example, natural language processing is used to identify "positive" or "negative" sentiment. This sentiment data is reflected in the message content and becomes input for analysis by the generative artificial intelligence.
[0368] Step 7:
[0369] The server receives messages containing emotional data obtained from the emotion analysis engine, and the generative AI analyzes them and sends them to the legislator in an appropriate format, allowing the legislator to understand the user's emotions and generate a reply accordingly.
[0370] Step 8:
[0371] Users select the generated report card or score from the dashboard and click the share button. The device generates a link and a share message and posts it via social media APIs (e.g., Twitter API, Facebook API). This process allows feedback, including the legislator's activity evaluation and the user's sentiment, to be widely shared.
[0372] (Application example 2)
[0373] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0374] It is difficult to grasp, evaluate, and improve the efficiency and results of work activities at work sites such as factories. Furthermore, in environments with multinational workers, language barriers are a major obstacle to communication and supervision. Furthermore, it is difficult to exchange appropriate feedback between on-site supervisors and workers in real time.
[0375] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for converting collected voice data into text data using generative artificial intelligence, means for automatically summarizing the text data, and means for calculating scores by analyzing multiple metrics for evaluating the efficiency and results of work activities based on the summarized data. This enables objective evaluation of work efficiency and results within a factory. The generative artificial intelligence also includes means for generating a report card evaluation of work activities based on the summarized data and translating it into multiple languages, means for the work supervisor to input and send feedback messages to workers via a display terminal, and means for storing the report card and summary data in a data management system. This enables communication across language barriers and appropriate feedback in real time.
[0376] "Generative AI" refers to AI that automatically generates, analyzes, and generates summaries and evaluations based on collected data.
[0377] "Voice data" refers to audio information about work collected in factories, work sites, etc.
[0378] "Text data" refers to character information converted from voice data using a voice recognition model.
[0379] A "summary" refers to content that has been created by using generative artificial intelligence to shorten text data and extract only the important information.
[0380] "Work activities" refers to specific work processes involving workers and robots in factories or on-site.
[0381] "Efficiency" refers to an indicator of how effectively work activities are being carried out.
[0382] "Outcomes" refers to the products obtained as a result of work activities or indicators of success.
[0383] "Metrics" refers to specific indicators or benchmarks for evaluating the efficiency and results of work activities.
[0384] "Score" refers to a numerical rating calculated based on metrics.
[0385] A "report card" refers to an evaluation report prepared based on calculated scores.
[0386] "Multilingual translation" refers to translating generated text data or report card contents into multiple languages.
[0387] "Display terminal" refers to a device for visually displaying data, such as smart glasses.
[0388] A "feedback message" refers to a message sent by a work supervisor to a worker for evaluation or guidance.
[0389] A "data management system" refers to a system that stores report cards and summary data and manages them so that they can be viewed later.
[0390] This invention is a system for improving the efficiency and evaluation of work activities in factories and on-site. It uses generative artificial intelligence to convert collected voice data into text data and automatically summarize it. The efficiency and results of work activities are then evaluated based on the summarized data. One specific embodiment of the invention is described in detail below.
[0391] Hardware and Software Configuration
[0392] The server includes the following hardware and software:
[0393] Speech recognition server: Converts collected voice data into text data. For example, the speech recognition server uses a "speech recognition model (e.g., Google Speech-to-Text API)."
[0394] Summarization server: Automatically summarizes text data. As a specific example, the summary generation server uses a generative AI model (e.g., a BERT-based natural language processing model).
[0395] Scoring Server: Evaluates the efficiency and performance of work activities based on summary data and calculates a score.
[0396] Multilingual translation server: Translates the generated data into multiple languages. A specific example is a cloud-based translation service (e.g., Google Translate API).
[0397] Data management system: Stores report cards and summary data and manages them for later review.
[0398] The display terminal is a device such as smart glasses. As a specific example, we will use "smart glasses (e.g., Microsoft (registered trademark) HoloLens (registered trademark))."
[0399] Data processing and calculation processing
[0400] Audio data collection and conversion:
[0401] The user (work supervisor) uses smart glasses to collect voices while working on-site. The voice data is transmitted to a voice recognition server via wireless communication (e.g., Bluetooth, Wi-Fi).
[0402] The server (speech recognition server) uses a speech recognition model to convert the voice data into text data, using a generative AI model.
[0403] Summary generation:
[0404] The server (summarization server) summarizes text data using a generative artificial intelligence model, which utilizes a BERT-based natural language processing model.
[0405] Scoring and Rating:
[0406] The server (scoring server) analyzes metrics that evaluate the efficiency and results of work activities based on the summary data and calculates a score.
[0407] The calculated scores are generated as a report card and stored in a data management system.
[0408] Multilingual Translation:
[0409] The server (multilingual translation server) translates report cards and summary data into multiple languages and provides them to foreign workers as needed.
[0410] Send a feedback message:
[0411] The user (work supervisor) uses a display terminal (smart glasses) to input and send feedback messages to the workers.
[0412] The server (generative artificial intelligence) analyzes the feedback message and sends it to the worker in an appropriate format.
[0413] Examples of concrete examples and prompts
[0414] As a concrete example, the following prompts are used by a supervisor wearing smart glasses to record the robot's movements and the workers' tasks on-site:
[0415] Example prompt (summary generation): "Please summarize this text data: [text data from speech recognition results]"
[0416] Example prompt (sentiment analysis): "Analyze the sentiment of this message: [Supervisor's feedback]"
[0417] Example prompt (multilingual translation): "Please translate this text into [language]: [summary data / assessment results]"
[0418] In this way, the present invention provides improved efficiency, effective evaluation, and appropriate feedback at the workplace.
[0419] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0420] Step 1:
[0421] Collection and transmission of voice data
[0422] The user uses smart glasses to collect voices from the field. The collected voice data is sent to a voice recognition server via wireless communication (Bluetooth or Wi-Fi). The input is the voice data from the field, and the output is the data sent to the voice recognition server.
[0423] Step 2:
[0424] Converting audio data to text
[0425] The server (speech recognition server) converts the received voice data into text data using a voice recognition model (e.g., Google Speech-to-Text API). The input is voice data, and the output is text data based on this.
[0426] Step 3:
[0427] Summarizing text data
[0428] The server (summarization server) automatically summarizes the converted text data using a generative AI model (e.g., a BERT-based natural language processing model). The input is text data, and the output is summarized text data.
[0429] Step 4:
[0430] Work activity evaluation and scoring
[0431] The server (scoring server) analyzes metrics that evaluate the efficiency and results of work activities based on the summary data and calculates a score. The input is the summary data, and the output is the calculated score. The specific operation of scoring is to analyze multiple metrics such as the number of comments and the task completion rate.
[0432] Step 5:
[0433] Generate evaluation report cards and save data
[0434] The server (scoring server) generates a report card based on the calculated score and stores it in a data management system. The input is the score, and the output is the generated report card and its storage.
[0435] Step 6:
[0436] Multilingual Translation
[0437] The server (multilingual translation server) translates the generated report card and summary data into multiple languages. As a concrete example, we use the Google Translate API. The input is the report card and summary data, and the output is the translated data.
[0438] Step 7:
[0439] Enter and send your feedback message
[0440] The user (work supervisor) inputs a feedback message using a display terminal (smart glasses). The input message is sent to the generative AI model. The input is the user's feedback message, and the output is the message sent to the generative AI model.
[0441] Step 8:
[0442] Parsing and delivering feedback messages
[0443] The server (generative artificial intelligence) analyzes the feedback message and sends it to the worker in an appropriate format. The input is the analyzed feedback message, and the output is the appropriate message to the worker.
[0444] This series of steps enables the system to evaluate on-site work efficiency and provide feedback in real time.
[0445] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0446] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0447] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0448] [Second embodiment]
[0449] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0450] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0451] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0452] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0453] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0455] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0456] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0457] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0458] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0459] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0460] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0461] This invention is a system that uses generative artificial intelligence to summarize the speeches and Q&A sessions of Diet members, and evaluates their activity and performance based on the summaries. This system provides integrated functions for voice data conversion, summary generation, scoring, multilingual translation, message sending, and social media sharing.
[0462] Overall system overview
[0463] The system mainly consists of the following components:
[0464] 1. Voice data collection terminal
[0465] 2. Speech Recognition Server
[0466] 3. Abstract Generation Server
[0467] 4. Scoring Server
[0468] 5. Multilingual Translation Server
[0469] 6. Message sending function
[0470] 7. Dashboard
[0471] 8. Social Media Integration
[0472] Audio data collection and conversion
[0473] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal.
[0474] The device sends the uploaded voice data to a speech recognition server, which uses generative artificial intelligence to convert the voice data into text data. The text data returned from the speech recognition server is sent to a summary generation server.
[0475] Summary Generation
[0476] The server receives the text data and generates a summary of the text data using generative artificial intelligence (generative artificial intelligence), using a BERT-based natural language processing model, and then sends the summarized text data to a scoring server.
[0477] Scoring and report card generation
[0478] The server analyzes the summary data and calculates scores based on multiple metrics (e.g., number of comments, number of proposals, number of votes) that evaluate the activity and performance of each member of parliament. A score is generated based on each metric, and an overall evaluation score is calculated. A document visually summarizing these evaluation results is generated as a report card and saved in a database.
[0479] Multilingual Translation
[0480] The server translates report cards and summary data into multiple languages using generative artificial intelligence models or cloud-based translation services (e.g., Google Translate API). The translation results are displayed on a dashboard for users to view.
[0481] Sending a message
[0482] Users input messages to their legislators on the dashboard. The device then sends the input messages to the generative AI, which analyzes the contents of the messages. Based on the analysis results, the messages are sent to the legislators in an appropriate format.
[0483] Share on social media
[0484] To share a report card or score on social media, a user clicks the share button from the dashboard. The device generates a link and a share message and posts it using the social media API.
[0485] Specific examples
[0486] For example, a speech and question-and-answer session by a member of the National Diet is recorded and uploaded to the system as audio data. The speech is converted into text by a speech recognition server, and then summarized by a summary generation server as "a speech in the Diet regarding an increase in the education budget." The scoring server then evaluates the impact of the speech and adds the summarized data to metrics related to education policy. As a result, the member's score for education policy is calculated as 85 points, which is reflected in his / her report card.
[0487] In this way, by using the system of the present invention, it is possible to increase the transparency of the activities and statements of Diet members and promote political participation by ordinary voters.
[0488] The processing flow will be explained below.
[0489] Step 1:
[0490] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal by accessing the system's dedicated upload page, selecting the audio file, and clicking the upload button.
[0491] Step 2:
[0492] The device sends the uploaded voice data to the voice recognition server. At the same time as receiving the voice data, the device sends the voice file using the voice recognition API. This API converts the voice data into text.
[0493] Step 3:
[0494] The server analyzes the voice data received through the voice recognition API and converts it into text data. The voice recognition server then converts the spoken content into text using a matching algorithm, and sends the converted text data to the summary generation server.
[0495] Step 4:
[0496] The server receives the text data at the summary generation server and generates a summary of the text data using generative artificial intelligence. The summary generation server activates the generative artificial intelligence model, extracts important points and keywords, and generates a summary sentence.
[0497] Step 5:
[0498] The server sends the summarized text data to a scoring server, which analyzes metrics that evaluate the legislator's activity and performance. The scoring server analyzes the content, frequency, and activity of each statement, and calculates a score for each metric. The calculated scores are stored in a database.
[0499] Step 6:
[0500] The server generates a report card with the evaluation of each legislator. The scoring server calculates an overall evaluation score for each legislator based on the saved scores. The evaluation scores are compiled into a visually easy-to-understand report card and recorded in a database.
[0501] Step 7:
[0502] The server translates the generated report cards and summary data into multiple languages. The multilingual translation server uses generative artificial intelligence or a cloud translation API to convert data into different languages, and stores the translation results in a database. The results are then reflected on the dashboard.
[0503] Step 8:
[0504] Users access the dashboard and enter messages to their legislators by entering text directly into the message entry form on the dashboard and clicking the send button.
[0505] Step 9:
[0506] The terminal sends the input message to the generative AI, which analyzes the message content. The generative AI analyzes the message and sends the content to the legislator in an appropriate format.
[0507] Step 10:
[0508] Users can share their report cards and scores on social media by clicking the share button on the dashboard. When a user clicks the share button, the device generates a link and a share message and posts it via the corresponding social media API.
[0509] Through the above processing steps, a system will be realized in which the activities and statements of Diet members are managed transparently, leading to a deeper understanding among ordinary voters.
[0510] Example 1
[0511] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0512] In today's political environment, there is a need to increase transparency in the statements and activities of Diet members and provide reliable information that allows voters to evaluate them. However, the current situation makes it difficult to obtain detailed information about Diet member statements and question and answer sessions, which hinders fair evaluation of Diet members' activities and achievements. Furthermore, opportunities for political participation are limited by a lack of information provision in multiple languages, voter feedback, and the means to share that information on social media.
[0513] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0514] In this invention, the server includes means for converting collected voice data into text data, means for automatically summarizing the text data, means for calculating scores by analyzing multiple indicators for evaluating the activities and achievements of Diet members based on the summarized data, means for generating a report card evaluation of each Diet member based on the calculated score, means for translating the generated report card and summary data into multiple languages, means for voters to input and send messages to Diet members via a dashboard, means for sharing the report card and summary data on social media, means for translating data into multiple languages using a cloud-based translation service, means for posting data using a social media API, and means for providing dedicated terminals for uploading voice data. This increases the transparency of Diet members' statements and activities, enabling voters to properly evaluate Diet members, obtain information in multiple languages, send feedback, and further share that information widely.
[0515] "Generative AI" is an AI system that uses natural language processing technology to analyze and generate data.
[0516] "Audio data" refers to data recorded in digital format containing audio, including statements by members of parliament and question and answer sessions.
[0517] "Text data" refers to data obtained by converting voice data into character information.
[0518] A "summary" is a document that extracts key information from long text data and summarizes it in a short form.
[0519] "Multiple indicators" are criteria for evaluating a member of parliament's activity and achievements, such as the number of times they speak, the number of proposals they make, and the number of votes they pass.
[0520] The "score" is a numerical evaluation of a member of parliament's activity and achievements based on multiple indicators.
[0521] A "report card" is a document that visually summarizes a legislator's evaluation, including calculated scores.
[0522] "Multilingual translation" is the process of converting data or documents into multiple languages.
[0523] A "dashboard" is an interface that allows voters to view and interact with information.
[0524] "Social media" is an online platform for sharing information and interacting.
[0525] A "voice recognition model" is an algorithm for converting voice data into text data.
[0526] A "cloud-based translation service" is a service that uses translation functions provided via the Internet.
[0527] A "social media API" is an interface for accessing and operating social media functions from outside.
[0528] A "dedicated terminal" is a specific hardware device used to upload audio data.
[0529] This invention is a system that uses generative artificial intelligence to summarize the speeches and questions and answers of Diet members, and evaluates their activity and performance based on the summaries. This system integrates the following main components and functions:
[0530] Audio data collection and uploading
[0531] Users collect audio data of Diet deliberations using a recording device and upload it to a dedicated terminal. This dedicated terminal can be a general digital device such as a smartphone or PC. Users use these devices to record audio data and upload it to a voice recognition server on the cloud.
[0532] Converting audio data to text
[0533] The device sends the uploaded voice data to a voice recognition server. The server converts the voice data into text data using a voice recognition model (e.g., a cloud-based voice recognition service). Specifically, Google Cloud Speech-to-Text API is used. At this stage, the voice data is converted into text format.
[0534] Summarizing text data
[0535] The server sends the text data obtained from the speech recognition server to the summary generation server, which automatically summarizes the text data using a BERT-based natural language processing model (e.g., Hugging Face's transformers library). In this step, redundant information is removed, leaving the key information in a compact form.
[0536] Scoring summary data
[0537] The server receives the summary data and sends it to a scoring server. The scoring server then analyzes multiple indicators (e.g., number of statements, number of proposals, number of votes) that evaluate the activity and performance of each legislator based on the summary data, and calculates a score. Data analysis tools such as the Python pandas library are used for scoring.
[0538] Multilingual translation of scores and summary data
[0539] The server translates the calculated scores and summary data into multiple languages using a cloud-based translation service (e.g., Google Translate API). The translated data is then reflected on a dashboard, allowing users to view their assessment results in multiple languages.
[0540] Send a message to your legislators
[0541] Users can input feedback messages to legislators on the dashboard. The input messages are sent to the generative AI via the terminal, where the content is analyzed. The generative AI analyzes the content of the messages and sends them to legislators in an appropriate format.
[0542] Share on social media
[0543] Users can share their assessment results and report cards on social media from the dashboard. When a user clicks the share button, the device generates a link and shareable text and posts it using a social media API (e.g., Twitter API).
[0544] Specific examples
[0545] For example, if you upload audio data containing a congressman's speech about the education budget, the system processes it as follows: First, it converts the audio data into text using a cloud-based speech recognition service, and then summarizes the text using a BERT-based natural language processing model. The scoring server then analyzes the summarized data and calculates a score (e.g., 85 points) for the congressman's education policy. These results are then translated into multiple languages and displayed on a dashboard.
[0546] Prompt Sentence Examples
[0547] "Use this speech recognition system to analyze recent questions and answers from members of Congress regarding the education budget. Generate a summary of the results and an evaluation of the members' performance, display it in multiple languages, and make the results shareable on social media."
[0548] In this way, the system of the present invention increases the transparency of the activities and statements of Diet members and makes it possible to promote political participation among ordinary voters.
[0549] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0550] Step 1:
[0551] Users collect audio data of Diet deliberations and upload it to a dedicated device. Specifically, users use smartphones or digital recorders to record Diet members' remarks and Q&A sessions, and then upload the audio files to the device. The input is the audio data (e.g., a Diet member's statement, "We need to increase the education budget"), and the output is the saving of the audio file on the device.
[0552] Step 2:
[0553] The device sends the uploaded voice data to a voice recognition server on the cloud. The input is the voice data stored on the device, and the output is the data transfer to the voice recognition server. Specifically, the device uploads the voice data to the server via an internet connection.
[0554] Step 3:
[0555] The server uses a speech recognition model (e.g., a cloud-based speech recognition service) to convert the voice data into text data. The input is voice data, and the output is text data (e.g., "We need to increase the education budget"). Specifically, it uses services such as the Google Cloud Speech-to-Text API to perform highly accurate voice analysis.
[0556] Step 4:
[0557] The server transmits the text data acquired from the speech recognition server to the summary generation server. The input is text data, and the output is data transfer to the summary generation server. Specifically, data is transferred securely between the servers.
[0558] Step 5:
[0559] The server uses a summary generation server to automatically summarize text data using a BERT-based natural language processing model. The input is text data, and the output is summarized text data (e.g., "An increase in the education budget was discussed"). Specifically, natural language processing is performed using the Hugging Face transformers library.
[0560] Step 6:
[0561] The server sends the summary data to the scoring server, which analyzes multiple indicators and calculates a score. The input is the summary data, and the output is a score based on each indicator (e.g., 85 points). Specifically, analysis is performed using Python's pandas library based on data such as the number of comments, number of proposals, and number of votes.
[0562] Step 7:
[0563] The server translates the generated scores and summary data into multiple languages. The input is the scores and summary data, and the output is the translated data. Specifically, it uses a cloud-based translation service such as Google Translate API. The translated data is reflected in a dashboard and can be viewed by users.
[0564] Step 8:
[0565] The user inputs a feedback message to the legislator on the dashboard. The input is the message entered by the user on the dashboard (e.g., "I agree with the opinion on the education budget"), and the output is the message sent to the generative AI model. Specifically, the generative AI model analyzes the content of the message and sends it to the legislator in an appropriate format.
[0566] Step 9:
[0567] Users can share their assessment results and report cards on social media from the dashboard. The input is a click on the share button, and the output is a link and text for sharing. Specifically, the device generates the link and text for sharing and posts it to social media using the Twitter API or similar.
[0568] This series of processes will increase transparency in the statements and activities of Diet members, enable voters to properly evaluate their members, obtain information in multiple languages, send feedback, and share that information widely.
[0569] (Application example 1)
[0570] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0571] Since there is no means to evaluate the driving performance of autonomous vehicles, it is difficult to analyze and evaluate the appropriateness of decisions and actions while driving. Conventional systems do not adequately collect and analyze voice data and behavioral data while driving, making it difficult to improve autonomous driving technology and evaluate driver performance.
[0572] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0573] In this invention, the server includes means for converting collected voice data into text data using generative artificial intelligence, means for automatically summarizing the text data, means for calculating a score by analyzing multiple measurement indicators for evaluating the activity level and achievements of the subject based on the summarized data, means for generating a visualized document of the subject's evaluation based on the calculated score, means for translating the generated evaluation document and summary data into multiple languages, means for a user to input and send a message to the subject through a dashboard, means for sharing the evaluation document and summary data on social media, means for collecting driving data, analyzing and evaluating the data, and means for providing a management screen for visualizing and displaying the evaluation results. This makes it possible to effectively evaluate the driving performance of autonomous vehicles and provide transparent information to stakeholders.
[0574] "Generative AI" is a type of AI that has the ability to learn patterns based on large amounts of data and generate or predict new data.
[0575] "Text data" refers to character string data that has been converted into a natural language format from voice data or other input data.
[0576] "Summarization" refers to extracting important information from the original text data and presenting it in a short, concise form.
[0577] "Metrics" are the standards or measures used to evaluate the activities and outcomes of an evaluation.
[0578] A "score" is the result of quantifying the performance of an evaluation target based on measurement indicators.
[0579] A "visualized document" is a document that visually displays information such as evaluation results using graphs and charts.
[0580] "Multilingual translation" is the process of converting text data or evaluation documents into multiple languages.
[0581] A "dashboard" is an interface that allows users to enter information and view results.
[0582] A "message" is a text-based communication sent by a user to a target of evaluation.
[0583] "Social media" refers to online platforms for sharing information and engaging in two-way communication with the community.
[0584] "Driving data" refers to data about the vehicle's movements and surrounding conditions that is collected by an autonomous vehicle while it is driving.
[0585] The "management screen" is an interface for visually displaying and managing information such as evaluation results.
[0586] This invention relates to a system for evaluating the driving performance of autonomous vehicles. This system analyzes speech data and driving data, and provides the function of summarizing and evaluating the results. The entire system consists of the following components:
[0587] Hardware Configuration
[0588] 1. Voice data collection terminal
[0589] The device has a built-in microphone that collects conversations and instructions that occur inside the self-driving vehicle.
[0590] 2. Speech Recognition Server
[0591] The collected voice data is sent to this server and converted into text data using voice recognition technology.
[0592] 3. Abstract Generation Server
[0593] Text data received from a speech recognition server is summarized using generative artificial intelligence.
[0594] 4. Scoring Server
[0595] Based on the summarized data, multiple measurement indicators are analyzed to evaluate driving performance and a score is calculated.
[0596] 5. Multilingual Translation Server
[0597] Translate the generated summary and evaluation documents into multiple languages.
[0598] 6. Message sending function
[0599] Users enter messages through a dashboard, and generative artificial intelligence converts them into the appropriate format and sends them.
[0600] 7. Dashboard
[0601] This interface is used to visually display and manage assessment results and summaries.
[0602] 8. Social Media Integration
[0603] Users can easily share assessment results and summaries on social media from the dashboard.
[0604] Software Configuration
[0605] 1. Generative AI Model
[0606] It uses Hugging Face's BART model and a BERT-based natural language processing model.
[0607] 2. Voice Recognition Software
[0608] Convert audio data into text data using the Google Speech Recognition API or similar.
[0609] 3. Translation Services
[0610] Use a cloud-based translation service such as the Google Translate API.
[0611] 4. Data Visualization Tools
[0612] Display the evaluation results as a graph using seaborn and matplotlib.
[0613] Specific examples
[0614] For example, a voice command for an autonomous vehicle to turn left at an intersection is collected. This voice data is converted into text data by a speech recognition server, and summarized as a "left turn command" by a summary generation server. The scoring server then evaluates the appropriateness of this command and calculates an "accuracy score of 85 points." The evaluation results are translated into other languages by a multilingual translation server and displayed on a dashboard.
[0615] Prompt Sentence Examples
[0616] For example, use the following prompt:
[0617] "Collect audio data while driving, summarize and evaluate the data, translate the evaluation results into other languages, send them by email, and display the evaluation results as graphs."
[0618] In this way, by using the system of the present invention, it is possible to effectively evaluate the driving performance of an autonomous vehicle and provide transparent information.
[0619] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0620] Step 1:
[0621] Audio data collection
[0622] The device collects audio data using a microphone installed inside the autonomous vehicle.
[0623] Input: Voice data of conversations and instructions inside the vehicle
[0624] Output: Collected audio data file
[0625] Step 2:
[0626] Audio data conversion
[0627] The server converts the collected voice data files into text data using voice recognition software (e.g., Google Speech Recognition API).
[0628] Input: Audio data file
[0629] Output: Converted text data
[0630] What it does: Speech recognition software analyzes the audio waveform and outputs the corresponding text.
[0631] Step 3:
[0632] Summarizing text data
[0633] The server uses a generative artificial intelligence model (e.g., the BART model of Hugging Face) to summarize the text data.
[0634] Input: Converted text data
[0635] Output: Summary text
[0636] How it works: A generative artificial intelligence model extracts important parts of text data and generates a concise summary.
[0637] Step 4:
[0638] Driving performance scoring
[0639] The server evaluates driving performance using a BERT-based model based on the summarized text data and calculates a score.
[0640] Input: Summary text
[0641] Output: Performance score
[0642] How it works: A BERT-based model analyzes the summarized text and generates a score based on the evaluation metrics.
[0643] Step 5:
[0644] Evaluation document generation and multilingual translation
[0645] The server generates an evaluation document based on the calculated score and translates it into multiple languages using the Google Translate API or similar.
[0646] Input: Performance score, summary text
[0647] Output: Evaluation documents translated into multiple languages
[0648] Specific operation: Automatically generate evaluation documents and convert them into other languages using the translation API.
[0649] Step 6:
[0650] Sending a message
[0651] Users input messages to the subject of evaluation through the dashboard, and the generative artificial intelligence analyzes them and sends the messages in an appropriate format.
[0652] Input: The message entered by the user
[0653] Output: Message sent
[0654] Specific operation: Generative AI analyzes the content of the message, converts it into a format appropriate for the recipient, and then sends it.
[0655] Step 7:
[0656] Displaying evaluation results on the dashboard
[0657] The server visually displays the evaluation results on a dashboard.
[0658] Input: Performance score and summary text
[0659] Output: Visualized evaluation results (graphs and charts)
[0660] Specific operation: Use data visualization tools such as seaborn and matplotlib to display the evaluation results as graphs and charts.
[0661] Step 8:
[0662] Social media integration
[0663] Users can share evaluation documents and summary text on social media from the dashboard, and the device posts using social media APIs.
[0664] Input: Evaluation document and summary text
[0665] Output: Social media posts
[0666] Specific operation: Evaluation results are automatically posted via social media API.
[0667] The above processing steps make it possible to effectively evaluate the driving performance of an autonomous vehicle and provide transparent information.
[0668] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0669] This invention is a system that uses generative artificial intelligence to summarize the speeches and Q&A sessions of Diet members, and evaluates their activity levels and achievements based on those summaries. In particular, by combining it with an emotion engine that recognizes and analyzes the user's emotions, this invention realizes message exchange that reflects the user's emotional expressions.
[0670] Overall system overview
[0671] The system mainly consists of the following components:
[0672] 1. Voice data collection terminal
[0673] 2. Speech Recognition Server
[0674] 3. Abstract Generation Server
[0675] 4. Scoring Server
[0676] 5. Multilingual Translation Server
[0677] 6. Messaging and emotion engine
[0678] 7. Dashboard
[0679] 8. Social Media Integration
[0680] Audio data collection and conversion
[0681] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal by accessing the system's dedicated upload page, selecting the audio file, and clicking the upload button.
[0682] The terminal sends the uploaded voice data to a voice recognition server, which converts the voice data into text data and sends it to a summary generation server.
[0683] Summary Generation
[0684] The server receives the text data and generates a summary of the text data using generative artificial intelligence. The summary is generated using a BERT-based natural language processing model. The generated summary data is sent to a scoring server.
[0685] Scoring and report card generation
[0686] The server analyzes the summary data and calculates a score based on multiple metrics (e.g., number of comments, number of proposals, number of votes) that evaluate the activity and performance of each member of parliament. Based on the analyzed scores, a report card is generated and stored in a database.
[0687] Multilingual Translation
[0688] The server translates report cards and summary data into multiple languages using generative artificial intelligence models or cloud-based translation services (e.g., Google Translate API). The translation results are stored in a database and displayed on a dashboard.
[0689] Messaging and Sentiment Analysis
[0690] Users can input messages to lawmakers on the dashboard. As they input their messages, the emotion engine analyzes their emotions and reflects the results in the message content.
[0691] The device sends the input message and analyzed emotion data to the generative AI, which analyzes the message and sends it to the legislator in an appropriate format. The emotion data is used by the legislator to understand the user's emotions and suggest an appropriate reply.
[0692] Message examples
[0693] For example, when a user types, "I strongly desire an increase in the education budget," the emotion engine recognizes from the phrase "strongly desire" that the user has a high positive emotion. This emotion data is parsed as part of the message content and sent to the legislator. The legislator receives feedback with a high positive emotion and generates an appropriate reply based on that emotion.
[0694] Share on social media
[0695] Users select a report card or score from the dashboard and click the share button, and the device generates a link and share message that is posted via social media APIs.
[0696] As described above, the present invention is a system that transparently manages the activities and statements of Diet members, and further promotes richer communication by recognizing the user's emotions.
[0697] The processing flow will be explained below.
[0698] Step 1:
[0699] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal by accessing the system's dedicated upload page, selecting the audio file, and clicking the upload button.
[0700] Step 2:
[0701] The device sends the uploaded voice data to the voice recognition server. At the same time, the device receives the voice data and sends the voice file using the voice recognition API. This API converts the voice data into text.
[0702] Step 3:
[0703] The server analyzes the voice data received through the voice recognition API and converts it into text data. The voice recognition server converts the voice data into text data and sends the converted text data to the summary generation server.
[0704] Step 4:
[0705] The server receives text data at the summary generation server and generates a summary of the text data using generative artificial intelligence. The summary generation server activates the generative artificial intelligence model, extracts important points and keywords, and generates a summary. The generated summary data is sent to the scoring server.
[0706] Step 5:
[0707] The server sends the summarized text data to a scoring server, which analyzes metrics that evaluate the legislator's activity and performance. The scoring server analyzes the content, frequency, and activity of each statement, and calculates a score for each metric. The calculated scores are stored in a database.
[0708] Step 6:
[0709] The server generates a report card with the evaluation of each legislator. The scoring server calculates an overall evaluation score for each legislator based on the saved scores. The evaluation scores are compiled into a visually easy-to-understand report card and recorded in a database.
[0710] Step 7:
[0711] The server translates the generated report cards and summary data into multiple languages. The multilingual translation server uses generative artificial intelligence or a cloud translation API to convert data into different languages. The translation results are stored in a database and then reflected on the dashboard.
[0712] Step 8:
[0713] Users access the dashboard and input messages to legislators. When inputting a message, the emotion engine analyzes the user's emotions and reflects the results in the message content. Users enter text in the message input form on the dashboard and click the send button.
[0714] Step 9:
[0715] The device sends the input message and analyzed emotion data to the generative AI, which analyzes the message and sends the content to the legislator in an appropriate format. The emotion data is used by the legislator to understand the user's emotion and generate a reply accordingly.
[0716] Step 10:
[0717] Users can share their report cards or scores on social media by clicking the share button on the dashboard. The device generates a link and a share message and posts it via the social media API.
[0718] Through this series of processing steps, a system will be realized that transparently manages the activities and statements of Diet members, deepening the understanding of ordinary voters. In addition, the addition of an emotion engine will allow users' emotions to be reflected in messages, enabling richer communication.
[0719] Example 2
[0720] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0721] In conventional systems for evaluating the activities of legislators, the collection and analysis of voice data is performed manually, making it difficult to efficiently process large amounts of data. Furthermore, there are few ways for voters to provide direct, emotional feedback to legislators, preventing improvements in transparency and deeper communication. This has led to challenges in the accuracy of legislator activity evaluations and two-way communication with voters.
[0722] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0723] In this invention, the server includes: means for converting collected voice data into text data using generative artificial intelligence; means for automatically summarizing the text data; means for calculating scores by analyzing multiple metrics that evaluate the activity and performance of legislators based on the summarized data; means for translating the generated report card and summary data into multiple languages; means for analyzing messages from users using an emotion analysis engine and reflecting the emotion data in messages to legislators; means for voters to input and send messages to legislators via a dashboard; means for the generative artificial intelligence to analyze the input message and emotion data and send them to legislators in an appropriate format; and means for sharing the report card and summary data on social media. This enables efficient analysis and summarization of voice data, automated evaluation of legislator activity, and two-way communication with voters that reflects their emotions.
[0724] "Generative AI" refers to AI that has the ability to generate new information and data based on large amounts of data.
[0725] "Audio data" refers to recorded sounds or speech stored in digital format.
[0726] "Text data" refers to digital data that has been converted from audio data into text information.
[0727] A "summary" refers to information that has been shortened by extracting important information from long text data.
[0728] "Metrics" refers to the indicators and standards used to evaluate the activities of lawmakers.
[0729] "Sentiment analysis engine" refers to software or technology for analyzing and recognizing emotions from input text or voice.
[0730] "Dashboard" refers to the interface that allows system users to access and operate data and functions.
[0731] A "report card" is a report that evaluates a member of parliament's activities and achievements and compiles information such as scores.
[0732] "Multilingual translation" refers to the process or function of converting text from one language into another.
[0733] "Social media" refers to a platform on the Internet that allows users to share and interact with each other.
[0734] This invention is a system that uses generative artificial intelligence to summarize the speeches and Q&A sessions of legislators, and evaluates their activity levels and achievements based on the summaries. In particular, by combining it with an emotion engine that recognizes and analyzes the user's emotions, the system realizes message exchanges that reflect the user's emotional expressions.
[0735] System configuration
[0736] The system consists of the following components:
[0737] 1. Voice data collection terminal
[0738] 2. Speech Recognition Server
[0739] 3. Abstract Generation Server
[0740] 4. Scoring Server
[0741] 5. Multilingual Translation Server
[0742] 6. Messaging and emotion engine
[0743] 7. Dashboard
[0744] 8. Social Media Integration
[0745] Hardware and software used
[0746] The Google Cloud Speech-to-Text API is used as the speech recognition server.
[0747] The summary generation server uses a BERT-based natural language processing model.
[0748] The multilingual translation server uses the Google Translate API.
[0749] The Sentiment Analysis tool is used for sentiment analysis.
[0750] Audio data collection
[0751] Users record audio data of Diet deliberations and upload it to an audio data collection terminal. This process is performed using a dedicated upload page. Users access the upload page, select the audio file, and click the "Upload" button.
[0752] Converting audio data to text
[0753] The device sends the uploaded voice data to the speech recognition server, which receives the voice data and converts it into text data using the Google Cloud Speech-to-Text API, which then sends the converted text data to the summary generation server.
[0754] Summary Generation
[0755] The server receives the text data and generates a summary of the text data using generative artificial intelligence (e.g., a BERT-based natural language processing model). The summary generation process extracts important statements and keywords and shortens them while preserving the context. The generated summary data is sent to the scoring server.
[0756] Activity evaluation and scoring
[0757] The server analyzes the summary data to evaluate the activity and performance of each member of parliament (e.g., number of comments, number of proposals, number of votes). It calculates a score based on the analyzed data and generates a report card based on that score. The calculated report card is stored in a database.
[0758] Multilingual Translation
[0759] The server translates the generated report cards and summary data into multiple languages using the Google Translate API. The translated data is stored in a database and displayed on the dashboard.
[0760] Sentiment analysis and messaging
[0761] The user inputs a message to the legislator on the dashboard. At this time, the emotion engine analyzes the input message and recognizes the user's emotions. For example, Sentiment Analysis can be used to identify "highly positive emotions" and "negative emotions." The device then sends the analyzed emotion data to the generative AI, which then analyzes the message and sends it to the legislator in an appropriate format.
[0762] Message examples
[0763] For example, if a user types, "I strongly desire an increase in the education budget," the emotion engine recognizes from the phrase "strongly desire" that the user has a high positive emotion. This emotion data is parsed as part of the message content and sent to the legislator. The legislator receives feedback with a high positive emotion and generates an appropriate reply based on that emotion.
[0764] Share on social media
[0765] Users select a report card or score from the dashboard and click the share button, and the device generates a link and a share message, which is posted via social media APIs (such as Twitter API or Facebook API).
[0766] Example prompts for generative AI models
[0767] The user can provide prompts to the generative AI model, such as:
[0768] "We will create a report card to evaluate the activities of politicians. We will convert audio data into text, summarize the text, and analyze the number of statements and proposals made by politicians. As an example, please summarize and evaluate the following statements:
[0769] "There have been many problems with education in Japan over the past few years. I strongly hope that the education budget will be increased."
[0770] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0771] Step 1:
[0772] Users record audio data of Diet deliberations and upload it to the audio data collection terminal by accessing the upload page, selecting the audio file, and clicking the "Upload" button. This inputs the audio data into the system.
[0773] Step 2:
[0774] The device sends the uploaded voice data to the speech recognition server, which receives the voice data and converts it into text data using the Google Cloud Speech-to-Text API. Through this process, the voice data is converted into text data.
[0775] Step 3:
[0776] The server receives the text data and generates a summary of the text data using a BERT-based natural language processing model. The summary generation process extracts important information from the text data and generates a compact summary. The generated summary data is then sent to the scoring server.
[0777] Step 4:
[0778] The server analyzes the summary data to evaluate the activity and performance of each member of parliament using multiple metrics (e.g., number of comments, number of proposals, number of votes). It calculates a score based on the analyzed data and generates a report card based on that score. This report card is stored in a database.
[0779] Step 5:
[0780] The server translates the generated report cards and summary data into multiple languages. The generated data is converted into other languages using the Google Translate API, etc. The translated data is stored in a database and displayed on a dashboard.
[0781] Step 6:
[0782] Users input messages to their legislators on the dashboard. The input message is analyzed by a sentiment analysis engine to identify the user's sentiment. For example, natural language processing is used to identify "positive" or "negative" sentiment. This sentiment data is reflected in the message content and becomes input for analysis by the generative artificial intelligence.
[0783] Step 7:
[0784] The server receives messages containing emotional data obtained from the emotion analysis engine, and the generative AI analyzes them and sends them to the legislator in an appropriate format, allowing the legislator to understand the user's emotions and generate a reply accordingly.
[0785] Step 8:
[0786] Users select the generated report card or score from the dashboard and click the share button. The device generates a link and a share message and posts it via social media APIs (e.g., Twitter API, Facebook API). This process allows feedback, including the legislator's activity evaluation and the user's sentiment, to be widely shared.
[0787] (Application example 2)
[0788] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0789] It is difficult to grasp, evaluate, and improve the efficiency and results of work activities at work sites such as factories. Furthermore, in environments with multinational workers, language barriers are a major obstacle to communication and supervision. Furthermore, it is difficult to exchange appropriate feedback between on-site supervisors and workers in real time.
[0790] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for converting collected voice data into text data using generative artificial intelligence, means for automatically summarizing the text data, and means for calculating scores by analyzing multiple metrics for evaluating the efficiency and results of work activities based on the summarized data. This enables objective evaluation of work efficiency and results within a factory. The generative artificial intelligence also includes means for generating a report card evaluation of work activities based on the summarized data and translating it into multiple languages, means for the work supervisor to input and send feedback messages to workers via a display terminal, and means for storing the report card and summary data in a data management system. This enables communication across language barriers and appropriate feedback in real time.
[0791] "Generative AI" refers to AI that automatically generates, analyzes, and generates summaries and evaluations based on collected data.
[0792] "Voice data" refers to audio information about work collected in factories, work sites, etc.
[0793] "Text data" refers to character information converted from voice data using a voice recognition model.
[0794] A "summary" refers to content that has been created by using generative artificial intelligence to shorten text data and extract only the important information.
[0795] "Work activities" refers to specific work processes involving workers and robots in factories or on-site.
[0796] "Efficiency" refers to an indicator of how effectively work activities are being carried out.
[0797] "Outcomes" refers to the products obtained as a result of work activities or indicators of success.
[0798] "Metrics" refers to specific indicators or benchmarks for evaluating the efficiency and results of work activities.
[0799] "Score" refers to a numerical rating calculated based on metrics.
[0800] A "report card" refers to an evaluation report prepared based on calculated scores.
[0801] "Multilingual translation" refers to translating generated text data or report card contents into multiple languages.
[0802] "Display terminal" refers to a device for visually displaying data, such as smart glasses.
[0803] A "feedback message" refers to a message sent by a work supervisor to a worker for evaluation or guidance.
[0804] A "data management system" refers to a system that stores report cards and summary data and manages them so that they can be viewed later.
[0805] This invention is a system for improving the efficiency and evaluation of work activities in factories and on-site. It uses generative artificial intelligence to convert collected voice data into text data and automatically summarize it. The efficiency and results of work activities are then evaluated based on the summarized data. One specific embodiment of the invention is described in detail below.
[0806] Hardware and Software Configuration
[0807] The server includes the following hardware and software:
[0808] Speech recognition server: Converts collected voice data into text data. For example, the speech recognition server uses a "speech recognition model (e.g., Google Speech-to-Text API)."
[0809] Summarization server: Automatically summarizes text data. As a specific example, the summary generation server uses a generative AI model (e.g., a BERT-based natural language processing model).
[0810] Scoring Server: Evaluates the efficiency and performance of work activities based on summary data and calculates a score.
[0811] Multilingual translation server: Translates the generated data into multiple languages. A specific example is a cloud-based translation service (e.g., Google Translate API).
[0812] Data management system: Stores report cards and summary data and manages them for later review.
[0813] The display terminal is a device such as smart glasses. As a specific example, we will use "smart glasses (e.g., Microsoft HoloLens)."
[0814] Data processing and calculation processing
[0815] Audio data collection and conversion:
[0816] The user (work supervisor) uses smart glasses to collect voices while working on-site. The voice data is transmitted to a voice recognition server via wireless communication (e.g., Bluetooth, Wi-Fi).
[0817] The server (speech recognition server) uses a speech recognition model to convert the voice data into text data, using a generative AI model.
[0818] Summary generation:
[0819] The server (summarization server) summarizes text data using a generative artificial intelligence model, which utilizes a BERT-based natural language processing model.
[0820] Scoring and Rating:
[0821] The server (scoring server) analyzes metrics that evaluate the efficiency and results of work activities based on the summary data and calculates a score.
[0822] The calculated scores are generated as a report card and stored in a data management system.
[0823] Multilingual Translation:
[0824] The server (multilingual translation server) translates report cards and summary data into multiple languages and provides them to foreign workers as needed.
[0825] Send a feedback message:
[0826] The user (work supervisor) uses a display terminal (smart glasses) to input and send feedback messages to the workers.
[0827] The server (generative artificial intelligence) analyzes the feedback message and sends it to the worker in an appropriate format.
[0828] Examples of concrete examples and prompts
[0829] As a concrete example, the following prompts are used by a supervisor wearing smart glasses to record the robot's movements and the workers' tasks on-site:
[0830] Example prompt (summary generation): "Please summarize this text data: [text data from speech recognition results]"
[0831] Example prompt (sentiment analysis): "Analyze the sentiment of this message: [Supervisor's feedback]"
[0832] Example prompt (multilingual translation): "Please translate this text into [language]: [summary data / assessment results]"
[0833] In this way, the present invention provides improved efficiency, effective evaluation, and appropriate feedback at the workplace.
[0834] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0835] Step 1:
[0836] Collection and transmission of voice data
[0837] The user uses smart glasses to collect voices from the field. The collected voice data is sent to a voice recognition server via wireless communication (Bluetooth or Wi-Fi). The input is the voice data from the field, and the output is the data sent to the voice recognition server.
[0838] Step 2:
[0839] Converting audio data to text
[0840] The server (speech recognition server) converts the received voice data into text data using a voice recognition model (e.g., Google Speech-to-Text API). The input is voice data, and the output is text data based on this.
[0841] Step 3:
[0842] Summarizing text data
[0843] The server (summarization server) automatically summarizes the converted text data using a generative AI model (e.g., a BERT-based natural language processing model). The input is text data, and the output is summarized text data.
[0844] Step 4:
[0845] Work activity evaluation and scoring
[0846] The server (scoring server) analyzes metrics that evaluate the efficiency and results of work activities based on the summary data and calculates a score. The input is the summary data, and the output is the calculated score. The specific operation of scoring is to analyze multiple metrics such as the number of comments and the task completion rate.
[0847] Step 5:
[0848] Generate evaluation report cards and save data
[0849] The server (scoring server) generates a report card based on the calculated score and stores it in a data management system. The input is the score, and the output is the generated report card and its storage.
[0850] Step 6:
[0851] Multilingual Translation
[0852] The server (multilingual translation server) translates the generated report card and summary data into multiple languages. As a concrete example, we use the Google Translate API. The input is the report card and summary data, and the output is the translated data.
[0853] Step 7:
[0854] Enter and send your feedback message
[0855] The user (work supervisor) inputs a feedback message using a display terminal (smart glasses). The input message is sent to the generative AI model. The input is the user's feedback message, and the output is the message sent to the generative AI model.
[0856] Step 8:
[0857] Parsing and delivering feedback messages
[0858] The server (generative artificial intelligence) analyzes the feedback message and sends it to the worker in an appropriate format. The input is the analyzed feedback message, and the output is the appropriate message to the worker.
[0859] This series of steps enables the system to evaluate on-site work efficiency and provide feedback in real time.
[0860] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0861] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0862] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0863] [Third embodiment]
[0864] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0865] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0866] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0867] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0868] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0869] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0870] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0871] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0872] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0873] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0874] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0875] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0876] This invention is a system that uses generative artificial intelligence to summarize the speeches and Q&A sessions of Diet members, and evaluates their activity and performance based on the summaries. This system provides integrated functions for voice data conversion, summary generation, scoring, multilingual translation, message sending, and social media sharing.
[0877] Overall system overview
[0878] The system mainly consists of the following components:
[0879] 1. Voice data collection terminal
[0880] 2. Speech Recognition Server
[0881] 3. Abstract Generation Server
[0882] 4. Scoring Server
[0883] 5. Multilingual Translation Server
[0884] 6. Message sending function
[0885] 7. Dashboard
[0886] 8. Social Media Integration
[0887] Audio data collection and conversion
[0888] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal.
[0889] The device sends the uploaded voice data to a speech recognition server, which uses generative artificial intelligence to convert the voice data into text data. The text data returned from the speech recognition server is sent to a summary generation server.
[0890] Summary Generation
[0891] The server receives the text data and generates a summary of the text data using generative artificial intelligence (generative artificial intelligence), using a BERT-based natural language processing model, and then sends the summarized text data to a scoring server.
[0892] Scoring and report card generation
[0893] The server analyzes the summary data and calculates scores based on multiple metrics (e.g., number of comments, number of proposals, number of votes) that evaluate the activity and performance of each member of parliament. A score is generated based on each metric, and an overall evaluation score is calculated. A document visually summarizing these evaluation results is generated as a report card and saved in a database.
[0894] Multilingual Translation
[0895] The server translates report cards and summary data into multiple languages using generative artificial intelligence models or cloud-based translation services (e.g., Google Translate API). The translation results are displayed on a dashboard for users to view.
[0896] Sending a message
[0897] Users input messages to their legislators on the dashboard. The device then sends the input messages to the generative AI, which analyzes the contents of the messages. Based on the analysis results, the messages are sent to the legislators in an appropriate format.
[0898] Share on social media
[0899] To share a report card or score on social media, a user clicks the share button from the dashboard. The device generates a link and a share message and posts it using the social media API.
[0900] Specific examples
[0901] For example, a speech and question-and-answer session by a member of the National Diet is recorded and uploaded to the system as audio data. The speech is converted into text by a speech recognition server, and then summarized by a summary generation server as "a speech in the Diet regarding an increase in the education budget." The scoring server then evaluates the impact of the speech and adds the summarized data to metrics related to education policy. As a result, the member's score for education policy is calculated as 85 points, which is reflected in his / her report card.
[0902] In this way, by using the system of the present invention, it is possible to increase the transparency of the activities and statements of Diet members and promote political participation by ordinary voters.
[0903] The processing flow will be explained below.
[0904] Step 1:
[0905] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal by accessing the system's dedicated upload page, selecting the audio file, and clicking the upload button.
[0906] Step 2:
[0907] The device sends the uploaded voice data to the voice recognition server. At the same time as receiving the voice data, the device sends the voice file using the voice recognition API. This API converts the voice data into text.
[0908] Step 3:
[0909] The server analyzes the voice data received through the voice recognition API and converts it into text data. The voice recognition server then converts the spoken content into text using a matching algorithm, and sends the converted text data to the summary generation server.
[0910] Step 4:
[0911] The server receives the text data at the summary generation server and generates a summary of the text data using generative artificial intelligence. The summary generation server activates the generative artificial intelligence model, extracts important points and keywords, and generates a summary sentence.
[0912] Step 5:
[0913] The server sends the summarized text data to a scoring server, which analyzes metrics that evaluate the legislator's activity and performance. The scoring server analyzes the content, frequency, and activity of each statement, and calculates a score for each metric. The calculated scores are stored in a database.
[0914] Step 6:
[0915] The server generates a report card with the evaluation of each legislator. The scoring server calculates an overall evaluation score for each legislator based on the saved scores. The evaluation scores are compiled into a visually easy-to-understand report card and recorded in a database.
[0916] Step 7:
[0917] The server translates the generated report cards and summary data into multiple languages. The multilingual translation server uses generative artificial intelligence or a cloud translation API to convert data into different languages, and stores the translation results in a database. The results are then reflected on the dashboard.
[0918] Step 8:
[0919] Users access the dashboard and enter messages to their legislators by entering text directly into the message entry form on the dashboard and clicking the send button.
[0920] Step 9:
[0921] The terminal sends the input message to the generative AI, which analyzes the message content. The generative AI analyzes the message and sends the content to the legislator in an appropriate format.
[0922] Step 10:
[0923] Users can share their report cards and scores on social media by clicking the share button on the dashboard. When a user clicks the share button, the device generates a link and a share message and posts it via the corresponding social media API.
[0924] Through the above processing steps, a system will be realized in which the activities and statements of Diet members are managed transparently, leading to a deeper understanding among ordinary voters.
[0925] Example 1
[0926] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0927] In today's political environment, there is a need to increase transparency in the statements and activities of Diet members and provide reliable information that allows voters to evaluate them. However, the current situation makes it difficult to obtain detailed information about Diet member statements and question and answer sessions, which hinders fair evaluation of Diet members' activities and achievements. Furthermore, opportunities for political participation are limited by a lack of information provision in multiple languages, voter feedback, and the means to share that information on social media.
[0928] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0929] In this invention, the server includes means for converting collected voice data into text data, means for automatically summarizing the text data, means for calculating scores by analyzing multiple indicators for evaluating the activities and achievements of Diet members based on the summarized data, means for generating a report card evaluation of each Diet member based on the calculated score, means for translating the generated report card and summary data into multiple languages, means for voters to input and send messages to Diet members via a dashboard, means for sharing the report card and summary data on social media, means for translating data into multiple languages using a cloud-based translation service, means for posting data using a social media API, and means for providing dedicated terminals for uploading voice data. This increases the transparency of Diet members' statements and activities, enabling voters to properly evaluate Diet members, obtain information in multiple languages, send feedback, and further share that information widely.
[0930] "Generative AI" is an AI system that uses natural language processing technology to analyze and generate data.
[0931] "Audio data" refers to data recorded in digital format containing audio, including statements by members of parliament and question and answer sessions.
[0932] "Text data" refers to data obtained by converting voice data into character information.
[0933] A "summary" is a document that extracts key information from long text data and summarizes it in a short form.
[0934] "Multiple indicators" are criteria for evaluating a member of parliament's activity and achievements, such as the number of times they speak, the number of proposals they make, and the number of votes they pass.
[0935] The "score" is a numerical evaluation of a member of parliament's activity and achievements based on multiple indicators.
[0936] A "report card" is a document that visually summarizes a legislator's evaluation, including calculated scores.
[0937] "Multilingual translation" is the process of converting data or documents into multiple languages.
[0938] A "dashboard" is an interface that allows voters to view and interact with information.
[0939] "Social media" is an online platform for sharing information and interacting.
[0940] A "voice recognition model" is an algorithm for converting voice data into text data.
[0941] A "cloud-based translation service" is a service that uses translation functions provided via the Internet.
[0942] A "social media API" is an interface for accessing and operating social media functions from outside.
[0943] A "dedicated terminal" is a specific hardware device used to upload audio data.
[0944] This invention is a system that uses generative artificial intelligence to summarize the speeches and questions and answers of Diet members, and evaluates their activity and performance based on the summaries. This system integrates the following main components and functions:
[0945] Audio data collection and uploading
[0946] Users collect audio data of Diet deliberations using a recording device and upload it to a dedicated terminal. This dedicated terminal can be a general digital device such as a smartphone or PC. Users use these devices to record audio data and upload it to a voice recognition server on the cloud.
[0947] Converting audio data to text
[0948] The device sends the uploaded voice data to a voice recognition server. The server converts the voice data into text data using a voice recognition model (e.g., a cloud-based voice recognition service). Specifically, Google Cloud Speech-to-Text API is used. At this stage, the voice data is converted into text format.
[0949] Summarizing text data
[0950] The server sends the text data obtained from the speech recognition server to the summary generation server, which automatically summarizes the text data using a BERT-based natural language processing model (e.g., Hugging Face's transformers library). In this step, redundant information is removed, leaving the key information in a compact form.
[0951] Scoring summary data
[0952] The server receives the summary data and sends it to a scoring server. The scoring server then analyzes multiple indicators (e.g., number of statements, number of proposals, number of votes) that evaluate the activity and performance of each legislator based on the summary data, and calculates a score. Data analysis tools such as the Python pandas library are used for scoring.
[0953] Multilingual translation of scores and summary data
[0954] The server translates the calculated scores and summary data into multiple languages using a cloud-based translation service (e.g., Google Translate API). The translated data is then reflected on a dashboard, allowing users to view their assessment results in multiple languages.
[0955] Send a message to your legislators
[0956] Users can input feedback messages to legislators on the dashboard. The input messages are sent to the generative AI via the terminal, where the content is analyzed. The generative AI analyzes the content of the messages and sends them to legislators in an appropriate format.
[0957] Share on social media
[0958] Users can share their assessment results and report cards on social media from the dashboard. When a user clicks the share button, the device generates a link and shareable text and posts it using a social media API (e.g., Twitter API).
[0959] Specific examples
[0960] For example, if you upload audio data containing a congressman's speech about the education budget, the system processes it as follows: First, it converts the audio data into text using a cloud-based speech recognition service, and then summarizes the text using a BERT-based natural language processing model. The scoring server then analyzes the summarized data and calculates a score (e.g., 85 points) for the congressman's education policy. These results are then translated into multiple languages and displayed on a dashboard.
[0961] Prompt Sentence Examples
[0962] "Use this speech recognition system to analyze recent questions and answers from members of Congress regarding the education budget. Generate a summary of the results and an evaluation of the members' performance, display it in multiple languages, and make the results shareable on social media."
[0963] In this way, the system of the present invention increases the transparency of the activities and statements of Diet members and makes it possible to promote political participation among ordinary voters.
[0964] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0965] Step 1:
[0966] Users collect audio data of Diet deliberations and upload it to a dedicated device. Specifically, users use smartphones or digital recorders to record Diet members' remarks and Q&A sessions, and then upload the audio files to the device. The input is the audio data (e.g., a Diet member's statement, "We need to increase the education budget"), and the output is the saving of the audio file on the device.
[0967] Step 2:
[0968] The device sends the uploaded voice data to a voice recognition server on the cloud. The input is the voice data stored on the device, and the output is the data transfer to the voice recognition server. Specifically, the device uploads the voice data to the server via an internet connection.
[0969] Step 3:
[0970] The server uses a speech recognition model (e.g., a cloud-based speech recognition service) to convert the voice data into text data. The input is voice data, and the output is text data (e.g., "We need to increase the education budget"). Specifically, it uses services such as the Google Cloud Speech-to-Text API to perform highly accurate voice analysis.
[0971] Step 4:
[0972] The server transmits the text data acquired from the speech recognition server to the summary generation server. The input is text data, and the output is data transfer to the summary generation server. Specifically, data is transferred securely between the servers.
[0973] Step 5:
[0974] The server uses a summary generation server to automatically summarize text data using a BERT-based natural language processing model. The input is text data, and the output is summarized text data (e.g., "An increase in the education budget was discussed"). Specifically, natural language processing is performed using the Hugging Face transformers library.
[0975] Step 6:
[0976] The server sends the summary data to the scoring server, which analyzes multiple indicators and calculates a score. The input is the summary data, and the output is a score based on each indicator (e.g., 85 points). Specifically, analysis is performed using Python's pandas library based on data such as the number of comments, number of proposals, and number of votes.
[0977] Step 7:
[0978] The server translates the generated scores and summary data into multiple languages. The input is the scores and summary data, and the output is the translated data. Specifically, it uses a cloud-based translation service such as Google Translate API. The translated data is reflected in a dashboard and can be viewed by users.
[0979] Step 8:
[0980] The user inputs a feedback message to the legislator on the dashboard. The input is the message entered by the user on the dashboard (e.g., "I agree with the opinion on the education budget"), and the output is the message sent to the generative AI model. Specifically, the generative AI model analyzes the content of the message and sends it to the legislator in an appropriate format.
[0981] Step 9:
[0982] Users can share their assessment results and report cards on social media from the dashboard. The input is a click on the share button, and the output is a link and text for sharing. Specifically, the device generates the link and text for sharing and posts it to social media using the Twitter API or similar.
[0983] This series of processes will increase transparency in the statements and activities of Diet members, enable voters to properly evaluate their members, obtain information in multiple languages, send feedback, and share that information widely.
[0984] (Application example 1)
[0985] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0986] Since there is no means to evaluate the driving performance of autonomous vehicles, it is difficult to analyze and evaluate the appropriateness of decisions and actions while driving. Conventional systems do not adequately collect and analyze voice data and behavioral data while driving, making it difficult to improve autonomous driving technology and evaluate driver performance.
[0987] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0988] In this invention, the server includes means for converting collected voice data into text data using generative artificial intelligence, means for automatically summarizing the text data, means for calculating a score by analyzing multiple measurement indicators for evaluating the activity level and achievements of the subject based on the summarized data, means for generating a visualized document of the subject's evaluation based on the calculated score, means for translating the generated evaluation document and summary data into multiple languages, means for a user to input and send a message to the subject through a dashboard, means for sharing the evaluation document and summary data on social media, means for collecting driving data, analyzing and evaluating the data, and means for providing a management screen for visualizing and displaying the evaluation results. This makes it possible to effectively evaluate the driving performance of autonomous vehicles and provide transparent information to stakeholders.
[0989] "Generative AI" is a type of AI that has the ability to learn patterns based on large amounts of data and generate or predict new data.
[0990] "Text data" refers to character string data that has been converted into a natural language format from voice data or other input data.
[0991] "Summarization" refers to extracting important information from the original text data and presenting it in a short, concise form.
[0992] "Metrics" are the standards or measures used to evaluate the activities and outcomes of an evaluation.
[0993] A "score" is the result of quantifying the performance of an evaluation target based on measurement indicators.
[0994] A "visualized document" is a document that visually displays information such as evaluation results using graphs and charts.
[0995] "Multilingual translation" is the process of converting text data or evaluation documents into multiple languages.
[0996] A "dashboard" is an interface that allows users to enter information and view results.
[0997] A "message" is a text-based communication sent by a user to a target of evaluation.
[0998] "Social media" refers to online platforms for sharing information and engaging in two-way communication with the community.
[0999] "Driving data" refers to data about the vehicle's movements and surrounding conditions that is collected by an autonomous vehicle while it is driving.
[1000] The "management screen" is an interface for visually displaying and managing information such as evaluation results.
[1001] This invention relates to a system for evaluating the driving performance of autonomous vehicles. This system analyzes speech data and driving data, and provides the function of summarizing and evaluating the results. The entire system consists of the following components:
[1002] Hardware Configuration
[1003] 1. Voice data collection terminal
[1004] The device has a built-in microphone that collects conversations and instructions that occur inside the self-driving vehicle.
[1005] 2. Speech Recognition Server
[1006] The collected voice data is sent to this server and converted into text data using voice recognition technology.
[1007] 3. Abstract Generation Server
[1008] Text data received from a speech recognition server is summarized using generative artificial intelligence.
[1009] 4. Scoring Server
[1010] Based on the summarized data, multiple measurement indicators are analyzed to evaluate driving performance and a score is calculated.
[1011] 5. Multilingual Translation Server
[1012] Translate the generated summary and evaluation documents into multiple languages.
[1013] 6. Message sending function
[1014] Users enter messages through a dashboard, and generative artificial intelligence converts them into the appropriate format and sends them.
[1015] 7. Dashboard
[1016] This interface is used to visually display and manage assessment results and summaries.
[1017] 8. Social Media Integration
[1018] Users can easily share assessment results and summaries on social media from the dashboard.
[1019] Software Configuration
[1020] 1. Generative AI Model
[1021] It uses Hugging Face's BART model and a BERT-based natural language processing model.
[1022] 2. Voice Recognition Software
[1023] Convert audio data into text data using the Google Speech Recognition API or similar.
[1024] 3. Translation Services
[1025] Use a cloud-based translation service such as the Google Translate API.
[1026] 4. Data Visualization Tools
[1027] Display the evaluation results as a graph using seaborn and matplotlib.
[1028] Specific examples
[1029] For example, a voice command for an autonomous vehicle to turn left at an intersection is collected. This voice data is converted into text data by a speech recognition server, and summarized as a "left turn command" by a summary generation server. The scoring server then evaluates the appropriateness of this command and calculates an "accuracy score of 85 points." The evaluation results are translated into other languages by a multilingual translation server and displayed on a dashboard.
[1030] Prompt Sentence Examples
[1031] For example, use the following prompt:
[1032] "Collect audio data while driving, summarize and evaluate the data, translate the evaluation results into other languages, send them by email, and display the evaluation results as graphs."
[1033] In this way, by using the system of the present invention, it is possible to effectively evaluate the driving performance of an autonomous vehicle and provide transparent information.
[1034] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1035] Step 1:
[1036] Audio data collection
[1037] The device collects audio data using a microphone installed inside the autonomous vehicle.
[1038] Input: Voice data of conversations and instructions inside the vehicle
[1039] Output: Collected audio data file
[1040] Step 2:
[1041] Audio data conversion
[1042] The server converts the collected voice data files into text data using voice recognition software (e.g., Google Speech Recognition API).
[1043] Input: Audio data file
[1044] Output: Converted text data
[1045] What it does: Speech recognition software analyzes the audio waveform and outputs the corresponding text.
[1046] Step 3:
[1047] Summarizing text data
[1048] The server uses a generative artificial intelligence model (e.g., the BART model of Hugging Face) to summarize the text data.
[1049] Input: Converted text data
[1050] Output: Summary text
[1051] How it works: A generative artificial intelligence model extracts important parts of text data and generates a concise summary.
[1052] Step 4:
[1053] Driving performance scoring
[1054] The server evaluates driving performance using a BERT-based model based on the summarized text data and calculates a score.
[1055] Input: Summary text
[1056] Output: Performance score
[1057] How it works: A BERT-based model analyzes the summarized text and generates a score based on the evaluation metrics.
[1058] Step 5:
[1059] Evaluation document generation and multilingual translation
[1060] The server generates an evaluation document based on the calculated score and translates it into multiple languages using the Google Translate API or similar.
[1061] Input: Performance score, summary text
[1062] Output: Evaluation documents translated into multiple languages
[1063] Specific operation: Automatically generate evaluation documents and convert them into other languages using the translation API.
[1064] Step 6:
[1065] Sending a message
[1066] Users input messages to the subject of evaluation through the dashboard, and the generative artificial intelligence analyzes them and sends the messages in an appropriate format.
[1067] Input: The message entered by the user
[1068] Output: Message sent
[1069] Specific operation: Generative AI analyzes the content of the message, converts it into a format appropriate for the recipient, and then sends it.
[1070] Step 7:
[1071] Displaying evaluation results on the dashboard
[1072] The server visually displays the evaluation results on a dashboard.
[1073] Input: Performance score and summary text
[1074] Output: Visualized evaluation results (graphs and charts)
[1075] Specific operation: Use data visualization tools such as seaborn and matplotlib to display the evaluation results as graphs and charts.
[1076] Step 8:
[1077] Social media integration
[1078] Users can share evaluation documents and summary text on social media from the dashboard, and the device posts using social media APIs.
[1079] Input: Evaluation document and summary text
[1080] Output: Social media posts
[1081] Specific operation: Evaluation results are automatically posted via social media API.
[1082] The above processing steps make it possible to effectively evaluate the driving performance of an autonomous vehicle and provide transparent information.
[1083] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1084] This invention is a system that uses generative artificial intelligence to summarize the speeches and Q&A sessions of Diet members, and evaluates their activity levels and achievements based on those summaries. In particular, by combining it with an emotion engine that recognizes and analyzes the user's emotions, this invention realizes message exchange that reflects the user's emotional expressions.
[1085] Overall system overview
[1086] The system mainly consists of the following components:
[1087] 1. Voice data collection terminal
[1088] 2. Speech Recognition Server
[1089] 3. Abstract Generation Server
[1090] 4. Scoring Server
[1091] 5. Multilingual Translation Server
[1092] 6. Messaging and emotion engine
[1093] 7. Dashboard
[1094] 8. Social Media Integration
[1095] Audio data collection and conversion
[1096] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal by accessing the system's dedicated upload page, selecting the audio file, and clicking the upload button.
[1097] The terminal sends the uploaded voice data to a voice recognition server, which converts the voice data into text data and sends it to a summary generation server.
[1098] Summary Generation
[1099] The server receives the text data and generates a summary of the text data using generative artificial intelligence. The summary is generated using a BERT-based natural language processing model. The generated summary data is sent to a scoring server.
[1100] Scoring and report card generation
[1101] The server analyzes the summary data and calculates a score based on multiple metrics (e.g., number of comments, number of proposals, number of votes) that evaluate the activity and performance of each member of parliament. Based on the analyzed scores, a report card is generated and stored in a database.
[1102] Multilingual Translation
[1103] The server translates report cards and summary data into multiple languages using generative artificial intelligence models or cloud-based translation services (e.g., Google Translate API). The translation results are stored in a database and displayed on a dashboard.
[1104] Messaging and Sentiment Analysis
[1105] Users can input messages to lawmakers on the dashboard. As they input their messages, the emotion engine analyzes their emotions and reflects the results in the message content.
[1106] The device sends the input message and analyzed emotion data to the generative AI, which analyzes the message and sends it to the legislator in an appropriate format. The emotion data is used by the legislator to understand the user's emotions and suggest an appropriate reply.
[1107] Message examples
[1108] For example, when a user types, "I strongly desire an increase in the education budget," the emotion engine recognizes from the phrase "strongly desire" that the user has a high positive emotion. This emotion data is parsed as part of the message content and sent to the legislator. The legislator receives feedback with a high positive emotion and generates an appropriate reply based on that emotion.
[1109] Share on social media
[1110] Users select a report card or score from the dashboard and click the share button, and the device generates a link and share message that is posted via social media APIs.
[1111] As described above, the present invention is a system that transparently manages the activities and statements of Diet members, and further promotes richer communication by recognizing the user's emotions.
[1112] The processing flow will be explained below.
[1113] Step 1:
[1114] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal by accessing the system's dedicated upload page, selecting the audio file, and clicking the upload button.
[1115] Step 2:
[1116] The device sends the uploaded voice data to the voice recognition server. At the same time, the device receives the voice data and sends the voice file using the voice recognition API. This API converts the voice data into text.
[1117] Step 3:
[1118] The server analyzes the voice data received through the voice recognition API and converts it into text data. The voice recognition server converts the voice data into text data and sends the converted text data to the summary generation server.
[1119] Step 4:
[1120] The server receives text data at the summary generation server and generates a summary of the text data using generative artificial intelligence. The summary generation server activates the generative artificial intelligence model, extracts important points and keywords, and generates a summary. The generated summary data is sent to the scoring server.
[1121] Step 5:
[1122] The server sends the summarized text data to a scoring server, which analyzes metrics that evaluate the legislator's activity and performance. The scoring server analyzes the content, frequency, and activity of each statement, and calculates a score for each metric. The calculated scores are stored in a database.
[1123] Step 6:
[1124] The server generates a report card with the evaluation of each legislator. The scoring server calculates an overall evaluation score for each legislator based on the saved scores. The evaluation scores are compiled into a visually easy-to-understand report card and recorded in a database.
[1125] Step 7:
[1126] The server translates the generated report cards and summary data into multiple languages. The multilingual translation server uses generative artificial intelligence or a cloud translation API to convert data into different languages. The translation results are stored in a database and then reflected on the dashboard.
[1127] Step 8:
[1128] Users access the dashboard and input messages to legislators. When inputting a message, the emotion engine analyzes the user's emotions and reflects the results in the message content. Users enter text in the message input form on the dashboard and click the send button.
[1129] Step 9:
[1130] The device sends the input message and analyzed emotion data to the generative AI, which analyzes the message and sends the content to the legislator in an appropriate format. The emotion data is used by the legislator to understand the user's emotion and generate a reply accordingly.
[1131] Step 10:
[1132] Users can share their report cards or scores on social media by clicking the share button on the dashboard. The device generates a link and a share message and posts it via the social media API.
[1133] Through this series of processing steps, a system will be realized that transparently manages the activities and statements of Diet members, deepening the understanding of ordinary voters. In addition, the addition of an emotion engine will allow users' emotions to be reflected in messages, enabling richer communication.
[1134] Example 2
[1135] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1136] In conventional systems for evaluating the activities of legislators, the collection and analysis of voice data is performed manually, making it difficult to efficiently process large amounts of data. Furthermore, there are few ways for voters to provide direct, emotional feedback to legislators, preventing improvements in transparency and deeper communication. This has led to challenges in the accuracy of legislator activity evaluations and two-way communication with voters.
[1137] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1138] In this invention, the server includes: means for converting collected voice data into text data using generative artificial intelligence; means for automatically summarizing the text data; means for calculating scores by analyzing multiple metrics that evaluate the activity and performance of legislators based on the summarized data; means for translating the generated report card and summary data into multiple languages; means for analyzing messages from users using an emotion analysis engine and reflecting the emotion data in messages to legislators; means for voters to input and send messages to legislators via a dashboard; means for the generative artificial intelligence to analyze the input message and emotion data and send them to legislators in an appropriate format; and means for sharing the report card and summary data on social media. This enables efficient analysis and summarization of voice data, automated evaluation of legislator activity, and two-way communication with voters that reflects their emotions.
[1139] "Generative AI" refers to AI that has the ability to generate new information and data based on large amounts of data.
[1140] "Audio data" refers to recorded sounds or speech stored in digital format.
[1141] "Text data" refers to digital data that has been converted from audio data into text information.
[1142] A "summary" refers to information that has been shortened by extracting important information from long text data.
[1143] "Metrics" refers to the indicators and standards used to evaluate the activities of lawmakers.
[1144] "Sentiment analysis engine" refers to software or technology for analyzing and recognizing emotions from input text or voice.
[1145] "Dashboard" refers to the interface that allows system users to access and operate data and functions.
[1146] A "report card" is a report that evaluates a member of parliament's activities and achievements and compiles information such as scores.
[1147] "Multilingual translation" refers to the process or function of converting text from one language into another.
[1148] "Social media" refers to a platform on the Internet that allows users to share and interact with each other.
[1149] This invention is a system that uses generative artificial intelligence to summarize the speeches and Q&A sessions of legislators, and evaluates their activity levels and achievements based on the summaries. In particular, by combining it with an emotion engine that recognizes and analyzes the user's emotions, the system realizes message exchanges that reflect the user's emotional expressions.
[1150] System configuration
[1151] The system consists of the following components:
[1152] 1. Voice data collection terminal
[1153] 2. Speech Recognition Server
[1154] 3. Abstract Generation Server
[1155] 4. Scoring Server
[1156] 5. Multilingual Translation Server
[1157] 6. Messaging and emotion engine
[1158] 7. Dashboard
[1159] 8. Social Media Integration
[1160] Hardware and software used
[1161] The Google Cloud Speech-to-Text API is used as the speech recognition server.
[1162] The summary generation server uses a BERT-based natural language processing model.
[1163] The multilingual translation server uses the Google Translate API.
[1164] The Sentiment Analysis tool is used for sentiment analysis.
[1165] Audio data collection
[1166] Users record audio data of Diet deliberations and upload it to an audio data collection terminal. This process is performed using a dedicated upload page. Users access the upload page, select the audio file, and click the "Upload" button.
[1167] Converting audio data to text
[1168] The device sends the uploaded voice data to the speech recognition server, which receives the voice data and converts it into text data using the Google Cloud Speech-to-Text API, which then sends the converted text data to the summary generation server.
[1169] Summary Generation
[1170] The server receives the text data and generates a summary of the text data using generative artificial intelligence (e.g., a BERT-based natural language processing model). The summary generation process extracts important statements and keywords and shortens them while preserving the context. The generated summary data is sent to the scoring server.
[1171] Activity evaluation and scoring
[1172] The server analyzes the summary data to evaluate the activity and performance of each member of parliament (e.g., number of comments, number of proposals, number of votes). It calculates a score based on the analyzed data and generates a report card based on that score. The calculated report card is stored in a database.
[1173] Multilingual Translation
[1174] The server translates the generated report cards and summary data into multiple languages using the Google Translate API. The translated data is stored in a database and displayed on the dashboard.
[1175] Sentiment analysis and messaging
[1176] The user inputs a message to the legislator on the dashboard. At this time, the emotion engine analyzes the input message and recognizes the user's emotions. For example, Sentiment Analysis can be used to identify "highly positive emotions" and "negative emotions." The device then sends the analyzed emotion data to the generative AI, which then analyzes the message and sends it to the legislator in an appropriate format.
[1177] Message examples
[1178] For example, if a user types, "I strongly desire an increase in the education budget," the emotion engine recognizes from the phrase "strongly desire" that the user has a high positive emotion. This emotion data is parsed as part of the message content and sent to the legislator. The legislator receives feedback with a high positive emotion and generates an appropriate reply based on that emotion.
[1179] Share on social media
[1180] Users select a report card or score from the dashboard and click the share button, and the device generates a link and a share message, which is posted via social media APIs (such as Twitter API or Facebook API).
[1181] Example prompts for generative AI models
[1182] The user can provide prompts to the generative AI model, such as:
[1183] "We will create a report card to evaluate the activities of politicians. We will convert audio data into text, summarize the text, and analyze the number of statements and proposals made by politicians. As an example, please summarize and evaluate the following statements:
[1184] "There have been many problems with education in Japan over the past few years. I strongly hope that the education budget will be increased."
[1185] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1186] Step 1:
[1187] Users record audio data of Diet deliberations and upload it to the audio data collection terminal by accessing the upload page, selecting the audio file, and clicking the "Upload" button. This inputs the audio data into the system.
[1188] Step 2:
[1189] The device sends the uploaded voice data to the speech recognition server, which receives the voice data and converts it into text data using the Google Cloud Speech-to-Text API. Through this process, the voice data is converted into text data.
[1190] Step 3:
[1191] The server receives the text data and generates a summary of the text data using a BERT-based natural language processing model. The summary generation process extracts important information from the text data and generates a compact summary. The generated summary data is then sent to the scoring server.
[1192] Step 4:
[1193] The server analyzes the summary data to evaluate the activity and performance of each member of parliament using multiple metrics (e.g., number of comments, number of proposals, number of votes). It calculates a score based on the analyzed data and generates a report card based on that score. This report card is stored in a database.
[1194] Step 5:
[1195] The server translates the generated report cards and summary data into multiple languages. The generated data is converted into other languages using the Google Translate API, etc. The translated data is stored in a database and displayed on a dashboard.
[1196] Step 6:
[1197] Users input messages to their legislators on the dashboard. The input message is analyzed by a sentiment analysis engine to identify the user's sentiment. For example, natural language processing is used to identify "positive" or "negative" sentiment. This sentiment data is reflected in the message content and becomes input for analysis by the generative artificial intelligence.
[1198] Step 7:
[1199] The server receives messages containing emotional data obtained from the emotion analysis engine, and the generative AI analyzes them and sends them to the legislator in an appropriate format, allowing the legislator to understand the user's emotions and generate a reply accordingly.
[1200] Step 8:
[1201] Users select the generated report card or score from the dashboard and click the share button. The device generates a link and a share message and posts it via social media APIs (e.g., Twitter API, Facebook API). This process allows feedback, including the legislator's activity evaluation and the user's sentiment, to be widely shared.
[1202] (Application example 2)
[1203] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1204] It is difficult to grasp, evaluate, and improve the efficiency and results of work activities at work sites such as factories. Furthermore, in environments with multinational workers, language barriers are a major obstacle to communication and supervision. Furthermore, it is difficult to exchange appropriate feedback between on-site supervisors and workers in real time.
[1205] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for converting collected voice data into text data using generative artificial intelligence, means for automatically summarizing the text data, and means for calculating scores by analyzing multiple metrics for evaluating the efficiency and results of work activities based on the summarized data. This enables objective evaluation of work efficiency and results within a factory. The generative artificial intelligence also includes means for generating a report card evaluation of work activities based on the summarized data and translating it into multiple languages, means for the work supervisor to input and send feedback messages to workers via a display terminal, and means for storing the report card and summary data in a data management system. This enables communication across language barriers and appropriate feedback in real time.
[1206] "Generative AI" refers to AI that automatically generates, analyzes, and generates summaries and evaluations based on collected data.
[1207] "Voice data" refers to audio information about work collected in factories, work sites, etc.
[1208] "Text data" refers to character information converted from voice data using a voice recognition model.
[1209] A "summary" refers to content that has been created by using generative artificial intelligence to shorten text data and extract only the important information.
[1210] "Work activities" refers to specific work processes involving workers and robots in factories or on-site.
[1211] "Efficiency" refers to an indicator of how effectively work activities are being carried out.
[1212] "Outcomes" refers to the products obtained as a result of work activities or indicators of success.
[1213] "Metrics" refers to specific indicators or benchmarks for evaluating the efficiency and results of work activities.
[1214] "Score" refers to a numerical rating calculated based on metrics.
[1215] A "report card" refers to an evaluation report prepared based on calculated scores.
[1216] "Multilingual translation" refers to translating generated text data or report card contents into multiple languages.
[1217] "Display terminal" refers to a device for visually displaying data, such as smart glasses.
[1218] A "feedback message" refers to a message sent by a work supervisor to a worker for evaluation or guidance.
[1219] A "data management system" refers to a system that stores report cards and summary data and manages them so that they can be viewed later.
[1220] This invention is a system for improving the efficiency and evaluation of work activities in factories and on-site. It uses generative artificial intelligence to convert collected voice data into text data and automatically summarize it. The efficiency and results of work activities are then evaluated based on the summarized data. One specific embodiment of the invention is described in detail below.
[1221] Hardware and Software Configuration
[1222] The server includes the following hardware and software:
[1223] Speech recognition server: Converts collected voice data into text data. For example, the speech recognition server uses a "speech recognition model (e.g., Google Speech-to-Text API)."
[1224] Summarization server: Automatically summarizes text data. As a specific example, the summary generation server uses a generative AI model (e.g., a BERT-based natural language processing model).
[1225] Scoring Server: Evaluates the efficiency and performance of work activities based on summary data and calculates a score.
[1226] Multilingual translation server: Translates the generated data into multiple languages. A specific example is a cloud-based translation service (e.g., Google Translate API).
[1227] Data management system: Stores report cards and summary data and manages them for later review.
[1228] The display terminal is a device such as smart glasses. As a specific example, we will use "smart glasses (e.g., Microsoft HoloLens)."
[1229] Data processing and calculation processing
[1230] Audio data collection and conversion:
[1231] The user (work supervisor) uses smart glasses to collect voices while working on-site. The voice data is transmitted to a voice recognition server via wireless communication (e.g., Bluetooth, Wi-Fi).
[1232] The server (speech recognition server) uses a speech recognition model to convert the voice data into text data, using a generative AI model.
[1233] Summary generation:
[1234] The server (summarization server) summarizes text data using a generative artificial intelligence model, which utilizes a BERT-based natural language processing model.
[1235] Scoring and Rating:
[1236] The server (scoring server) analyzes metrics that evaluate the efficiency and results of work activities based on the summary data and calculates a score.
[1237] The calculated scores are generated as a report card and stored in a data management system.
[1238] Multilingual Translation:
[1239] The server (multilingual translation server) translates report cards and summary data into multiple languages and provides them to foreign workers as needed.
[1240] Send a feedback message:
[1241] The user (work supervisor) uses a display terminal (smart glasses) to input and send feedback messages to the workers.
[1242] The server (generative artificial intelligence) analyzes the feedback message and sends it to the worker in an appropriate format.
[1243] Examples of concrete examples and prompts
[1244] As a concrete example, the following prompts are used by a supervisor wearing smart glasses to record the robot's movements and the workers' tasks on-site:
[1245] Example prompt (summary generation): "Please summarize this text data: [text data from speech recognition results]"
[1246] Example prompt (sentiment analysis): "Analyze the sentiment of this message: [Supervisor's feedback]"
[1247] Example prompt (multilingual translation): "Please translate this text into [language]: [summary data / assessment results]"
[1248] In this way, the present invention provides improved efficiency, effective evaluation, and appropriate feedback at the workplace.
[1249] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1250] Step 1:
[1251] Collection and transmission of voice data
[1252] The user uses smart glasses to collect voices from the field. The collected voice data is sent to a voice recognition server via wireless communication (Bluetooth or Wi-Fi). The input is the voice data from the field, and the output is the data sent to the voice recognition server.
[1253] Step 2:
[1254] Converting audio data to text
[1255] The server (speech recognition server) converts the received voice data into text data using a voice recognition model (e.g., Google Speech-to-Text API). The input is voice data, and the output is text data based on this.
[1256] Step 3:
[1257] Summarizing text data
[1258] The server (summarization server) automatically summarizes the converted text data using a generative AI model (e.g., a BERT-based natural language processing model). The input is text data, and the output is summarized text data.
[1259] Step 4:
[1260] Work activity evaluation and scoring
[1261] The server (scoring server) analyzes metrics that evaluate the efficiency and results of work activities based on the summary data and calculates a score. The input is the summary data, and the output is the calculated score. The specific operation of scoring is to analyze multiple metrics such as the number of comments and the task completion rate.
[1262] Step 5:
[1263] Generate evaluation report cards and save data
[1264] The server (scoring server) generates a report card based on the calculated score and stores it in a data management system. The input is the score, and the output is the generated report card and its storage.
[1265] Step 6:
[1266] Multilingual Translation
[1267] The server (multilingual translation server) translates the generated report card and summary data into multiple languages. As a concrete example, we use the Google Translate API. The input is the report card and summary data, and the output is the translated data.
[1268] Step 7:
[1269] Enter and send your feedback message
[1270] The user (work supervisor) inputs a feedback message using a display terminal (smart glasses). The input message is sent to the generative AI model. The input is the user's feedback message, and the output is the message sent to the generative AI model.
[1271] Step 8:
[1272] Parsing and delivering feedback messages
[1273] The server (generative artificial intelligence) analyzes the feedback message and sends it to the worker in an appropriate format. The input is the analyzed feedback message, and the output is the appropriate message to the worker.
[1274] This series of steps enables the system to evaluate on-site work efficiency and provide feedback in real time.
[1275] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1276] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1277] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1278] [Fourth embodiment]
[1279] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1280] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1281] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1282] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1283] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1284] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1285] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1286] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1287] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1288] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1289] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1290] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1291] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1292] This invention is a system that uses generative artificial intelligence to summarize the speeches and Q&A sessions of Diet members, and evaluates their activity and performance based on the summaries. This system provides integrated functions for voice data conversion, summary generation, scoring, multilingual translation, message sending, and social media sharing.
[1293] Overall system overview
[1294] The system mainly consists of the following components:
[1295] 1. Voice data collection terminal
[1296] 2. Speech Recognition Server
[1297] 3. Abstract Generation Server
[1298] 4. Scoring Server
[1299] 5. Multilingual Translation Server
[1300] 6. Message sending function
[1301] 7. Dashboard
[1302] 8. Social Media Integration
[1303] Audio data collection and conversion
[1304] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal.
[1305] The device sends the uploaded voice data to a speech recognition server, which uses generative artificial intelligence to convert the voice data into text data. The text data returned from the speech recognition server is sent to a summary generation server.
[1306] Summary Generation
[1307] The server receives the text data and generates a summary of the text data using generative artificial intelligence (generative artificial intelligence), using a BERT-based natural language processing model, and then sends the summarized text data to a scoring server.
[1308] Scoring and report card generation
[1309] The server analyzes the summary data and calculates scores based on multiple metrics (e.g., number of comments, number of proposals, number of votes) that evaluate the activity and performance of each member of parliament. A score is generated based on each metric, and an overall evaluation score is calculated. A document visually summarizing these evaluation results is generated as a report card and saved in a database.
[1310] Multilingual Translation
[1311] The server translates report cards and summary data into multiple languages using generative artificial intelligence models or cloud-based translation services (e.g., Google Translate API). The translation results are displayed on a dashboard for users to view.
[1312] Sending a message
[1313] Users input messages to their legislators on the dashboard. The device then sends the input messages to the generative AI, which analyzes the contents of the messages. Based on the analysis results, the messages are sent to the legislators in an appropriate format.
[1314] Share on social media
[1315] To share a report card or score on social media, a user clicks the share button from the dashboard. The device generates a link and a share message and posts it using the social media API.
[1316] Specific examples
[1317] For example, a speech and question-and-answer session by a member of the National Diet is recorded and uploaded to the system as audio data. The speech is converted into text by a speech recognition server, and then summarized by a summary generation server as "a speech in the Diet regarding an increase in the education budget." The scoring server then evaluates the impact of the speech and adds the summarized data to metrics related to education policy. As a result, the member's score for education policy is calculated as 85 points, which is reflected in his / her report card.
[1318] In this way, by using the system of the present invention, it is possible to increase the transparency of the activities and statements of Diet members and promote political participation by ordinary voters.
[1319] The processing flow will be explained below.
[1320] Step 1:
[1321] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal by accessing the system's dedicated upload page, selecting the audio file, and clicking the upload button.
[1322] Step 2:
[1323] The device sends the uploaded voice data to the voice recognition server. At the same time as receiving the voice data, the device sends the voice file using the voice recognition API. This API converts the voice data into text.
[1324] Step 3:
[1325] The server analyzes the voice data received through the voice recognition API and converts it into text data. The voice recognition server then converts the spoken content into text using a matching algorithm, and sends the converted text data to the summary generation server.
[1326] Step 4:
[1327] The server receives the text data at the summary generation server and generates a summary of the text data using generative artificial intelligence. The summary generation server activates the generative artificial intelligence model, extracts important points and keywords, and generates a summary sentence.
[1328] Step 5:
[1329] The server sends the summarized text data to a scoring server, which analyzes metrics that evaluate the legislator's activity and performance. The scoring server analyzes the content, frequency, and activity of each statement, and calculates a score for each metric. The calculated scores are stored in a database.
[1330] Step 6:
[1331] The server generates a report card with the evaluation of each legislator. The scoring server calculates an overall evaluation score for each legislator based on the saved scores. The evaluation scores are compiled into a visually easy-to-understand report card and recorded in a database.
[1332] Step 7:
[1333] The server translates the generated report cards and summary data into multiple languages. The multilingual translation server uses generative artificial intelligence or a cloud translation API to convert data into different languages, and stores the translation results in a database. The results are then reflected on the dashboard.
[1334] Step 8:
[1335] Users access the dashboard and enter messages to their legislators by entering text directly into the message entry form on the dashboard and clicking the send button.
[1336] Step 9:
[1337] The terminal sends the input message to the generative AI, which analyzes the message content. The generative AI analyzes the message and sends the content to the legislator in an appropriate format.
[1338] Step 10:
[1339] Users can share their report cards and scores on social media by clicking the share button on the dashboard. When a user clicks the share button, the device generates a link and a share message and posts it via the corresponding social media API.
[1340] Through the above processing steps, a system will be realized in which the activities and statements of Diet members are managed transparently, leading to a deeper understanding among ordinary voters.
[1341] Example 1
[1342] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1343] In today's political environment, there is a need to increase transparency in the statements and activities of Diet members and provide reliable information that allows voters to evaluate them. However, the current situation makes it difficult to obtain detailed information about Diet member statements and question and answer sessions, which hinders fair evaluation of Diet members' activities and achievements. Furthermore, opportunities for political participation are limited by a lack of information provision in multiple languages, voter feedback, and the means to share that information on social media.
[1344] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1345] In this invention, the server includes means for converting collected voice data into text data, means for automatically summarizing the text data, means for calculating scores by analyzing multiple indicators for evaluating the activities and achievements of Diet members based on the summarized data, means for generating a report card evaluation of each Diet member based on the calculated score, means for translating the generated report card and summary data into multiple languages, means for voters to input and send messages to Diet members via a dashboard, means for sharing the report card and summary data on social media, means for translating data into multiple languages using a cloud-based translation service, means for posting data using a social media API, and means for providing dedicated terminals for uploading voice data. This increases the transparency of Diet members' statements and activities, enabling voters to properly evaluate Diet members, obtain information in multiple languages, send feedback, and further share that information widely.
[1346] "Generative AI" is an AI system that uses natural language processing technology to analyze and generate data.
[1347] "Audio data" refers to data recorded in digital format containing audio, including statements by members of parliament and question and answer sessions.
[1348] "Text data" refers to data obtained by converting voice data into character information.
[1349] A "summary" is a document that extracts key information from long text data and summarizes it in a short form.
[1350] "Multiple indicators" are criteria for evaluating a member of parliament's activity and achievements, such as the number of times they speak, the number of proposals they make, and the number of votes they pass.
[1351] The "score" is a numerical evaluation of a member of parliament's activity and achievements based on multiple indicators.
[1352] A "report card" is a document that visually summarizes a legislator's evaluation, including calculated scores.
[1353] "Multilingual translation" is the process of converting data or documents into multiple languages.
[1354] A "dashboard" is an interface that allows voters to view and interact with information.
[1355] "Social media" is an online platform for sharing information and interacting.
[1356] A "voice recognition model" is an algorithm for converting voice data into text data.
[1357] A "cloud-based translation service" is a service that uses translation functions provided via the Internet.
[1358] A "social media API" is an interface for accessing and operating social media functions from outside.
[1359] A "dedicated terminal" is a specific hardware device used to upload audio data.
[1360] This invention is a system that uses generative artificial intelligence to summarize the speeches and questions and answers of Diet members, and evaluates their activity and performance based on the summaries. This system integrates the following main components and functions:
[1361] Audio data collection and uploading
[1362] Users collect audio data of Diet deliberations using a recording device and upload it to a dedicated terminal. This dedicated terminal can be a general digital device such as a smartphone or PC. Users use these devices to record audio data and upload it to a voice recognition server on the cloud.
[1363] Converting audio data to text
[1364] The device sends the uploaded voice data to a voice recognition server. The server converts the voice data into text data using a voice recognition model (e.g., a cloud-based voice recognition service). Specifically, Google Cloud Speech-to-Text API is used. At this stage, the voice data is converted into text format.
[1365] Summarizing text data
[1366] The server sends the text data obtained from the speech recognition server to the summary generation server, which automatically summarizes the text data using a BERT-based natural language processing model (e.g., Hugging Face's transformers library). In this step, redundant information is removed, leaving the key information in a compact form.
[1367] Scoring summary data
[1368] The server receives the summary data and sends it to a scoring server. The scoring server then analyzes multiple indicators (e.g., number of statements, number of proposals, number of votes) that evaluate the activity and performance of each legislator based on the summary data, and calculates a score. Data analysis tools such as the Python pandas library are used for scoring.
[1369] Multilingual translation of scores and summary data
[1370] The server translates the calculated scores and summary data into multiple languages using a cloud-based translation service (e.g., Google Translate API). The translated data is then reflected on a dashboard, allowing users to view their assessment results in multiple languages.
[1371] Send a message to your legislators
[1372] Users can input feedback messages to legislators on the dashboard. The input messages are sent to the generative AI via the terminal, where the content is analyzed. The generative AI analyzes the content of the messages and sends them to legislators in an appropriate format.
[1373] Share on social media
[1374] Users can share their assessment results and report cards on social media from the dashboard. When a user clicks the share button, the device generates a link and shareable text and posts it using a social media API (e.g., Twitter API).
[1375] Specific examples
[1376] For example, if you upload audio data containing a congressman's speech about the education budget, the system processes it as follows: First, it converts the audio data into text using a cloud-based speech recognition service, and then summarizes the text using a BERT-based natural language processing model. The scoring server then analyzes the summarized data and calculates a score (e.g., 85 points) for the congressman's education policy. These results are then translated into multiple languages and displayed on a dashboard.
[1377] Prompt Sentence Examples
[1378] "Use this speech recognition system to analyze recent questions and answers from members of Congress regarding the education budget. Generate a summary of the results and an evaluation of the members' performance, display it in multiple languages, and make the results shareable on social media."
[1379] In this way, the system of the present invention increases the transparency of the activities and statements of Diet members and makes it possible to promote political participation among ordinary voters.
[1380] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1381] Step 1:
[1382] Users collect audio data of Diet deliberations and upload it to a dedicated device. Specifically, users use smartphones or digital recorders to record Diet members' remarks and Q&A sessions, and then upload the audio files to the device. The input is the audio data (e.g., a Diet member's statement, "We need to increase the education budget"), and the output is the saving of the audio file on the device.
[1383] Step 2:
[1384] The device sends the uploaded voice data to a voice recognition server on the cloud. The input is the voice data stored on the device, and the output is the data transfer to the voice recognition server. Specifically, the device uploads the voice data to the server via an internet connection.
[1385] Step 3:
[1386] The server uses a speech recognition model (e.g., a cloud-based speech recognition service) to convert the voice data into text data. The input is voice data, and the output is text data (e.g., "We need to increase the education budget"). Specifically, it uses services such as the Google Cloud Speech-to-Text API to perform highly accurate voice analysis.
[1387] Step 4:
[1388] The server transmits the text data acquired from the speech recognition server to the summary generation server. The input is text data, and the output is data transfer to the summary generation server. Specifically, data is transferred securely between the servers.
[1389] Step 5:
[1390] The server uses a summary generation server to automatically summarize text data using a BERT-based natural language processing model. The input is text data, and the output is summarized text data (e.g., "An increase in the education budget was discussed"). Specifically, natural language processing is performed using the Hugging Face transformers library.
[1391] Step 6:
[1392] The server sends the summary data to the scoring server, which analyzes multiple indicators and calculates a score. The input is the summary data, and the output is a score based on each indicator (e.g., 85 points). Specifically, analysis is performed using Python's pandas library based on data such as the number of comments, number of proposals, and number of votes.
[1393] Step 7:
[1394] The server translates the generated scores and summary data into multiple languages. The input is the scores and summary data, and the output is the translated data. Specifically, it uses a cloud-based translation service such as Google Translate API. The translated data is reflected in a dashboard and can be viewed by users.
[1395] Step 8:
[1396] The user inputs a feedback message to the legislator on the dashboard. The input is the message entered by the user on the dashboard (e.g., "I agree with the opinion on the education budget"), and the output is the message sent to the generative AI model. Specifically, the generative AI model analyzes the content of the message and sends it to the legislator in an appropriate format.
[1397] Step 9:
[1398] Users can share their assessment results and report cards on social media from the dashboard. The input is a click on the share button, and the output is a link and text for sharing. Specifically, the device generates the link and text for sharing and posts it to social media using the Twitter API or similar.
[1399] This series of processes will increase transparency in the statements and activities of Diet members, enable voters to properly evaluate their members, obtain information in multiple languages, send feedback, and share that information widely.
[1400] (Application example 1)
[1401] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1402] Since there is no means to evaluate the driving performance of autonomous vehicles, it is difficult to analyze and evaluate the appropriateness of decisions and actions while driving. Conventional systems do not adequately collect and analyze voice data and behavioral data while driving, making it difficult to improve autonomous driving technology and evaluate driver performance.
[1403] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1404] In this invention, the server includes means for converting collected voice data into text data using generative artificial intelligence, means for automatically summarizing the text data, means for calculating a score by analyzing multiple measurement indicators for evaluating the activity level and achievements of the subject based on the summarized data, means for generating a visualized document of the subject's evaluation based on the calculated score, means for translating the generated evaluation document and summary data into multiple languages, means for a user to input and send a message to the subject through a dashboard, means for sharing the evaluation document and summary data on social media, means for collecting driving data, analyzing and evaluating the data, and means for providing a management screen for visualizing and displaying the evaluation results. This makes it possible to effectively evaluate the driving performance of autonomous vehicles and provide transparent information to stakeholders.
[1405] "Generative AI" is a type of AI that has the ability to learn patterns based on large amounts of data and generate or predict new data.
[1406] "Text data" refers to character string data that has been converted into a natural language format from voice data or other input data.
[1407] "Summarization" refers to extracting important information from the original text data and presenting it in a short, concise form.
[1408] "Metrics" are the standards or measures used to evaluate the activities and outcomes of an evaluation.
[1409] A "score" is the result of quantifying the performance of an evaluation target based on measurement indicators.
[1410] A "visualized document" is a document that visually displays information such as evaluation results using graphs and charts.
[1411] "Multilingual translation" is the process of converting text data or evaluation documents into multiple languages.
[1412] A "dashboard" is an interface that allows users to enter information and view results.
[1413] A "message" is a text-based communication sent by a user to a target of evaluation.
[1414] "Social media" refers to online platforms for sharing information and engaging in two-way communication with the community.
[1415] "Driving data" refers to data about the vehicle's movements and surrounding conditions that is collected by an autonomous vehicle while it is driving.
[1416] The "management screen" is an interface for visually displaying and managing information such as evaluation results.
[1417] This invention relates to a system for evaluating the driving performance of autonomous vehicles. This system analyzes speech data and driving data, and provides the function of summarizing and evaluating the results. The entire system consists of the following components:
[1418] Hardware Configuration
[1419] 1. Voice data collection terminal
[1420] The device has a built-in microphone that collects conversations and instructions that occur inside the self-driving vehicle.
[1421] 2. Speech Recognition Server
[1422] The collected voice data is sent to this server and converted into text data using voice recognition technology.
[1423] 3. Abstract Generation Server
[1424] Text data received from a speech recognition server is summarized using generative artificial intelligence.
[1425] 4. Scoring Server
[1426] Based on the summarized data, multiple measurement indicators are analyzed to evaluate driving performance and a score is calculated.
[1427] 5. Multilingual Translation Server
[1428] Translate the generated summary and evaluation documents into multiple languages.
[1429] 6. Message sending function
[1430] Users enter messages through a dashboard, and generative artificial intelligence converts them into the appropriate format and sends them.
[1431] 7. Dashboard
[1432] This interface is used to visually display and manage assessment results and summaries.
[1433] 8. Social Media Integration
[1434] Users can easily share assessment results and summaries on social media from the dashboard.
[1435] Software Configuration
[1436] 1. Generative AI Model
[1437] It uses Hugging Face's BART model and a BERT-based natural language processing model.
[1438] 2. Voice Recognition Software
[1439] Convert audio data into text data using the Google Speech Recognition API or similar.
[1440] 3. Translation Services
[1441] Use a cloud-based translation service such as the Google Translate API.
[1442] 4. Data Visualization Tools
[1443] Display the evaluation results as a graph using seaborn and matplotlib.
[1444] Specific examples
[1445] For example, a voice command for an autonomous vehicle to turn left at an intersection is collected. This voice data is converted into text data by a speech recognition server, and summarized as a "left turn command" by a summary generation server. The scoring server then evaluates the appropriateness of this command and calculates an "accuracy score of 85 points." The evaluation results are translated into other languages by a multilingual translation server and displayed on a dashboard.
[1446] Prompt Sentence Examples
[1447] For example, use the following prompt:
[1448] "Collect audio data while driving, summarize and evaluate the data, translate the evaluation results into other languages, send them by email, and display the evaluation results as graphs."
[1449] In this way, by using the system of the present invention, it is possible to effectively evaluate the driving performance of an autonomous vehicle and provide transparent information.
[1450] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1451] Step 1:
[1452] Audio data collection
[1453] The device collects audio data using a microphone installed inside the autonomous vehicle.
[1454] Input: Voice data of conversations and instructions inside the vehicle
[1455] Output: Collected audio data file
[1456] Step 2:
[1457] Audio data conversion
[1458] The server converts the collected voice data files into text data using voice recognition software (e.g., Google Speech Recognition API).
[1459] Input: Audio data file
[1460] Output: Converted text data
[1461] What it does: Speech recognition software analyzes the audio waveform and outputs the corresponding text.
[1462] Step 3:
[1463] Summarizing text data
[1464] The server uses a generative artificial intelligence model (e.g., the BART model of Hugging Face) to summarize the text data.
[1465] Input: Converted text data
[1466] Output: Summary text
[1467] How it works: A generative artificial intelligence model extracts important parts of text data and generates a concise summary.
[1468] Step 4:
[1469] Driving performance scoring
[1470] The server evaluates driving performance using a BERT-based model based on the summarized text data and calculates a score.
[1471] Input: Summary text
[1472] Output: Performance score
[1473] How it works: A BERT-based model analyzes the summarized text and generates a score based on the evaluation metrics.
[1474] Step 5:
[1475] Evaluation document generation and multilingual translation
[1476] The server generates an evaluation document based on the calculated score and translates it into multiple languages using the Google Translate API or similar.
[1477] Input: Performance score, summary text
[1478] Output: Evaluation documents translated into multiple languages
[1479] Specific operation: Automatically generate evaluation documents and convert them into other languages using the translation API.
[1480] Step 6:
[1481] Sending a message
[1482] Users input messages to the subject of evaluation through the dashboard, and the generative artificial intelligence analyzes them and sends the messages in an appropriate format.
[1483] Input: The message entered by the user
[1484] Output: Message sent
[1485] Specific operation: Generative AI analyzes the content of the message, converts it into a format appropriate for the recipient, and then sends it.
[1486] Step 7:
[1487] Displaying evaluation results on the dashboard
[1488] The server visually displays the evaluation results on a dashboard.
[1489] Input: Performance score and summary text
[1490] Output: Visualized evaluation results (graphs and charts)
[1491] Specific operation: Use data visualization tools such as seaborn and matplotlib to display the evaluation results as graphs and charts.
[1492] Step 8:
[1493] Social media integration
[1494] Users can share evaluation documents and summary text on social media from the dashboard, and the device posts using social media APIs.
[1495] Input: Evaluation document and summary text
[1496] Output: Social media posts
[1497] Specific operation: Evaluation results are automatically posted via social media API.
[1498] The above processing steps make it possible to effectively evaluate the driving performance of an autonomous vehicle and provide transparent information.
[1499] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1500] This invention is a system that uses generative artificial intelligence to summarize the speeches and Q&A sessions of Diet members, and evaluates their activity levels and achievements based on those summaries. In particular, by combining it with an emotion engine that recognizes and analyzes the user's emotions, this invention realizes message exchange that reflects the user's emotional expressions.
[1501] Overall system overview
[1502] The system mainly consists of the following components:
[1503] 1. Voice data collection terminal
[1504] 2. Speech Recognition Server
[1505] 3. Abstract Generation Server
[1506] 4. Scoring Server
[1507] 5. Multilingual Translation Server
[1508] 6. Messaging and emotion engine
[1509] 7. Dashboard
[1510] 8. Social Media Integration
[1511] Audio data collection and conversion
[1512] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal by accessing the system's dedicated upload page, selecting the audio file, and clicking the upload button.
[1513] The terminal sends the uploaded voice data to a voice recognition server, which converts the voice data into text data and sends it to a summary generation server.
[1514] Summary Generation
[1515] The server receives the text data and generates a summary of the text data using generative artificial intelligence. The summary is generated using a BERT-based natural language processing model. The generated summary data is sent to a scoring server.
[1516] Scoring and report card generation
[1517] The server analyzes the summary data and calculates a score based on multiple metrics (e.g., number of comments, number of proposals, number of votes) that evaluate the activity and performance of each member of parliament. Based on the analyzed scores, a report card is generated and stored in a database.
[1518] Multilingual Translation
[1519] The server translates report cards and summary data into multiple languages using generative artificial intelligence models or cloud-based translation services (e.g., Google Translate API). The translation results are stored in a database and displayed on a dashboard.
[1520] Messaging and Sentiment Analysis
[1521] Users can input messages to lawmakers on the dashboard. As they input their messages, the emotion engine analyzes their emotions and reflects the results in the message content.
[1522] The device sends the input message and analyzed emotion data to the generative AI, which analyzes the message and sends it to the legislator in an appropriate format. The emotion data is used by the legislator to understand the user's emotions and suggest an appropriate reply.
[1523] Message examples
[1524] For example, when a user types, "I strongly desire an increase in the education budget," the emotion engine recognizes from the phrase "strongly desire" that the user has a high positive emotion. This emotion data is parsed as part of the message content and sent to the legislator. The legislator receives feedback with a high positive emotion and generates an appropriate reply based on that emotion.
[1525] Share on social media
[1526] Users select a report card or score from the dashboard and click the share button, and the device generates a link and share message that is posted via social media APIs.
[1527] As described above, the present invention is a system that transparently manages the activities and statements of Diet members, and further promotes richer communication by recognizing the user's emotions.
[1528] The processing flow will be explained below.
[1529] Step 1:
[1530] Users collect audio data of Diet deliberations using a recording device and upload it to an audio data collection terminal by accessing the system's dedicated upload page, selecting the audio file, and clicking the upload button.
[1531] Step 2:
[1532] The device sends the uploaded voice data to the voice recognition server. At the same time, the device receives the voice data and sends the voice file using the voice recognition API. This API converts the voice data into text.
[1533] Step 3:
[1534] The server analyzes the voice data received through the voice recognition API and converts it into text data. The voice recognition server converts the voice data into text data and sends the converted text data to the summary generation server.
[1535] Step 4:
[1536] The server receives text data at the summary generation server and generates a summary of the text data using generative artificial intelligence. The summary generation server activates the generative artificial intelligence model, extracts important points and keywords, and generates a summary. The generated summary data is sent to the scoring server.
[1537] Step 5:
[1538] The server sends the summarized text data to a scoring server, which analyzes metrics that evaluate the legislator's activity and performance. The scoring server analyzes the content, frequency, and activity of each statement, and calculates a score for each metric. The calculated scores are stored in a database.
[1539] Step 6:
[1540] The server generates a report card with the evaluation of each legislator. The scoring server calculates an overall evaluation score for each legislator based on the saved scores. The evaluation scores are compiled into a visually easy-to-understand report card and recorded in a database.
[1541] Step 7:
[1542] The server translates the generated report cards and summary data into multiple languages. The multilingual translation server uses generative artificial intelligence or a cloud translation API to convert data into different languages. The translation results are stored in a database and then reflected on the dashboard.
[1543] Step 8:
[1544] Users access the dashboard and input messages to legislators. When inputting a message, the emotion engine analyzes the user's emotions and reflects the results in the message content. Users enter text in the message input form on the dashboard and click the send button.
[1545] Step 9:
[1546] The device sends the input message and analyzed emotion data to the generative AI, which analyzes the message and sends the content to the legislator in an appropriate format. The emotion data is used by the legislator to understand the user's emotion and generate a reply accordingly.
[1547] Step 10:
[1548] Users can share their report cards or scores on social media by clicking the share button on the dashboard. The device generates a link and a share message and posts it via the social media API.
[1549] Through this series of processing steps, a system will be realized that transparently manages the activities and statements of Diet members, deepening the understanding of ordinary voters. In addition, the addition of an emotion engine will allow users' emotions to be reflected in messages, enabling richer communication.
[1550] Example 2
[1551] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1552] In conventional systems for evaluating the activities of legislators, the collection and analysis of voice data is performed manually, making it difficult to efficiently process large amounts of data. Furthermore, there are few ways for voters to provide direct, emotional feedback to legislators, preventing improvements in transparency and deeper communication. This has led to challenges in the accuracy of legislator activity evaluations and two-way communication with voters.
[1553] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1554] In this invention, the server includes: means for converting collected voice data into text data using generative artificial intelligence; means for automatically summarizing the text data; means for calculating scores by analyzing multiple metrics that evaluate the activity and performance of legislators based on the summarized data; means for translating the generated report card and summary data into multiple languages; means for analyzing messages from users using an emotion analysis engine and reflecting the emotion data in messages to legislators; means for voters to input and send messages to legislators via a dashboard; means for the generative artificial intelligence to analyze the input message and emotion data and send them to legislators in an appropriate format; and means for sharing the report card and summary data on social media. This enables efficient analysis and summarization of voice data, automated evaluation of legislator activity, and two-way communication with voters that reflects their emotions.
[1555] "Generative AI" refers to AI that has the ability to generate new information and data based on large amounts of data.
[1556] "Audio data" refers to recorded sounds or speech stored in digital format.
[1557] "Text data" refers to digital data that has been converted from audio data into text information.
[1558] A "summary" refers to information that has been shortened by extracting important information from long text data.
[1559] "Metrics" refers to the indicators and standards used to evaluate the activities of lawmakers.
[1560] "Sentiment analysis engine" refers to software or technology for analyzing and recognizing emotions from input text or voice.
[1561] "Dashboard" refers to the interface that allows system users to access and operate data and functions.
[1562] A "report card" is a report that evaluates a member of parliament's activities and achievements and compiles information such as scores.
[1563] "Multilingual translation" refers to the process or function of converting text from one language into another.
[1564] "Social media" refers to a platform on the Internet that allows users to share and interact with each other.
[1565] This invention is a system that uses generative artificial intelligence to summarize the speeches and Q&A sessions of legislators, and evaluates their activity levels and achievements based on the summaries. In particular, by combining it with an emotion engine that recognizes and analyzes the user's emotions, the system realizes message exchanges that reflect the user's emotional expressions.
[1566] System configuration
[1567] The system consists of the following components:
[1568] 1. Voice data collection terminal
[1569] 2. Speech Recognition Server
[1570] 3. Abstract Generation Server
[1571] 4. Scoring Server
[1572] 5. Multilingual Translation Server
[1573] 6. Messaging and emotion engine
[1574] 7. Dashboard
[1575] 8. Social Media Integration
[1576] Hardware and software used
[1577] The Google Cloud Speech-to-Text API is used as the speech recognition server.
[1578] The summary generation server uses a BERT-based natural language processing model.
[1579] The multilingual translation server uses the Google Translate API.
[1580] The Sentiment Analysis tool is used for sentiment analysis.
[1581] Audio data collection
[1582] Users record audio data of Diet deliberations and upload it to an audio data collection terminal. This process is performed using a dedicated upload page. Users access the upload page, select the audio file, and click the "Upload" button.
[1583] Converting audio data to text
[1584] The device sends the uploaded voice data to the speech recognition server, which receives the voice data and converts it into text data using the Google Cloud Speech-to-Text API, which then sends the converted text data to the summary generation server.
[1585] Summary Generation
[1586] The server receives the text data and generates a summary of the text data using generative artificial intelligence (e.g., a BERT-based natural language processing model). The summary generation process extracts important statements and keywords and shortens them while preserving the context. The generated summary data is sent to the scoring server.
[1587] Activity evaluation and scoring
[1588] The server analyzes the summary data to evaluate the activity and performance of each member of parliament (e.g., number of comments, number of proposals, number of votes). It calculates a score based on the analyzed data and generates a report card based on that score. The calculated report card is stored in a database.
[1589] Multilingual Translation
[1590] The server translates the generated report cards and summary data into multiple languages using the Google Translate API. The translated data is stored in a database and displayed on the dashboard.
[1591] Sentiment analysis and messaging
[1592] The user inputs a message to the legislator on the dashboard. At this time, the emotion engine analyzes the input message and recognizes the user's emotions. For example, Sentiment Analysis can be used to identify "highly positive emotions" and "negative emotions." The device then sends the analyzed emotion data to the generative AI, which then analyzes the message and sends it to the legislator in an appropriate format.
[1593] Message examples
[1594] For example, if a user types, "I strongly desire an increase in the education budget," the emotion engine recognizes from the phrase "strongly desire" that the user has a high positive emotion. This emotion data is parsed as part of the message content and sent to the legislator. The legislator receives feedback with a high positive emotion and generates an appropriate reply based on that emotion.
[1595] Share on social media
[1596] Users select a report card or score from the dashboard and click the share button, and the device generates a link and a share message, which is posted via social media APIs (such as Twitter API or Facebook API).
[1597] Example prompts for generative AI models
[1598] The user can provide prompts to the generative AI model, such as:
[1599] "We will create a report card to evaluate the activities of politicians. We will convert audio data into text, summarize the text, and analyze the number of statements and proposals made by politicians. As an example, please summarize and evaluate the following statements:
[1600] "There have been many problems with education in Japan over the past few years. I strongly hope that the education budget will be increased."
[1601] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1602] Step 1:
[1603] Users record audio data of Diet deliberations and upload it to the audio data collection terminal by accessing the upload page, selecting the audio file, and clicking the "Upload" button. This inputs the audio data into the system.
[1604] Step 2:
[1605] The device sends the uploaded voice data to the speech recognition server, which receives the voice data and converts it into text data using the Google Cloud Speech-to-Text API. Through this process, the voice data is converted into text data.
[1606] Step 3:
[1607] The server receives the text data and generates a summary of the text data using a BERT-based natural language processing model. The summary generation process extracts important information from the text data and generates a compact summary. The generated summary data is then sent to the scoring server.
[1608] Step 4:
[1609] The server analyzes the summary data to evaluate the activity and performance of each member of parliament using multiple metrics (e.g., number of comments, number of proposals, number of votes). It calculates a score based on the analyzed data and generates a report card based on that score. This report card is stored in a database.
[1610] Step 5:
[1611] The server translates the generated report cards and summary data into multiple languages. The generated data is converted into other languages using the Google Translate API, etc. The translated data is stored in a database and displayed on a dashboard.
[1612] Step 6:
[1613] Users input messages to their legislators on the dashboard. The input message is analyzed by a sentiment analysis engine to identify the user's sentiment. For example, natural language processing is used to identify "positive" or "negative" sentiment. This sentiment data is reflected in the message content and becomes input for analysis by the generative artificial intelligence.
[1614] Step 7:
[1615] The server receives messages containing emotional data obtained from the emotion analysis engine, and the generative AI analyzes them and sends them to the legislator in an appropriate format, allowing the legislator to understand the user's emotions and generate a reply accordingly.
[1616] Step 8:
[1617] Users select the generated report card or score from the dashboard and click the share button. The device generates a link and a share message and posts it via social media APIs (e.g., Twitter API, Facebook API). This process allows feedback, including the legislator's activity evaluation and the user's sentiment, to be widely shared.
[1618] (Application example 2)
[1619] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1620] It is difficult to grasp, evaluate, and improve the efficiency and results of work activities at work sites such as factories. Furthermore, in environments with multinational workers, language barriers are a major obstacle to communication and supervision. Furthermore, it is difficult to exchange appropriate feedback between on-site supervisors and workers in real time.
[1621] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for converting collected voice data into text data using generative artificial intelligence, means for automatically summarizing the text data, and means for calculating scores by analyzing multiple metrics for evaluating the efficiency and results of work activities based on the summarized data. This enables objective evaluation of work efficiency and results within a factory. The generative artificial intelligence also includes means for generating a report card evaluation of work activities based on the summarized data and translating it into multiple languages, means for the work supervisor to input and send feedback messages to workers via a display terminal, and means for storing the report card and summary data in a data management system. This enables communication across language barriers and appropriate feedback in real time.
[1622] "Generative AI" refers to AI that automatically generates, analyzes, and generates summaries and evaluations based on collected data.
[1623] "Voice data" refers to audio information about work collected in factories, work sites, etc.
[1624] "Text data" refers to character information converted from voice data using a voice recognition model.
[1625] A "summary" refers to content that has been created by using generative artificial intelligence to shorten text data and extract only the important information.
[1626] "Work activities" refers to specific work processes involving workers and robots in factories or on-site.
[1627] "Efficiency" refers to an indicator of how effectively work activities are being carried out.
[1628] "Outcomes" refers to the products obtained as a result of work activities or indicators of success.
[1629] "Metrics" refers to specific indicators or benchmarks for evaluating the efficiency and results of work activities.
[1630] "Score" refers to a numerical rating calculated based on metrics.
[1631] A "report card" refers to an evaluation report prepared based on calculated scores.
[1632] "Multilingual translation" refers to translating generated text data or report card contents into multiple languages.
[1633] "Display terminal" refers to a device for visually displaying data, such as smart glasses.
[1634] A "feedback message" refers to a message sent by a work supervisor to a worker for evaluation or guidance.
[1635] A "data management system" refers to a system that stores report cards and summary data and manages them so that they can be viewed later.
[1636] This invention is a system for improving the efficiency and evaluation of work activities in factories and on-site. It uses generative artificial intelligence to convert collected voice data into text data and automatically summarize it. The efficiency and results of work activities are then evaluated based on the summarized data. One specific embodiment of the invention is described in detail below.
[1637] Hardware and Software Configuration
[1638] The server includes the following hardware and software:
[1639] Speech recognition server: Converts collected voice data into text data. For example, the speech recognition server uses a "speech recognition model (e.g., Google Speech-to-Text API)."
[1640] Summarization server: Automatically summarizes text data. As a specific example, the summary generation server uses a generative AI model (e.g., a BERT-based natural language processing model).
[1641] Scoring Server: Evaluates the efficiency and performance of work activities based on summary data and calculates a score.
[1642] Multilingual translation server: Translates the generated data into multiple languages. A specific example is a cloud-based translation service (e.g., Google Translate API).
[1643] Data management system: Stores report cards and summary data and manages them for later review.
[1644] The display terminal is a device such as smart glasses. As a specific example, we will use "smart glasses (e.g., Microsoft HoloLens)."
[1645] Data processing and calculation processing
[1646] Audio data collection and conversion:
[1647] The user (work supervisor) uses smart glasses to collect voices while working on-site. The voice data is transmitted to a voice recognition server via wireless communication (e.g., Bluetooth, Wi-Fi).
[1648] The server (speech recognition server) uses a speech recognition model to convert the voice data into text data, using a generative AI model.
[1649] Summary generation:
[1650] The server (summarization server) summarizes text data using a generative artificial intelligence model, which utilizes a BERT-based natural language processing model.
[1651] Scoring and Rating:
[1652] The server (scoring server) analyzes metrics that evaluate the efficiency and results of work activities based on the summary data and calculates a score.
[1653] The calculated scores are generated as a report card and stored in a data management system.
[1654] Multilingual Translation:
[1655] The server (multilingual translation server) translates report cards and summary data into multiple languages and provides them to foreign workers as needed.
[1656] Send a feedback message:
[1657] The user (work supervisor) uses a display terminal (smart glasses) to input and send feedback messages to the workers.
[1658] The server (generative artificial intelligence) analyzes the feedback message and sends it to the worker in an appropriate format.
[1659] Examples of concrete examples and prompts
[1660] As a concrete example, the following prompts are used by a supervisor wearing smart glasses to record the robot's movements and the workers' tasks on-site:
[1661] Example prompt (summary generation): "Please summarize this text data: [text data from speech recognition results]"
[1662] Example prompt (sentiment analysis): "Analyze the sentiment of this message: [Supervisor's feedback]"
[1663] Example prompt (multilingual translation): "Please translate this text into [language]: [summary data / assessment results]"
[1664] In this way, the present invention provides improved efficiency, effective evaluation, and appropriate feedback at the workplace.
[1665] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1666] Step 1:
[1667] Collection and transmission of voice data
[1668] The user uses smart glasses to collect voices from the field. The collected voice data is sent to a voice recognition server via wireless communication (Bluetooth or Wi-Fi). The input is the voice data from the field, and the output is the data sent to the voice recognition server.
[1669] Step 2:
[1670] Converting audio data to text
[1671] The server (speech recognition server) converts the received voice data into text data using a voice recognition model (e.g., Google Speech-to-Text API). The input is voice data, and the output is text data based on this.
[1672] Step 3:
[1673] Summarizing text data
[1674] The server (summarization server) automatically summarizes the converted text data using a generative AI model (e.g., a BERT-based natural language processing model). The input is text data, and the output is summarized text data.
[1675] Step 4:
[1676] Work activity evaluation and scoring
[1677] The server (scoring server) analyzes metrics that evaluate the efficiency and results of work activities based on the summary data and calculates a score. The input is the summary data, and the output is the calculated score. The specific operation of scoring is to analyze multiple metrics such as the number of comments and the task completion rate.
[1678] Step 5:
[1679] Generate evaluation report cards and save data
[1680] The server (scoring server) generates a report card based on the calculated score and stores it in a data management system. The input is the score, and the output is the generated report card and its storage.
[1681] Step 6:
[1682] Multilingual Translation
[1683] The server (multilingual translation server) translates the generated report card and summary data into multiple languages. As a concrete example, we use the Google Translate API. The input is the report card and summary data, and the output is the translated data.
[1684] Step 7:
[1685] Enter and send your feedback message
[1686] The user (work supervisor) inputs a feedback message using a display terminal (smart glasses). The input message is sent to the generative AI model. The input is the user's feedback message, and the output is the message sent to the generative AI model.
[1687] Step 8:
[1688] Parsing and delivering feedback messages
[1689] The server (generative artificial intelligence) analyzes the feedback message and sends it to the worker in an appropriate format. The input is the analyzed feedback message, and the output is the appropriate message to the worker.
[1690] This series of steps enables the system to evaluate on-site work efficiency and provide feedback in real time.
[1691] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1692] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1693] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1694] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1695] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1696] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1697] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1698] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1699] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1700] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1701] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1702] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1703] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1704] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1705] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1706] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1707] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1708] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1709] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1710] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1711] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1712] The following is further disclosed regarding the above embodiment.
[1713] (Claim 1)
[1714] A means for converting the collected voice data into text data using generative artificial intelligence;
[1715] a means for automatically summarizing text data;
[1716] A method for calculating scores based on the summarized data by analyzing multiple metrics that evaluate the activity and performance of legislators;
[1717] A means for generating a report card evaluation of each member of parliament based on the calculated score;
[1718] a means for translating the generated report card and summary data into multiple languages;
[1719] A dashboard that allows voters to type and send messages to their representatives;
[1720] A system that includes a means to share report cards and summary data on social media.
[1721] (Claim 2)
[1722] 10. The system of claim 1, wherein the means for converting the speech data into text data uses a speech recognition model.
[1723] (Claim 3)
[1724] 2. The system of claim 1, further comprising means for the generative artificial intelligence to analyze messages for the legislator and transmit them to the legislator in an appropriate format.
[1725] "Example 1"
[1726] (Claim 1)
[1727] A means for converting the collected voice data into text data using generative artificial intelligence;
[1728] a means for automatically summarizing text data;
[1729] A method for calculating scores by analyzing multiple indicators that evaluate the activity and performance of legislators based on the summarized data;
[1730] A means for generating a report card evaluation of each member of parliament based on the calculated score;
[1731] a means for translating the generated report card and summary data into multiple languages;
[1732] A dashboard that allows voters to type and send messages to their representatives;
[1733] A means to share report cards and summary data on social media;
[1734] A means to translate data into multiple languages using a cloud-based translation service;
[1735] A means of posting data using social media APIs;
[1736] A system including means for providing a dedicated terminal for uploading audio data.
[1737] (Claim 2)
[1738] 10. The system of claim 1, wherein the means for converting the speech data into text data uses a speech recognition model.
[1739] (Claim 3)
[1740] 2. The system of claim 1, further comprising means for the generative artificial intelligence to analyze messages for the legislator and transmit them to the legislator in an appropriate format.
[1741] "Application Example 1"
[1742] (Claim 1)
[1743] A means for converting the collected voice data into text data using generative artificial intelligence;
[1744] a means for automatically summarizing text data;
[1745] A means for calculating a score by analyzing a plurality of measurement indicators for evaluating the activity level and results of the subject based on the summarized data;
[1746] a means for generating a visualized document of the evaluation of the evaluation target based on the calculated score;
[1747] a means for translating the generated evaluation documents and summary data into multiple languages;
[1748] A means for users to input and send messages to the evaluation target through the dashboard;
[1749] A means to share assessment documents and summary data on social media;
[1750] A means for collecting operational data and analyzing and evaluating the data;
[1751] A means to provide a management screen for visualizing and displaying evaluation results
[1752] A system including:
[1753] (Claim 2)
[1754] 10. The system of claim 1, wherein the means for converting the speech data into text data uses a speech recognition model.
[1755] (Claim 3)
[1756] The system according to claim 1, further comprising means for the generative artificial intelligence to analyze a message to be evaluated and transmit it to the evaluation target in an appropriate format.
[1757] "Example 2: Combining Emotion Engines"
[1758] (Claim 1)
[1759] A means for converting the collected voice data into text data using generative artificial intelligence;
[1760] a means for automatically summarizing text data;
[1761] A method for calculating scores based on the summarized data by analyzing multiple metrics that evaluate the activity and performance of legislators;
[1762] A means for generating a report card evaluation of each member of parliament based on the calculated score;
[1763] a means for translating the generated report card and summary data into multiple languages;
[1764] A means for analyzing messages from users using an emotion analysis engine and reflecting the emotion data in messages to legislators;
[1765] A dashboard that allows voters to type and send messages to their representatives;
[1766] A means for a generative artificial intelligence to analyze the input message and emotional data and send it to the legislator in an appropriate format;
[1767] A system that includes a means to share report cards and summary data on social media.
[1768] (Claim 2)
[1769] 10. The system of claim 1, wherein the means for converting the speech data into text data uses a speech recognition model.
[1770] (Claim 3)
[1771] 2. The system of claim 1, further comprising means for a member of parliament to generate an appropriate reply using a result of sentiment analysis of the input message.
[1772] "Application example 2 when combining emotion engines"
[1773] (Claim 1)
[1774] A means for converting the collected voice data into text data using generative artificial intelligence;
[1775] a means for automatically summarizing text data;
[1776] a means for calculating a score based on the summarized data by analyzing multiple metrics that evaluate the efficiency and performance of work activities;
[1777] A means for generating an evaluation of each work activity as a report card based on the calculated score;
[1778] a means for translating the generated report card and summary data into multiple languages;
[1779] a means for a work supervisor to input and send a feedback message to a worker via a display terminal;
[1780] A system including a means for storing report cards and summary data in a data management system.
[1781] (Claim 2)
[1782] 10. The system of claim 1, wherein the means for converting the speech data into text data uses a speech recognition model.
[1783] (Claim 3)
[1784] 2. The system according to claim 1, further comprising means for the generative artificial intelligence to analyze messages to the worker and transmit them to the worker in an appropriate format. [Explanation of symbols]
[1785] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for converting the collected voice data into text data using generative artificial intelligence; a means for automatically summarizing text data; A method for calculating scores based on the summarized data by analyzing multiple metrics that evaluate the activity and performance of legislators; A means for generating a report card evaluation of each member of parliament based on the calculated score; a means for translating the generated report card and summary data into multiple languages; A means for voters to type and send messages to their representatives through a dashboard; A system that includes a means to share report cards and summary data on social media.
2. 2. The system of claim 1, wherein the means for converting speech data to text data uses a speech recognition model.
3. 2. The system of claim 1, further comprising means for the generative artificial intelligence to analyze messages for the legislators and transmit them to the legislators in an appropriate format.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A