Speech generating device, speech generating method, and speech generating program

The speech assistance device addresses the challenge of smooth communication by generating corrected speech data with user-specific attributes, improving conversation flow and reducing errors.

JP2025186060APending Publication Date: 2025-12-23NTT DOCOMO BUSINESS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024094637
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-11
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Conventional technologies struggle to facilitate smooth communication by considering background information such as user's gender, age, place of origin, and culture, leading to difficulties in understanding the flow of conversations.

Method used

A speech assistance device that generates corrected speech data by inputting user text into a generative model with prior knowledge, including user attributes and speech history, to correct punctuation, typos, and expressions, and outputs the corrected data.

Benefits of technology

Enables smooth communication by correcting speech errors and maintaining user speech characteristics, reducing misunderstandings and enhancing communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025186060000001_ABST
    Figure 2025186060000001_ABST
Patent Text Reader

Abstract

To enable smooth communication.SOLUTION: A speech generating device 100 disclosed herein is configured to input a prompt based on a natural language text received from a user into a generative model provided with samples of sentences created during work so as to generate corrected speech data by correcting speech indicated by the text according to the samples. The speech generating device 100 outputs the corrected speech data that was generated to the user.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a speech assistance device, a speech assistance method, and a speech assistance program. [Background technology]

[0002] In recent years, technologies for efficiently communicating with others have become known. For example, a conventional technology is known that displays messages posted on a chat timeline on a display screen so that the connection between a first message and a second message can be seen (see, for example, Patent Document 1). The above-mentioned conventional technology makes it easy to understand the flow of conversation even in chat communication between multiple people, enabling efficient communication. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-153881 Summary of the Invention [Problem to be solved by the invention]

[0004] However, conventional technologies have issues with implementing smooth communication. For example, conventional technologies clarify the connections between multiple messages displayed on a timeline, making it easier to understand the flow of a series of conversations, such as questions and answers. However, conventional technologies have difficulty realizing communication that takes into account background information such as the user's gender, age, place of origin, upbringing, and culture. [Means for solving the problem]

[0005] Therefore, in order to solve the above-mentioned problems and achieve the objectives, the speech assistance device of the present invention is characterized by having a generation unit that inputs a prompt based on natural language text received from a user into a generation model that is provided with a sample of a sentence to be created in business, and generates corrected speech data by correcting the speech shown in the text based on the sample, and an output unit that outputs the generated corrected speech data to the user. [Effects of the Invention]

[0006] The present invention has the effect of enabling smooth communication. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a diagram illustrating an overall view of the processing of the speech assistance device according to this embodiment. [Figure 2] FIG. 2 is a diagram showing the configuration of the speech assistance device according to this embodiment. [Figure 3] FIG. 3 is a table diagram showing an example of employee information according to this embodiment. [Figure 4] FIG. 4 is a table showing an example of text information according to this embodiment. [Figure 5] FIG. 5 is a table diagram showing an example of an utterance history according to this embodiment. [Figure 6] FIG. 6 is a table showing an example of an utterance example according to this embodiment. [Figure 7] FIG. 7 is a diagram showing an example of the speech assistance process according to this embodiment. [Figure 8] FIG. 8 is a diagram showing an example of the speech assistance process according to this embodiment. [Figure 9] FIG. 9 is a diagram illustrating an example of the speech assistance process according to this embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of the speech assistance process according to this embodiment. [Figure 11]FIG. 11 is a flowchart showing the speech assistance process according to this embodiment. [Figure 12] FIG. 12 is a flowchart showing the speech assistance process according to this embodiment. [Figure 13] FIG. 13 is a diagram illustrating tsuzumi. [Figure 14] FIG. 14 is a diagram illustrating tsuzumi. [Figure 15] FIG. 15 is a diagram illustrating tsuzumi. [Figure 16] FIG. 16 is a diagram illustrating IOWN. [Figure 17] FIG. 17 is a diagram illustrating an example of a computer that executes the speech assistance process according to this embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention (hereinafter referred to as "embodiments") will be described with reference to the drawings. Note that the embodiments are not limited to the following description.

[0009] <Overview> Fig. 1 is a diagram illustrating an overview of the processing of the speech assistance device 100 according to this embodiment. The speech assistance device 100 shown in Fig. 1 is an example of a computer that provides a technology for performing predetermined correction processing, such as correcting punctuation, typos, and expressions, on the content spoken by users when they converse with each other.

[0010] In the following sections, an "employee" is a person who belongs to a specific organization and performs tasks assigned to the organization. A "user" is an employee who uses the speech assistance device 100 according to this embodiment, or an administrator who manages the speech assistance device 100, in the above-mentioned specific organization.

[0011] (background) There are known reference technologies that make it easier to understand the flow of conversations and enable efficient communication even in chat communication between multiple people. However, with the above-mentioned reference technologies, it is difficult to realize smooth communication that takes into account background information such as the user's gender, age, birthplace, upbringing, and culture.

[0012] (Processing by speech assistance device 100) Therefore, in order to solve the above-mentioned problems, the speech assistance device 100 according to this embodiment uses a natural language text (hereinafter, may be simply referred to as "text information") related to a task received from a user to generate speech data (hereinafter, may be referred to as "corrected speech data") corrected for the user's utterance in a generative model, and outputs the corrected speech data to the user. Returning to Fig. 1, the speech assistance process by the speech assistance device 100 will now be described.

[0013] First, the speech support device 100 receives text information related to a task input by a user. Next, the speech support device 100 performs a conversion process for the text information ((1) in FIG. 1). Specifically, the speech support device 100 converts the received text information into a prompt to be input to a generative model ((1-1) and (1-2) in FIG. 1).

[0014] The speech support device 100 executes a process for generating corrected utterance data ((2) in FIG. 1). The speech support device 100 inputs a prompt ((2-1) in FIG. 1) based on text information received from a user into the generation model 10 ((2-2) in FIG. 1) that is provided with a sample of a sentence created in business (hereinafter, sometimes simply referred to as "sample data"), and generates corrected utterance data ((2-3) in FIG. 1) by correcting the utterance shown in the text information based on the sample data.

[0015] The speech assistance device 100 outputs the generated corrected utterance data to the user ((3) in FIG. 1). Specifically, the speech assistance device 100 outputs the corrected utterance data generated by a predetermined generative model to a terminal device 200 operated by the user. The terminal device 200 can then display the corrected utterance data output by the speech assistance device 100 to the user ((3-1) in FIG. 1).

[0016] In this way, the speech assistance device 100 according to this embodiment can generate corrected speech data in which corrections have been made to the user's speech, such as corrections to punctuation, typos, and expressions, and output the corrected speech data to the user. Therefore, by making the predetermined corrections as described above, the speech assistance device 100 can eliminate misunderstandings and communication errors between users, thereby enabling smooth communication.

[0017] <Description of speech assistive device 100> Next, the configuration of the speech assistance device 100 according to this embodiment will be described. FIG. 2 is a diagram showing the configuration of the speech assistance device 100 according to this embodiment. As shown in FIG. 2, the speech assistance device 100 has a communication unit 110, a storage unit 120, and a control unit 130. Although not shown in FIG. 2, the speech assistance device 100 may also have an input unit such as a keyboard or a mouse for receiving input such as operations by an administrator or the like. Furthermore, the speech assistance device 100 may also have a display unit such as a display for displaying information such as input text information, prompts set for the generative model, generated sample data, corrected speech data, etc. to the administrator or the like.

[0018] (Communication unit 110) The communication unit 110 performs data communication related to the input of text information input via the terminal device 200 operated by the user. The communication unit 110 also performs data communication related to the output of corrected utterance data.

[0019] The communication unit 110 is realized by a NIC (Network Interface Card) or the like, and controls communication via an electric communication line such as a LAN (Local Area Network), the Internet, etc. The communication unit 110 is connected to the network by wire or wirelessly as necessary, and can transmit and receive information bidirectionally with the generative model 10, the terminal device 200, etc.

[0020] (Storage unit 120) The storage unit 120 stores data and programs used for various processes by the control unit 130, as well as various data acquired by the operation of the control unit 130. The storage unit 120 is realized by a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. As shown in FIG. 2, the storage unit 120 has an employee information DB 121, a text information DB 122, and a sample data DB 123.

[0021] (Employee Information DB121) Employee information DB121 is a database that stores employee information including identification information and attribute information of employees who belong to a predetermined organization. Specifically, employee information DB121 stores information such as employee identification information based on a combination of predetermined characters, numbers, symbols, etc. that identify employees, and organizational attribute information, which is attribute information of employees within the organization.

[0022] Here, an example of employee information stored in the employee information DB 121 will be described with reference to a table diagram. Fig. 3 is a table diagram showing an example of employee information according to this embodiment.

[0023] 3, employee information DB121 stores employee identification information and organization attribute information in association with "No.", which is information identifying individual employee information. For example, employee information DB121 stores employee identification information "A" and organization attribute information "B," which are identified by No. "1."

[0024] The above-mentioned employee identification information "A" may include information for identifying the employee, such as the employee's name, nickname, handle name, identification number, identification symbol, identification character string, etc.

[0025] Organizational attribute information "B" may include gender, seniority, department, position, job rank, career history, talent, skills, areas of expertise, qualifications, acquaintances, communication tendencies and preferences, communication data, personality, chat data, self-introduction data (self-promotion), career sheet, keywords, and other information used for matching offers.

[0026] (Text information DB122) The text information DB 122 is a database that stores text information, which is information on business-related conversations and utterances input by employees. An example of the text information stored in the text information DB 122 will now be described using a table. Fig. 4 is a table showing an example of text information according to this embodiment.

[0027] 4, the text information DB 122 stores employee identification information and text in association with "No.", which is information that identifies individual text information. For example, the text information DB 122 stores employee identification information "A" and text "C," which are identified by No. "1."

[0028] The above-mentioned employee identification information "A" is information that identifies an employee, and is the same information as the employee identification information stored in employee information DB 121. That is, this means that the employee information stored in employee information DB 121 and the text information stored in text information DB 122 correspond to each other.

[0029] Text "C" is information on dialogues and speeches based on natural language input via the terminal device 200 operated by an employee. Specifically, the text includes business-related text such as text about purchasing, text about systems used in business, text about business schedules, text about business and sales activities, text about employee skills and evaluations, and text including general business-related conversations between employees.

[0030] (Sample data DB123) The sample data DB123 is a database that stores information about the user's past speech history (speech history), information about speech examples (speech examples) that are speech data from other employees, etc. that are used as reference examples, etc., as sample data to be used to generate corrected speech data.

[0031] First, an example of an utterance history stored in the sample data DB 123 will be described with reference to the table shown in Fig. 5. Fig. 5 is a table showing an example of an utterance history according to this embodiment.

[0032] 5, the sample data DB 123 stores employee identification information and a speech history, which is past speech data by the employee, in association with "No.", which is information identifying an individual speech history. For example, the sample data DB 123 stores employee identification information "A" and a speech history "D," which are identified by No. "1."

[0033] The employee identification information described above is information for identifying a user who is an employee, and is the same information as the employee identification information stored in the employee information DB 121. That is, this means that the employee information stored in the employee information DB 121 and the speech history stored in the sample data DB 123 are associated with each other.

[0034] The speech history is information about speech made by the user. For example, the speech history includes the content of speech made by employees stored as text data, voice data, image data, or other types of data.

[0035] Next, an example of an utterance example stored in the sample data DB 123 will be described with reference to the table shown in Fig. 6. Fig. 6 is a table showing an example of an utterance example according to this embodiment.

[0036] 6, the sample data DB 123 stores example utterances, which are utterance data made by employees and used as reference examples for the correction process, in association with "No.", which is information identifying individual example utterances. For example, the sample data DB 123 stores example utterance "E," which is identified by No. "1."

[0037] The above-mentioned example utterances are information about utterances made by employees, and are information from which the personal information of the employees has been deleted. For example, the utterance history includes the content of utterances made by employees stored as text data, audio data, image data, or other types of data.

[0038] (control unit 130) Here, we will return to Fig. 2 to continue the explanation. The control unit 130 has an internal memory for temporarily storing programs that define various processing procedures and the like of the speech assistance device 100 and processing data, and is realized by electronic circuits such as a CPU (Central Processing Unit) and an MPU (Micro Processing Unit), and integrated circuits such as an ASIC (Application Specific Integrated Circuit) and an FPGA (Field Programmable Gate Array). As shown in Fig. 2, the control unit 130 has a receiving unit 131, a generating unit 132, a storage unit 133, and an output unit 134.

[0039] (Reception Department 131) The receiving unit 131 receives text information input by the user via the terminal device 200 via the communication unit 110 described above.

[0040] (Generation unit 132) The generation unit 132 provides predetermined prior knowledge, such as sample data including an utterance history and example utterances related to the employee's work, to the generative model to be used. Then, the generation unit 132 inputs a prompt based on text information received from the user to the generative model to which the predetermined prior knowledge has been provided, thereby generating corrected utterance data. Note that the provision of the above-mentioned prior knowledge may be realized by prompt engineering, adapter tuning, or the like, which will be described later.

[0041] Specifically, the generation unit 132 inputs a prompt including an instruction to generate corrected utterance data to a generation model to which the user's utterance history has been further provided, and generates corrected utterance data.

[0042] For example, the generation unit 132 converts text information input by a user into a prompt. Next, the generation unit 132 inputs the converted prompt into a generation model that has been provided in advance with information such as "organization attribute information" stored in the employee information DB 121, and "utterance history" and "utterance examples" stored in the sample data DB 123. Then, the generation unit 132 generates corrected utterance data by correcting the user's utterance indicated in the text information input by the user.

[0043] Furthermore, the generation unit 132 can generate corrected utterance data that retains the speech characteristics of the target user. Specifically, the generation unit 132 inputs a prompt including an instruction to correct (generate) the utterance data of the user by imitating the speech characteristics of the user to a generation model to which the "organization attribute information" and "utterance history" related to the user who made the utterance are provided as prior knowledge, and generates corrected utterance data that retains or imitates the speech characteristics of the user.

[0044] Furthermore, the generation unit 132 performs at least one of the following on the utterance indicated in the text information: correcting punctuation marks, correcting specific symbols, correcting expressions that should be avoided, correcting typos or omissions, and correcting expressions according to time, place, and occasion (hereinafter, sometimes simply referred to as "TPO (Time, Place, Occasion)"). The generation unit 132 also generates corrected utterance data to which at least one of an explanation of an expression, an explanation of an expression in a specific culture, and an explanation of an expression with multiple meanings has been added. An example of the above-mentioned correction process will be described later in the section on example processing.

[0045] The generation unit 132 converts text information related to the employee's work, which is input by the employee on a daily basis, into a prompt including an instruction to generate sample data. Next, the generation unit 132 inputs the converted prompt into the generative model to which the sample data has been provided, thereby generating the sample data. The generated business support information is then stored in the memory unit 120 (sample data DB 123) by the storage unit 133, which will be described later. Note that an example of the above-mentioned correction process will be described in the section on an example of processing, which will be described later.

[0046] The generation unit 132 can use at least one of a large-scale language model having general-purpose knowledge and tsuzumi as a generative model. Furthermore, the generation unit 132 can use a generative model that communicates via a communication network related to IOWN (Innovative Optical and Wireless Network). Details of tsuzumi and IOWN will be explained later in the section on modified examples.

[0047] (storage section 133) The storage unit 133 stores the sample data generated by the generation unit 132 in the memory unit 120 (sample data DB 123). Specifically, the storage unit 133 stores the sample data including an utterance history, an utterance example, etc. in the memory unit 120 (sample data DB 123). Note that an example of the storage process by the storage unit 133 will be described in detail in the section describing an example of the process described later.

[0048] (output unit 134) The output unit 134 outputs to the user the corrected utterance data generated by the generation unit 132. An example of output by the output unit 134 will be described later in the section on an example of processing.

[0049] (Generative Model 10) The generative model 10 is a generative model such as a large-scale language model, and generates output information in response to a generation instruction in an input natural language. Specifically, the generative model 10 generates sample data and corrected utterance data in response to text information input by the speech assistance device 100, based on a large-scale language model or the like in which a prompt is set to generate sample data and corrected utterance data.

[0050] The generative model 10 according to this embodiment may be realized by an information processing device such as a server. The generative model 10 can generate an answer in response to a prompt based on prior knowledge that has been provided, input, learned, added, etc.

[0051] (Terminal device 200) The terminal device 200 is an information processing terminal device operated by a user. Specifically, the terminal device 200 receives text information from the user, transmits the received text information to the speech assistance device 100, and receives and displays corrected utterance data from the speech assistance device 100.

[0052] (Example of processing) An example of the speech assistance process by the speech assistance device 100 according to this embodiment will now be described with reference to Fig. 7 to Fig. 10. Fig. 7 to Fig. 10 are diagrams showing an example of the speech assistance process according to this embodiment.

[0053] 7 to 9 show an example of "processing for generating and outputting corrected utterance data" as a first example, and FIG. 10 shows an example of "processing for generating and storing sample data" as a second example.

[0054] (First example) The first example of "processing for generating and outputting corrected utterance data" will be described first with reference to Fig. 7. The first example is an example in which a prompt based on text information input by a user is input to a generative model provided with prior knowledge to generate corrected utterance data.

[0055] FIG. 7 shows a speech assistance device 100 that generates corrected utterance data, a terminal device 200 that is operated by a user who makes an utterance, and a generation model 10 that is used by the speech assistance device 100.

[0056] First, the terminal device 200 receives text information such as an utterance "xxx..." from the user. Then, the terminal device 200 transmits the text information received from the user to the speech assistance device 100 ((1-1) in FIG. 7).

[0057] The speech support device 100 converts the text information ((1-1) in FIG. 7) received from the terminal device 200 into a prompt ((1-2) in FIG. 7). For example, the speech support device 100 converts text information such as a user utterance "XXX..." into a prompt to be input to a generation model, such as "Please generate speech data that corrects the utterance "XXX..." input by the user based on the sample data."

[0058] The speech assistance device 100 executes a process for generating corrected utterance data ((2) in FIG. 7). Specifically, the speech assistance device 100 inputs the converted prompt to the generation model 10, which has been provided with prior knowledge for correcting the user's utterance data ((2-1) in FIG. 7), and generates "corrected utterance data" by correcting the utterance indicated in the text information input by the user.

[0059] For example, the speech assistance device 100 generates corrected speech data that has been corrected or added in terms of punctuation, emoticons, expressions to be avoided, typos, content that is in line with the time, place, and occasion, explanations of the meanings of expressions, explanations of slang terms, etc. ((2-2) in FIG. 7). Examples of the above-mentioned corrections and additions will be described later in conjunction with the description of FIGS. 8 and 9.

[0060] Here, the speech assistance device 100 executes a process of outputting the corrected utterance data to the user ((3) in FIG. 7). Then, the speech assistance device 100 transmits the corrected utterance data, which has been corrected or has additional content in terms of punctuation marks, emoticons, expressions to be avoided, typos, content in line with the time, place, and occasion, explanations of the meanings of expressions, explanations of slang words, etc., to the terminal device 200 operated by the user. Then, the terminal device 200 displays the received corrected utterance data to the user ((3-1) in FIG. 7).

[0061] Next, an example of the above-mentioned corrected utterance data will be described with reference to Fig. 8 and Fig. 9. First, a description will be given with reference to Fig. 8. Fig. 8 shows an example of utterance data that has been corrected and has additional content in terms of punctuation marks, emoticons, expressions to be avoided, typos, explanations of the meanings of expressions, explanations of slang words, etc., and is displayed on the terminal device 200.

[0062] 8 shows the speech assistance device 100 and a terminal device 200 that displays corrected speech data transmitted from the speech assistance device 100. The terminal device 200 displays a screen ((1) in FIG. 8) that displays the dialogue history between users and a screen ((2) in FIG. 8) that allows the user to input text information.

[0063] As an example, the terminal device 200 displays, on the screen shown in (1) of Fig. 8, the utterance history by user A, "I've been working with AI recently," and the utterance history by user B, "That's crazy. Are you using 'generative AI'?" Then, the terminal device 200 displays information based on the corrected utterance data transmitted from the speech assistance device 100 on the above-mentioned screen.

[0064] For example, the terminal device 200 displays an explanation of the polysemous word "yabai" uttered by the user B shown in (1) of FIG. 8. Specifically ... <2> " is likely to be the meaning of ", ", " <1> A dangerous or inconvenient situation. A state of being unwell. <2> Information explaining the meaning of the word, such as "very interesting" or "interesting," is displayed ((1-1) in Figure 8).

[0065] Furthermore, the terminal device 200 displays an explanation of the meaning of the expression about "generative AI" uttered by the user B shown in (1) of Fig. 8. Specifically, the terminal device 200 displays information explaining the meaning such as "It refers to a large-scale language model that generates information according to input" ((1-2) of Fig. 8).

[0066] In addition, the terminal device 200 displays corrections from the perspectives of "typos" and "punctuation" for "Mr. B, even if it's difficult, you know it, which is amazing (emoji)," spoken by user A shown in (2) of Figure 8.

[0067] Specifically, the terminal device 200 displays a message indicating that the word "nandemo" may be a misspelling of "nani demo" ("difficult to do"), such as "It may be a misspelling of "nani demo" ("whatever")" ((2-1) of FIG. 8). The terminal device 200 also displays a message indicating that the use of punctuation marks may be inappropriate, such as "The number and position of punctuation marks may be inappropriate" ((2-2) of FIG. 8).

[0068] Furthermore, the terminal device 200 displays corrections from the perspective of "expressions to avoid" for "I bet you're good at studying (emoticon)" uttered by the user A shown in (2) of Fig. 8. Specifically, the terminal device 200 displays content that conveys that there is a possibility that expressions to avoid, such as "Depending on the context, the other person may feel that you are being sarcastic," and "How about expressions such as 'You've studied a lot, haven't you?'," are included, along with alternative expressions ((2-3) of Fig. 8).

[0069] Additionally, the terminal device 200 displays a message such as "Are you using too many pictograms?" to inform users that the use of specific symbols or pictograms may be inappropriate ((2-4) in FIG. 8).

[0070] Next, a description will be given with reference to Fig. 9. Fig. 9 shows an example of displayed speech data in which "corrections and additions have been made in terms of typographical errors, content appropriate for the time, place, and occasion, etc." The terminal device 200 displays a screen ((1) in Fig. 9) for the user to input text information.

[0071] As an example, the terminal device 200 displays an email screen on which the user A inputs text, as shown in (1) of Fig. 9. Then, the terminal device 200 displays information based on the corrected utterance data transmitted from the speech support device 100 on the above-mentioned screen.

[0072] For example, the terminal device 200 displays a message pointing out typos and omissions in the "email text" entered by user A shown in (1) of Fig. 9. Specifically, the terminal device 200 displays information such as "'Dono' is a title used by a superior to a subordinate, and may be considered rude when used with a customer," and "'Job title + Dono (title)' is generally considered incorrect usage" ((1-1) of Fig. 9).

[0073] Furthermore, the terminal device 200 displays comments on the "email text" entered by user A shown in (1) of Fig. 9 from the viewpoint of content that is appropriate for the time, place, and occasion. Specifically, the terminal device 200 displays information such as "In business emails, it may be inappropriate to use expressions such as 'Thank you for your hard work' when writing to a superior from an external company," "For example, how about the following text? <For the first time> ``Thank you very much for your continued support,'' <For the second time or later> ``Thank you very much for your continued support,'' "In addition, if the relationship is close, the opening sentence may be omitted" ((1-2) of Fig. 9).

[0074] (Second example) Next, a second example of "sample data generation processing and storage processing" will be described with reference to Fig. 10. The second example is an example of processing in which sample data is generated based on text information input by a user and stored in storage unit 120 (sample data DB 123) even when no instruction is given by the user or the like.

[0075] FIG. 10 shows a speech assistance device 100 that generates sample data, a terminal device 200 operated by an employee, and a generative model 10 used by the speech assistance device 100.

[0076] First, the terminal device 200 receives text information such as an employee's utterance "XXX..." as a business conversation between the employee. Then, the terminal device 200 transmits the text information received from the user to the speech assistance device 100 ((1-1) in FIG. 10).

[0077] The speech assistance device 100 converts the text information ((1-1) in FIG. 10) received from the terminal device 200 into a prompt ((1-2) in FIG. 10). For example, the speech assistance device 100 converts text information such as an employee's utterance "XXX..." into a prompt to be input to a generation model, such as "Please generate a speech history and example utterances using 'XXX...'."

[0078] The speech support device 100 executes a process for generating sample data ((2) in FIG. 10). Specifically, the speech support device 100 inputs the converted prompt to the generative model 10, which has been provided with prior knowledge for generating sample data ((2-1) in FIG. 10), and generates "speech history," "example utterances," and the like as sample data ((2-2) in FIG. 10).

[0079] The speech assistance device 100 stores the generated sample data in the sample data DB 123 ((3) in FIG. 10). Specifically, the speech assistance device 100 stores the speech history "D" ((3-1) in FIG. 9) and the example utterance "E" ((3-2) in FIG. 10) in the sample data DB 123.

[0080] (Operational support processing procedure) An example of speech assistance processing by the speech assistance device 100 according to this embodiment will now be described with reference to Fig. 11 and Fig. 12. Fig. 11 and Fig. 12 are diagrams showing flowcharts of the speech assistance processing according to this embodiment.

[0081] 11 shows an example of a processing procedure for "processing for generating and outputting corrected utterance data." Also, FIG. 12 shows an example of a processing procedure for "processing for generating and storing sample data."

[0082] First, an example of the processing procedure for the "processing for generating and outputting corrected utterance data" will be described with reference to Fig. 11. The generation unit 132 provides prior knowledge to the generative model (S101). Specifically, the generation unit 132 inputs information such as "utterance history" and "example utterances" stored in the sample data DB 123, and a prompt including an instruction to store the sample data, to the generative model to be used. Through the above-described processing, the generation unit 132 causes the generative model to generate sample data.

[0083] If the user does not perform an operation such as an instruction to start the generation process, the speech assistance device 100 waits for the process (No in S102). On the other hand, if the user performs an operation such as an instruction to start the generation process (Yes in S102), the speech assistance device 100 starts the process of S103.

[0084] The receiving unit 131 receives text information from the user (S103). Next, the generating unit 132 converts the input text information into a prompt (S104). Next, the generating unit 132 inputs the converted prompt into a generation model and generates corrected utterance data based on the sample data (S105). Next, the output unit 134 outputs the corrected utterance data to the user (S106). Then, the speech assistance device 100 ends the process.

[0085] Next, an example of a processing procedure for "sample data generation processing and storage processing" will be described with reference to FIG. 12. As in the processing described with reference to FIG. 11, the generation unit 132 provides prior knowledge to the generative model (S201). Next, the reception unit 131 receives text information from the user (S202). Next, the generation unit 132 converts the text information into a prompt (S203). Next, the generation unit 132 inputs the converted prompt into the generative model to generate sample data (S204).

[0086] The storage unit 133 stores the generated sample data in the memory unit 120 (sample data DB 123) (S205). Then, the speech assistance device 100 ends the process.

[0087] (effect) Hereinafter, a description will be given of the effects achieved by the speech assistance device 100 according to this embodiment. Conventionally, reference technologies that enable efficient communication based on information and communication technology have been known, but there are problems in realizing smooth communication.

[0088] Therefore, the generation unit 132 of the speech assistance device 100 according to this embodiment inputs a prompt based on the text information received from the user to the generation model provided with the sample data, and generates corrected utterance data by correcting the utterance shown in the text information based on the sample. The output unit 134 of the speech assistance device 100 outputs the generated corrected utterance data to the user.

[0089] Through the above-described processing, the speech assistance device 100 according to this embodiment can automatically perform predetermined correction processing, such as correcting punctuation, typos, and expressions, on the speech data of a user contained in text information of business-related exchanges between the user and other employees during work, etc., to generate corrected speech data. As a result, the speech assistance device 100 according to this embodiment can appropriately correct the content of speech uttered by the user when the users converse with each other, thereby achieving the effect of enabling smooth communication between users.

[0090] Furthermore, the speech assistance device 100 according to this embodiment achieves predetermined effects by executing the processes described below.

[0091] The generation unit 132 inputs a prompt including an instruction to generate corrected utterance data to the generation model to which the user's utterance history has been further provided, and generates corrected utterance data.

[0092] Through the above-described process, the speech assistance device 100 can correct the speech of the user without changing the speech characteristics of the user based on the attribute information of the user who made the utterance, the past speech history, etc. As a result, the speech assistance device 100 can suppress the awkwardness of the speech content and expression caused by the correction by generating corrected speech data by imitating the speech characteristics of the user using the past speech history, etc. Therefore, the speech assistance device 100 has the effect of enabling smooth communication between users.

[0093] Furthermore, the generation unit 132 performs at least one of the following on the utterance indicated in the text information: correcting punctuation marks, correcting specific symbols, correcting expressions that should be avoided, correcting typos or omissions, and correcting expressions according to the time, place, and situation. The generation unit 132 also generates corrected utterance data to which at least one of an explanation of an expression, an explanation of an expression in a predetermined culture, and an explanation of an expression with multiple meanings has been added.

[0094] Through the above-described processing, the speech assistance device 100 can correct the user's speech by correcting grammar such as punctuation, the usage of expressions such as appropriate use of symbols and pictograms, deleting inappropriate expressions, etc. Therefore, the speech assistance device 100 appropriately corrects the content of speech uttered by a user when the users converse with each other, thereby achieving the effect of enabling smooth communication between the users.

[0095] The generation unit 132 converts text information related to the employee's work, which is input by the employee on a daily basis, into a prompt including an instruction to generate sample data. Next, the generation unit 132 inputs the converted prompt into the generative model to which the sample data was provided, thereby generating the sample data. The storage unit 133 then stores the sample data generated by the generation unit 132 in the memory unit 120 (sample data DB 123).

[0096] Through the above-described process, even when there is no instruction for generation processing from a user or the like, the speech assistance device 100 can generate sample data for correcting utterances using text information related to conversations between employees and store the sample data in the storage unit 120 (sample data DB 123). That is, the speech assistance device 100 can automatically generate and store sample data for correcting utterances based on daily business conversation data. As a result, the speech assistance device 100 can use the stored sample data to improve the accuracy of user utterance correction, thereby achieving the effect of enabling smooth communication between users.

[0097] <Modification> The following describes modified examples realized by the speech assistance device 100 according to this embodiment.

[0098] (Data, etc.) The text information, prompts, sample data, corrected speech data, names of functional parts of speech assistance device 100, steps, processes, names of steps or processes, etc. used in the description of the above embodiments are merely examples and can be changed as desired.

[0099] For example, the employee information DB121 stores employee identification information and organization attribute information in association with "No," which is information identifying individual employee information, but is not limited to this. Furthermore, the text information DB122 stores employee identification information and text in association with "No," which is information identifying individual text information, but is not limited to this. Furthermore, the sample data DB123 stores employee identification information and a speech history, which is data of past speeches by users, in association with "No," which is identification information for each individual speech history. Furthermore, the sample data DB123 stores speech examples, which are data of speech made by employees to be used as reference examples for correction processing, in association with "No," which is identification information for each individual speech example.

[0100] (Combination of processing examples, etc.) The first and second examples according to the present embodiment described above are merely examples, and the present invention is not limited to the described contents.

[0101] (An example of a generative model) The speech assistance device 100 according to this embodiment can use a large-scale language model such as ChatGPT (registered trademark) (see, for example, Reference 1) or a large-scale language model such as tsuzumi (registered trademark) (see, for example, Reference 2) as a generative model.

[0102] (Reference 1):ChatGPT(OpenAI),<URL:https: / / openai.com / chatgpt> ,<Searched on March 29, 2020> (Reference 2): NTT's large-scale language model "tsuzumi",<URL:https: / / www.rd.ntt / research / LLM_tsuzumi.html> ,<Searched on March 29, 2020>

[0103] From now on, tsuzumi will be described as an example of a generation model used by the speech assistance device 100 according to this embodiment. Figures 13 to 15 are diagrams for explaining tsuzumi.

[0104] First, the concept of tsuzumi will be explained using Figure 13. tsuzumi is a small, energy-efficient large-scale language model that achieves the same level of accuracy as ChatGPT, a huge single large-scale language model that consumes a lot of power.

[0105] Tsuzumi is a large-scale language model that is compact by focusing on a high-quality corpus and supporting only English and Japanese, rather than on the amount of training data. Furthermore, for domain specialization, Tsuzumi can be fine-tuned and can integrate with external data by combining search and generative AI (Artificial Intelligence).

[0106] As mentioned above, tsuzumi is a compact generative model, making it possible to create and operate multiple large-scale language models with specific specialized fields and diverse personalities. Furthermore, by linking small, large-scale language models for each specialized field, such as "medical care," "retail," "construction," "local government," "technology," "travel," "culture," "religion," "art," "education," "finance," "legal affairs," and "banking" based on a specified network, tsuzumi achieves the formation of generative models that are high-performance, efficient, fault-tolerant, and highly democratic compared to conventional single, huge, large-scale language models.

[0107] Furthermore, tsuzumi allows flexible tuning such as the fine tuning mentioned above. Specifically, tsuzumi allows tuning by "prompt engineering" as shown in Figure 14, "full fine tuning," and "adapter tuning."

[0108] For example, "prompt engineering" shown in (1) of Figure 14 is a tuning method that uses prompts with information about a specific field added when setting prompts for the base model. In tuning using prompt engineering, the base model itself is not changed and only the prompts that are set are changed, so the learning cost can be reduced compared to other methods.

[0109] For example, "full fine-tuning" shown in Figure 14 (2) is a tuning method in which the base model is additionally trained or retrained using training data related to a specific field. In full fine-tuning, a tuning model specialized for a target field can be constructed by training the base model using training data for the target field. Therefore, full fine-tuning can improve the accuracy of inference and generation compared to other methods.

[0110] For example, "adapter tuning" shown in Figure 14 (3) is a tuning method that adds a submodule related to a specific field to the base model. Adapter tuning can improve the accuracy of inference and generation by fine-tuning the base model using a submodule related to the target field. Furthermore, because adapter tuning does not require retraining the base model, it can improve accuracy while reducing training costs compared to full fine tuning.

[0111] Furthermore, tsuzumi can build a model based on a “multi-adapter.” Here, the multi-adapter will be explained using FIG.

[0112] As explained using (3) in Figure 14, tsuzumi allows fine-tuning of the model by adding adapters (sub-modules) to the base model. tsuzumi can also add a combination of one or more adapters to the base model.

[0113] For example, as shown in Fig. 15, tsuzumi fine-tunes the foundation model using an adapter specialized for organization A, an adapter specialized for organization B, and an adapter specialized for organization C. As a result, even when data with different characteristics, such as data from organization A, organization B, and organization C, is input, tsuzumi can accurately perform inference and generation processing based on the model fine-tuned by the adapters specialized for each organization.

[0114] For example, organization A could be the "research laboratory," organization B could be the "sales department," and organization C could be the "entire company." In other words, even if the characteristics and granularity of the organizations differ, tsuzumi can make fine adjustments using an adapter appropriate for the organization.

[0115] As mentioned above, tsuzumi is small and energy-efficient, and by combining small, large-scale language models specialized for specific fields, and by implementing flexible tuning, it is possible to achieve both accuracy and cost when performing inference and generation processing compared to conventional huge, large-scale language models.

[0116] In other words, tsuzumi can be used as an appropriate generation model for executing specific processing, such as speech assistance processing by the speech assistance device 100 of this embodiment, corresponding to the purpose, objective, and target field of the processing.

[0117] (IOWN technology) The generative model used by the speech assistance device 100 according to the present embodiment described above may be realized using technology related to the Innovative Optical and Wireless Network (IOWN) technology.

[0118] Here, we will explain the IOWN technology. Figure 16 is a diagram explaining IOWN. As shown in Figure 16, the IOWN technology consists of three main technology fields: "All-Photonics Network (APN)," "Digital Twin Computing (DTC)," and "Cognitive Foundation (CF (Cognitive Foundation (registered trademark))."

[0119] (All Photonics Network) The APN related to IOWN technology is a technology that enables the construction of high-speed networks by processing all network transfer functions in the optical domain. Specifically, the APN related to IOWN technology is a technology that realizes low-power, high-quality, large-capacity, and low-latency communications based on optical-based (photonics-based) technologies such as "photonics-electronic convergence technology," "large-capacity optical transmission system and device technology," "optical Ising machine," and "optical lattice clock network."

[0120] (Digital Twin Computing) DTC, which is related to IOWN technology, is a technology that maps individual objects in the real world onto a virtual space using the vast amount of data collected by devices connected to the APN described above.

[0121] Conventional digital twin frameworks are used by mapping individual objects, such as automobiles and robots, into a virtual space, performing analysis and predictions on them, and then mapping the results of the analysis and predictions back onto the real world.

[0122] On the other hand, DTC related to IOWN technology expands on the conventional concept of digital twins, freely combining digital twins of various industries, objects, and people to perform calculations, thereby reproducing with high accuracy the combination of multiple objects, such as people and automobiles in a city. Furthermore, DTC related to IOWN technology enables not only the expression of a person's external appearance, but also the digital expression of their internal state, such as consciousness and thoughts, by combining technologies that enable "speech recognition," "speech synthesis," "understanding of emotions and intentions," etc. to collect information and build a digital twin environment.

[0123] In this way, DTC related to IOWN technology is a technology that enables the creation of digital twins that do not exist in the real world by combining multiple entities that are single in the real world and replicating them as digital twins in a virtual space, or by exchanging or merging some of the components between multiple digital twins.

[0124] (Cognitive Foundation) CF related to IOWN technology is a technology that centrally performs the deployment, configuration, linkage, management, and operation of ICT (Information and Communication Technology) resources at different layers, from the cloud to edge computers, network services, user equipment, etc. Specifically, CF related to IOWN technology treats various targets as a group of virtualized ICT resources, and optimally integrates multiple resources at different layers using multi-orchestration functions as a hub.

[0125] Furthermore, as shown in FIG. 16, the IOWN technology provides high-value-added services by linking the above-mentioned APN, DTC, and network services provided by operators.

[0126] For example, as shown in (1) of Figure 16, IOWN technology provides a technology for transmitting information collected via APN to other terminal devices at high speed and with low latency. Also, as shown in (2) of Figure 16, IOWN technology provides a technology for collecting large amounts of information from terminal devices and outputting information such as analysis results from the service provided by a business operator at high speed and with low latency in services such as information analysis. Also, as shown in (3) of Figure 16, IOWN technology provides a technology for transmitting large amounts of information at high speed and with low latency, using information obtained from surveillance cameras, automobile sensors, etc. to build a digital twin environment, make future predictions, and output the prediction results to the user.

[0127] The speech assistance device 100 according to this embodiment can efficiently realize speech assistance processing based on the generative model of tsuzumi or the like that is configured based on the IOWN technology that transmits data at high speed and with low latency as described above.

[0128] For example, the speech assistance device 100 can generate sample data and corrected speech data with higher accuracy than conventional methods based on tsuzumi, which is trained using text information related to its own organization or other organizations that is collected in large quantities at high speed and with low latency via a network built based on IOWN.

[0129] (Regarding business-related natural language text) In this embodiment, it has been described that the speech assistance device 100 uses “business-related natural language text.” The “business-related natural language text” does not mean to include only business-related text, but also broadly includes, for example, daily conversations, business-related meetings, and other natural language conversations between employees.

[0130] (Flowcharts, etc.) The steps in the flowcharts may be interchanged as long as there is no contradiction, and some steps may not be performed. In addition, conjunctions such as "next," "continue," "further," "at this time," and "on this occasion" used in the explanation of the flowcharts do not limit the order or timing of the execution of the processes in the flowcharts.

[0131] <Hardware configuration> The components of each device shown in the figure are conceptual functional units and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown, and all or part of each device can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions performed by each device can be realized by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware using wired logic.

[0132] Furthermore, among the processes described in this embodiment, all or part of the processes described as being performed automatically can also be performed manually using known methods. In addition, the information including the processing procedures, control procedures, specific names, various data, and parameters shown in the drawings can be changed as desired unless otherwise specified.

[0133] <Program> In one embodiment, the various devices constituting the speech assistance device 100 can be implemented by installing a speech assistance program as package software or online software on a desired computer. For example, by executing the speech assistance program on an information processing device, the various devices constituting the speech assistance device 100 can function. The information processing device referred to here includes desktop and notebook personal computers. In addition, the information processing device also includes mobile communication terminals such as smartphones and mobile phones, and even slate terminals such as PDAs (Personal Digital Assistants).

[0134] 17 is a diagram showing an example of a computer that executes speech assistance processing according to this embodiment. The computer 1000 includes, for example, a memory 1010 and a CPU 1020. The computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0135] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores, for example, a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.

[0136] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, application programs 1092, program modules 1093, and program data 1094. That is, programs that define the processes of the various devices that make up the speech assistance device 100 are implemented as program modules 1093 in which computer-executable codes are written. The program modules 1093 are stored, for example, in the hard disk drive 1090. For example, program modules 1093 for executing processes similar to the functional configurations of the various devices that make up the speech assistance device 100 are stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced with an SSD (Solid State Drive).

[0137] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads the program module 1093 or the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary, and executes the processing of the above-described embodiment.

[0138] The program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may also be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a LAN or a WAN (Wide Area Network)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.

[0139] <Other> Although the present embodiment has been described above, the present embodiment is not limited by the descriptions and drawings that form part of the disclosure. In other words, other embodiments, examples, operational techniques, etc. that are made by those skilled in the art based on the present embodiment are all included in the scope of the present embodiment. [Explanation of symbols]

[0140] 10 Generative Model 100 Speech assistive devices 110 Communications Department 120 Storage section 121 Employee Information DB 122 Text Information DB 123 Sample Data DB 130 Control Unit 131 Reception 132 Generation part 133 Storage area 134 Output section 200 Terminal Device

Claims

1. a generation unit that inputs a prompt based on a natural language text received from a user to a generation model provided with a sample of a sentence created in a business, and generates corrected utterance data by correcting the utterance shown in the text based on the sample; an output unit that outputs the generated corrected utterance data to the user; A speech assistance device comprising:

2. The generation unit inputting the prompt, including an instruction to generate the corrected utterance data, to the generative model to which the user's utterance history has been further provided, thereby generating the corrected utterance data; 2. The speech aid device according to claim 1.

3. The generation unit performing at least one of the following on the utterances shown in the text: correction of punctuation marks, correction of specific symbols, correction of expressions to be avoided, correction of typos or omissions, and correction of expressions appropriate for the time, place, and situation; 3. The speech aid device according to claim 2.

4. The generation unit generating the modified speech data to which at least one of an explanation of an expression, an explanation of an expression in a predetermined culture, and an explanation of an expression having multiple meanings has been added; 4. The speech aid device according to claim 2 or 3.

5. The generation unit converting natural language text related to the employee's work into prompts containing instructions for generating sample sentences to be produced in the work; inputting the converted prompt into a generative model provided with a sample of a sentence created in a business to generate the sample; Further comprising a storage unit that stores the generated sample in a storage unit.

2. The speech aid device according to claim 1.

6. The generation unit As the generative model, at least one of a large-scale language model having general knowledge and Tsuzumi is used.

6. The speech aid device according to claim 1, 2, 3 or 5.

7. The generation unit Using the generative model to communicate over a communication network related to an Innovative Optical and Wireless Network (IOWN), 7. The speech aid device according to claim 6.

8. A speech assistance method executed by a speech assistance device, comprising: a generation step of inputting a prompt based on a natural language text received from a user into a generation model provided with a sample of a sentence created in a business, and correcting the utterance shown in the text based on the sample to generate corrected utterance data; an output step of outputting the generated corrected utterance data to the user; A speech assistance method comprising:

9. a generation step of inputting a prompt based on a natural language text received from a user to a generation model provided with a sample of a sentence created in a business, and correcting the utterance shown in the text based on the sample to generate corrected utterance data; an output step of outputting the generated corrected utterance data to the user; A speech assistance program that causes a computer to execute the following.

Citation Information

Patent Citations

  • Programs, methods and devices for message management and document generation on device

    JP2023153881A