Speech analysis system
Through the information interaction and analysis between multiple voice analysis devices, the problem of low recognition accuracy of long-distance users is solved, and high-precision voice recognition effect is achieved.
Patent Information
- Application Number
- CN202310427070.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-24
- Filing Date
- 2020-12-15
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-12-15
AI Technical Summary
It is difficult for existing speech analysis systems to accurately identify conversations between users far away from speech analysis terminals, resulting in a decrease in recognition accuracy.
By using the analyzed dialogue information between multiple voice analysis devices for speech recognition, information interaction and analysis between multiple voice analysis terminals is used, and the recognition accuracy is improved by using relevant words, topics and time and frequency analysis.
High-precision recognition of user conversations far away from the speech analysis terminal is realized, and the accuracy and reliability of speech recognition are improved.
Smart Images

Figure CN116434755B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application number 202080054980.7 and the invention title "Speech Analysis System", which was filed on January 28, 2022. Technical Field
[0002] The present invention relates to a speech analysis system. Background Art
[0003] A system is described in Japanese Patent Application Laid-Open No. 2002-259635, which displays keywords in the speech of discussion participants during a discussion through a combination of graphic objects and text.
[0004] A presentation evaluation device using a speech analysis terminal is described in Japanese Patent Application Laid-Open No. 2017-224052.
[0005] However, when using one speech analysis terminal to perform speech recognition on a conversation, there are the following problems: Although it is possible to perform speech analysis on the conversation of a user close to the speech analysis terminal with relatively high accuracy, it is not possible to accurately perform speech analysis on the conversation of a user far from the speech analysis terminal.
[0006] On the other hand, a retrieval document information storage device is described in Japanese Patent No. 6646184.
[0007] Prior Art Documents
[0008] Patent Documents
[0009] Patent Document 1: Japanese Patent Application Laid-Open No. 2002-259635
[0010] Patent Document 2: Japanese Patent Application Laid-Open No. 2017-224052
[0011] Patent Document 3: Japanese Patent No. 6646184 Summary of the Invention
[0012] Technical Problem to be Solved by the Invention
[0013] An object of the invention according to one aspect described in this specification is to provide a speech analysis system capable of performing speech recognition with higher accuracy.
[0014] Technical Solution for Solving the Technical Problem
[0015] The invention according to one aspect is basically based on the following insight: By using the analyzed conversation information mutually between multiple speech analysis devices to perform speech recognition, it is possible to perform speech recognition with higher accuracy.
[0016] One aspect of the invention described in this specification relates to a speech analysis system 1.
[0017] The speech analysis system 1 is a system including a first speech analysis terminal 3 and a second speech analysis terminal 5. Each terminal includes a computer, and each structural element described below is a structural element installed by the computer. The system may also include a server.
[0018] The first speech analysis terminal 3 is a terminal including a first term analysis unit 7, a first conversation storage unit 9, a first analysis unit 11, a presentation storage unit 13, a related word storage unit 15, a display unit 17, and a topic word storage unit 19.
[0019] The first term analysis unit 7 is a structural element for analyzing the words included in a conversation and obtaining first conversation information.
[0020] The first conversation storage unit 9 is a structural element for storing the first conversation information analyzed by the first term analysis unit 7.
[0021] The first analysis unit 11 is a structural element for analyzing the first conversation information stored in the first conversation storage unit 9.
[0022] The presentation storage unit 13 is a structural element for storing a plurality of presentation materials.
[0023] The related word storage unit 15 is a structural element for storing related words associated with each presentation material stored in the presentation storage unit 13.
[0024] The display unit 17 is a structural element capable of displaying any one of the presentation materials stored in the presentation storage unit 13.
[0025] The topic word storage unit 19 is a structural element for storing topic words associated with the terms in the conversation.
[0026] The second speech analysis terminal 5 is a terminal including a second term analysis unit 21 and a second conversation storage unit 23.
[0027] The second term analysis unit 21 is a structural element for analyzing the words included in a conversation and obtaining second conversation information. The second conversation storage unit 23 is a structural element for storing the second conversation information analyzed by the second term analysis unit 21.
[0028] The first speech analysis terminal 3 further has a conversation information receiving unit 25.
[0029] Moreover, the conversation information receiving unit 25 is a structural element for receiving the second conversation information from the second speech analysis terminal 5. And the first conversation storage unit also stores the second conversation information received by the conversation information receiving unit 25.
[0030] The first analysis unit 11 includes a specific presentation information acquisition unit 31, a first dialogue segment acquisition unit 33, a specific related word reading unit 35, a first in-dialogue term extraction unit 37, a first topic word extraction unit 39, a second in-dialogue term extraction unit 41, a second topic word extraction unit 43, and a dialogue segment adoption unit 45.
[0031] The specific presentation information acquisition unit 31 is a structural element for receiving information that a specific presentation material has been selected, where the specific presentation material is one of a plurality of presentation materials.
[0032] The first dialogue segment acquisition unit 33 is a structural element for analyzing dialogue segments in the first dialogue information to obtain one or more dialogue segments.
[0033] The specific related word reading unit 35 is a structural element for reading specific related words from the related word storage unit 15, where the specific related words are related words associated with the specific presentation material.
[0034] The first in-dialogue term extraction unit 37 is a structural element for extracting terms in the first dialogue, where the terms in the first dialogue are terms in the dialogue analyzed by the first analysis unit 11 included in the first dialogue segment, and the first dialogue segment is a certain dialogue segment in the first dialogue information.
[0035] The first topic word extraction unit 39 is a structural element for extracting the first topic word from the topic word storage unit 19, where the first topic word is a topic word associated with the terms in the first dialogue.
[0036] The second in-dialogue term extraction unit 41 is a structural element for extracting terms in the second dialogue, where the terms in the second dialogue are terms in the dialogue included in the second dialogue segment, and the second dialogue segment is the dialogue segment in the second dialogue information corresponding to the first dialogue segment.
[0037] The second topic word extraction unit 43 is a structural element for extracting the second topic word from the topic word storage unit 19, where the second topic word is a topic word associated with the terms in the second dialogue.
[0038] The dialogue segment adoption unit 45 is a structural element for using the relationship between the first topic word and the specific related word and the relationship between the second topic word and the specific related word to adopt the first dialogue segment or the second dialogue segment as the correct dialogue segment.
[0039] The dialogue segment adoption unit 45 can also be configured as follows. That is, when the first topic word and the second topic word are different, if the first topic word is a specific related word and the second topic word is not a specific related word, the dialogue segment adoption unit 45 adopts the first dialogue segment in the first dialogue information as the correct dialogue segment; if the first topic word is not a specific related word and the second topic word is a specific related word, the dialogue segment adoption unit 45 adopts the second dialogue segment in the second dialogue information as the correct dialogue segment.
[0040] The dialogue segment adoption unit 45 can also be configured as follows. That is, the dialogue segment adoption unit 45 compares the number of the first topic words that are specific related words with the number of the second topic words that are specific related words. If the former is more, it adopts the first dialogue segment as the correct dialogue segment; if the latter is more, it adopts the second dialogue segment as the correct dialogue segment.
[0041] A preferred mode of the speech analysis system 1 is that the first speech analysis terminal 3 further has a time storage unit 51, and this time storage unit 51 is used to store the time.
[0042] This system is such that the first dialogue information includes the words included in the dialogue and the time associated with each word. The first dialogue segment acquisition unit 33 uses the time information of each word to analyze the dialogue segment.
[0043] When the dialogue is interrupted, it can be known that the speaker has changed. Therefore, when there is a time interval between words, it can be known that the dialogue segment has changed.
[0044] A preferred mode of the speech analysis system 1 is that the first speech analysis terminal 3 further has a frequency analysis unit 53, and this frequency analysis unit 53 analyzes the frequency of the speech included in the dialogue.
[0045] This system is such that the first dialogue information includes the words included in the dialogue and the frequency of the speech associated with each word.
[0046] The first dialogue segment acquisition unit 33 uses the frequency of each word to analyze the dialogue segment.
[0047] When the pitch of the voice changes, it can be known that the speaker has changed. Therefore, by analyzing the frequency of the voice of each word, it can be known that the dialogue segment has changed.
[0048] A preferred mode of the speech analysis system 1 is that the related words stored in the related word storage unit 15 include presenter-related words and listener-related words. The first dialogue segment acquisition unit 33 uses the presenter-related words and listener-related words included in the dialogue information to analyze the dialogue segment.
[0049] The terms used by the presenter and the terms used by the listener in the presentation are different. Therefore, it is possible to analyze the dialogue segments using their respective terms.
[0050] A preferred mode of the voice analysis system 1 is that the first voice analysis terminal 3 further has a mis-converted term storage unit 55 that stores mis-converted terms associated with each of the multiple presentation materials.
[0051] Moreover, in the case where mis-converted terms related to a specific presentation material are included, from the terms included in the dialogue segments that are not adopted as correct dialogue segments among the respective dialogue segments, the first analysis unit 11 uses the terms corresponding to the mis-converted terms included in the correct dialogue segments to correct the terms included in the correct dialogue segments. The first voice analysis terminal 3 and the second voice analysis terminal 5 can obtain highly accurate analysis results by comparing information with each other.
[0052] Advantages of the Invention
[0053] According to the invention of one mode described in this specification, by performing voice recognition by mutually using the analyzed dialogue information among multiple voice analysis devices, voice recognition can be performed with higher accuracy. Brief Description of the Drawings
[0054] Figure 1 It is a block diagram showing a structural example of the voice analysis system.
[0055] Figure 2 It is a flowchart showing a processing example of the voice analysis system.
[0056] Figure 3 It is a conceptual diagram showing a processing example of the voice analysis system.
[0057] Figure 4 It is a conceptual diagram showing a second processing example of the voice analysis system.
[0058] Figure 5 It is a conceptual diagram showing a third processing example of the voice analysis system. Detailed Description of the Embodiment
[0059] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The present invention is not limited to the embodiments described below, and also includes embodiments appropriately modified within an obvious range by those skilled in the art based on the following embodiments.
[0060] One aspect of the invention described in this specification relates to a speech analysis system 1. The speech analysis system is used to receive speech information such as conversations as input information, analyze the received speech information, and obtain conversation sentences. The speech analysis system is installed by a computer. In addition, a system for converting speech information into text information is well-known, and the present invention can suitably use the structure of such a well-known system. This system can be installed by a mobile terminal (such as a mobile phone or other computer terminal), or can be installed by a computer or a server. The computer may also have a processor that implements various functions.
[0061] Figure 1 It is a block diagram showing a structural example of the speech analysis system. The speech analysis system 1 is a system including a first speech analysis terminal 3 and a second speech analysis terminal 5. Each terminal includes a computer, and each structural element described below is a structural element installed by the computer.
[0062] The computer has an input / output unit, a control unit, an arithmetic unit, and a storage unit. Each structural element is connected by a bus or the like to enable information transmission and reception. And, for example, the control unit reads a control program stored in the storage unit, uses the information stored in the storage unit or the information input from the input / output unit, and causes the arithmetic unit to perform various operations. The information calculated by the arithmetic unit is output from the input / output unit in addition to being stored in the storage unit. In this way, various arithmetic processes are performed. Each structural element described below may also correspond to any structural element of the computer.
[0063] The first speech analysis terminal 3 is a terminal including a first term analysis unit 7, a first conversation storage unit 9, a first analysis unit 11, a presentation storage unit 13, a related word storage unit 15, a display unit 17, a topic word storage unit 19, and a conversation information receiving unit 25.
[0064] The first term analysis unit 7 is a structural element for analyzing the words included in the dialogue to obtain the first dialogue information. For example, voice is input to the first voice analysis terminal 3 through a microphone. Then, the first voice analysis terminal 3 stores the dialogue (voice) in the storage unit. The first term analysis unit 7 analyzes the words included in the dialogue to obtain the first dialogue information. The first dialogue information is the information obtained by modifying the voice into voice information. An example of the voice information is "jie xia lai dui zuo wei guan yutang niao bing de xin yao ai ke si wai zei jin xing shuo ming,gai xin yaoneng jiang di xue tang zhi ma? ". For example, the digitized voice data can be read out from the storage unit of the computer, the program can be read out from the storage unit, and according to the instructions of the program, the arithmetic unit analyzes the read voice data.
[0065] The first dialogue storage unit 9 is a structural element for storing the first dialogue information analyzed by the first term analysis unit 7. For example, the storage unit of the computer functions as the first dialogue storage unit 9. The first dialogue storage unit 9 stores the above-mentioned dialogue information in the storage unit.
[0066] The first analysis unit 11 is a structural element for analyzing the first dialogue information stored in the first dialogue storage unit 9. The first analysis unit 11 reads out the voice information stored in the storage unit, retrieves it in the terms stored in the storage unit, and converts it into appropriate terms. At this time, in the case of having convertible terms (homophones), the conversion efficiency can also be improved by selecting the term with a high frequency of use together with other terms. For example, "tang niao bing" is converted to "diabetes". And the candidate conversion terms for "xin yao" are "new drug", "Xinyao", "essential requirement", "Xinyao". Select the "new drug" with a high frequency of co-occurrence with "diabetes" among them as the term included in the dialogue information. Then, the voice information stored in the storage unit is analyzed into a dialogue sentence like "Next, an explanation will be given about XYZ as a new drug for diabetes. Can this new drug lower blood sugar levels? ". Then, the analyzed dialogue sentence is stored in the storage unit.
[0067] The first analysis unit 11 can also use relevant words read out in association with the presentation materials to improve the analysis accuracy of the dialogue information. For example, in a part where there is dialogue information of "xin yao" and "new drug" is among the relevant words, this "xin yao" can be analyzed and "new drug" can be selected. In this way, the analysis accuracy can be improved. Additionally, it can be that when multiple readings are assigned to the relevant words and these readings are included in the dialogue information, the corresponding relevant words are selected. For example, regarding the relevant word "XYZ", the candidate readings are "ai ke si wai zei", "ai ke si wai zai de", "ai ke zi waizei", and "ai ke zi wai zai de".
[0068] The presentation storage unit 13 is a structural element for storing multiple presentation materials. For example, the storage unit of a computer functions as the presentation storage unit. Examples of presentation materials are the pages of PPT (Powerpoint: registered trademark). Presentation materials are stored in a computer and can be displayed on a display unit to give a presentation to a dialogue partner or an audience.
[0069] The relevant word storage unit 15 is a structural element for storing relevant words that are associated with each of the presentation materials stored in the presentation storage unit 13. The relevant words can be terms associated with an entire file as a presentation material, or can be terms associated with a certain page included in a file. For example, the storage unit of a computer functions as the relevant word storage mechanism. Examples of multiple relevant words associated with presentation materials are terms that may be used when explaining according to the pages of PPT. The storage unit stores multiple relevant words in association with presentation materials such as PPT. The storage unit stores multiple relevant words associated with a presentation material in association with the information of the presentation material (such as file ID, page number). Examples of relevant words are "diabetes", "new drug", "XYZ", "ABC" (names of other therapeutic drugs), "blood glucose level", "side effect", "blood sugar", "glaucoma", "retinopathy", "insulin", "DC Pharmaceutical", "attachment". Such relevant words can, for example, also be input into the computer by the user and stored by the storage unit. Additionally, such relevant words can also be automatically retrieved by the computer on websites regarding relevant words such as "XYZ", and the terms included in the retrieved websites are automatically stored in the storage unit to appropriately update the relevant words regarding a certain presentation material.
[0070] The display unit 17 is a structural element capable of displaying any of the presentation materials stored in the presentation storage unit 13. An example of the display unit 17 is the output unit of a computer, specifically, a monitor or a display. The computer reads information about the presentation materials stored in the storage unit and displays the presentation materials on the monitor or the screen. In this way, the presentation materials can be displayed to the conversation partner or the audience.
[0071] The topic word storage unit 19 is a structural element for storing topic words associated with the words used in the conversation. The words used in the conversation are, for example, the words that become keywords among the words used in the conversation. The topic word storage unit 19 is an institution for storing topic words associated with the words (keywords) used in the conversation. The topic word storage unit 19 can be implemented by a storage unit and a structural element (such as a control program) for reading information from the storage unit.
[0072] For example, in the topic word storage unit 19, the topic word "obesity" can be stored in association with keywords such as "obesity gene", "obesity", and "obesity experimental animals" that are assumed to be used in the conversation. The topic word can also be a word obtained by further unifying multiple keywords or a superordinate concept word. By using the topic word, retrieval can be performed more quickly. Examples of the topic word are disease names, drug names, active ingredient names, and pharmaceutical company names. That is, the topic word can be said to be the second conversion word for the words used in the conversation. The topic word can also be a word obtained by assigning words suitable for retrieval to multiple keywords. In addition, the topic word can also be a word about the message.
[0073] The second speech analysis terminal 5 is a terminal including a second word analysis unit 21 and a second conversation storage unit 23. For example, the first speech analysis terminal 3 is a laptop computer or the like carried by an explainer such as an MR, and is a device located near the person giving the explanation and used to reliably record the voice of the explainer. On the other hand, the second speech analysis terminal 5 is, for example, a device such as a microphone or a mobile terminal (mobile phone, smartphone, etc.) that is set at a position closer to the audience than the explainer, such as a position closer to the doctor than the MR, and is used to more reliably record the voice of the person listening to the explanation. The second speech analysis terminal 5 can communicate information with the first speech analysis terminal 3. For example, the first speech analysis terminal 3 and the second speech analysis terminal 5 can directly communicate information, or can communicate information through a server.
[0074] The second language analysis unit 21 is a structural element for analyzing the words included in a conversation to obtain second conversation information. An example of the second conversation information is "jie xia lai dui zuo wei guan yu tang niao bing de xin yaoai ke si wai zei jin xing shuo ming,gai xin yao neng jiang di xue tang zhima?". The second voice analysis terminal 5 stores the conversation input from a microphone or the like in a storage unit. Then, the second language analysis unit 21 reads the conversation from the storage unit and refers to the words stored in the storage unit to obtain conversation information. An example of the second conversation information is "Next, an explanation will be given regarding the new drug XYZ for diabetes. Can this new drug lower blood sugar levels?".
[0075] The second conversation storage unit 23 is a structural element for storing the second conversation information analyzed by the second language analysis unit 21. The storage unit functions as the second conversation storage unit 23. That is, the second conversation information is stored in the storage unit of the second voice analysis terminal 5. The second conversation information stored in the storage unit of the second voice analysis terminal 5 is sent to the first voice analysis terminal 3, for example, through an output unit such as an antenna of the second voice analysis terminal 5.
[0076] Then, the first voice analysis terminal 3 receives the second conversation information sent from the second voice analysis terminal 5. The conversation information receiving unit 25 of the first voice analysis terminal 3 is a structural element for receiving the second conversation information from the second voice analysis terminal 5. For example, the antenna of the first voice analysis terminal 3 functions as the conversation information receiving unit 25. The second voice conversation information is input into the first voice analysis terminal 3 through the conversation information receiving unit 25 and stored in the storage unit. At this time, for example, the first conversation storage unit may also store the second conversation information received by the conversation information receiving unit 25.
[0077] The first analysis unit 11 includes a specific demonstration information acquisition unit 31, a first conversation segment acquisition unit 33, a specific related word reading unit 35, a first in-conversation word extraction unit 37, a first topic word extraction unit 39, a second in-conversation word extraction unit 41, a second topic word extraction unit 43, and a conversation segment adoption unit 45.
[0078] The specific presentation information acquisition unit 31 is a component for receiving information regarding the selection of a specific presentation material, which is a presentation material from among multiple presentation materials. For example, a MR selects a PowerPoint (registered trademark) document about a new diabetes drug, XYZ. Information indicating that the page has been selected is then input into the computer via the computer's input device. This input information is then used as information regarding the selection of the specific presentation material.
[0079] The first conversation segment acquisition unit 33 is a component for analyzing the conversation segments in the first conversation information to obtain one or more conversation segments. The first conversation segment acquisition unit 33 may also analyze the conversation segments in the second conversation information to obtain one or more conversation segments. A conversation segment is typically a portion of a conversation demarcated by a period (.). A conversation segment can also be a sentence. Furthermore, a conversation segment can be modified when the speaker changes. Of course, depending on the conversation, the segment may not necessarily be identical to the written language.
[0080] For example, the sentence "Is there a problem with the quality of life in the university?" can be divided into two dialogue segments: "Is there a problem with the quality of life in the university?" and "Is there a problem with the quality of life in the university?" Alternatively, the sentence "Next, I will explain XYZ, a new diabetes drug. Can this new drug reduce the cost of schoolwork?" can be divided into two dialogue segments: "Next, I will explain XYZ, a new diabetes drug." and "Can this new drug reduce the cost of schoolwork?" Methods for obtaining such dialogue segments are well known.
[0081] The specific related word reading unit 35 is a structural element for reading specific related words from the related word storage unit 15, where the specific related words are related words associated with specific presentation materials. The storage unit of the computer functions as the related word storage unit 15. Also, the specific related word reading unit 35 uses information about specific presentation materials to read, from the storage unit acting as the related word storage unit 15, the related words stored in association with the specific presentation materials as specific related words. The specific related words can be one or more. The read specific related words can also be appropriately stored in the storage unit. For example, the arithmetic unit and the storage unit of the computer function as the specific related word reading unit 35.
[0082] For example, in the related word storage unit 15, “diabetes,” “new drug,” “XYZ,” “ABC” (the name of another therapeutic drug), “blood sugar level,” “side effects,” “blood glucose,” “glaucoma,” “retinopathy,” “insulin,” “DC Pharmaceutical,” and “attachment” are stored in association with a PPT (PowerPoint: registered trademark) material about XYZ, which is a new drug for a certain type of diabetes. Therefore, the specific related word reading unit 35 reads these terms associated with the PPT (PowerPoint: registered trademark) material about XYZ as specific related words and stores them in the storage unit.
[0083] The first dialogue phrase extraction unit 37 is a structural element for extracting phrases in the first dialogue, where the phrases in the first dialogue refer to the phrases in the dialogue analyzed by the first analysis unit 11 included in the first dialogue segment, and the first dialogue segment is a certain dialogue segment in the first dialogue information. The phrases in the dialogue are phrases included in the dialogue. Also, the phrases in the dialogue analyzed by the first analysis unit 11 and included in the first dialogue segment of the first dialogue information are the phrases in the first dialogue. The phrases in the dialogue analyzed by the first analysis unit 11 are, for example, stored in the storage unit. The first dialogue phrase extraction unit 37 can read the phrases in the dialogue included in the first dialogue segment from the phrases in the dialogue stored in the storage unit and store them in the storage unit. In this way, the phrases in the first dialogue can be extracted. For example, the arithmetic unit and the storage unit of the computer function as the first dialogue phrase extraction unit 37.
[0084] For example, take “Next, an explanation will be given about XYZ, which is a new drug for diabetes.” and “Can this new drug lower blood sugar?” As the first dialogue segment is “Next, an explanation will be given about XYZ, which is a new drug for diabetes.” And “Can this new drug lower blood sugar?” is the dialogue segment following the first dialogue segment. The first dialogue phrases “diabetes,” “new drug,” and “XYZ” are included in the first dialogue segment.
[0085] The first topic word extraction unit 39 is a structural element for extracting the first topic word from the topic word storage unit 19. The first topic word is a topic word associated with the terms used in the first conversation. Topic words associated with the terms used in the conversation are stored in the topic word storage unit 19. Therefore, the terms used in the first conversation are read out from the storage unit, and the first topic word, which is a topic word associated with the terms used in the first conversation, is extracted from the topic word storage unit 19 using the read-out terms used in the first conversation. For example, the arithmetic unit and the storage unit of a computer function as the first topic word extraction unit 39.
[0086] As described above, examples of the terms used in the first conversation are "diabetes", "new drug", and "XYZ". The common topic word for these terms is "XYZ". Multiple topic words can also be extracted according to the terms used in each conversation.
[0087] The second conversation term extraction unit 41 is a structural element for extracting the terms used in the second conversation. The terms used in the second conversation are the terms used in the conversation included in the second conversation segment, and the second conversation segment is the conversation segment corresponding to the first conversation segment in the second conversation information. For example, the first analysis unit 11 analyzes the terms used in the conversation included in the second conversation information. In addition, the second voice analysis terminal may also have a second analysis unit to analyze the terms used in the conversation included in the second conversation information. In this case, the second voice analysis terminal may also send the terms used in the conversation included in the analyzed second conversation information to the first voice analysis terminal. In addition, the terms used in the conversation included in the analyzed second conversation information may also be sent to the server.
[0088] An example of the second conversation segment is "Next, an explanation will be given of XYZ, which is a new drug related to Guan Yu Tang Niao Bing". Examples of the terms used in the second conversation are "Tang Niao", "Bing", "Guan Yu", "New Yao", "XYZ".
[0089] The second topic word extraction unit 43 is a structural element for extracting the second topic word from the topic word storage unit 19. The second topic word is a topic word associated with the terms used in the second conversation. The second topic word extraction unit 43 is the same as the first topic word extraction unit 39.
[0090] Examples of the terms used in the second conversation are "Tang Niao", "Bing", "Guan Yu", "New Yao", "XYZ". An example of the topic word common to more terms among the topic words related to these terms is "Bible".
[0091] The dialogue segment adoption unit 45 is a structural element that uses the relationships between the first topic word and specific related words, and between the second topic word and specific related words, to adopt either the first dialogue segment or the second dialogue segment as the correct dialogue segment. The first topic word, the second topic word, and the specific related words can be read from the storage unit, and their relationships are analyzed using the arithmetic unit. Based on the analysis results, either the first dialogue segment or the second dialogue segment is adopted as the correct dialogue segment. For example, the arithmetic unit and storage unit of a computer function as the dialogue segment adoption unit 45. In this way, instead of using the relationship between the terms and specific related words in the dialogue, the relationship between the topic words associated with the terms in the dialogue and the specific related words is used to adopt the correct dialogue segment, so that the correct dialogue segment can be adopted objectively and with high precision. Especially when using the terms in the dialogue, misconversions sometimes occur, but if topic words are used, their association with specific related words is stronger, so the correct dialogue segment can be adopted with higher precision.
[0092] For example, the first topic word is "XYZ", the second topic word is "Bible", and the specific related words are "diabetes", "new drug", "XYZ", "ABC" (the names of other therapeutic drugs), "blood glucose level", "side effect", "blood sugar", "glaucoma", "retinopathy", "insulin", "DC Pharmaceutical", "appendix". In this case, since the first topic word "XYZ" is identical to one of the specific related words, the first dialogue segment is adopted as the correct dialogue segment.
[0093] In addition, when both the first topic word and the second topic word are among the specific related words, coefficients (grades) are added to the specific related words and stored in the storage unit, and the dialogue segment with the topic word identical to the specific related word with a higher grade is adopted as the correct dialogue segment.
[0094] The dialogue segment adoption unit 45 can also be configured as follows. That is, when the first topic word and the second topic word are different, if the first topic word is a specific related word and the second topic word is not a specific related word, the dialogue segment adoption unit 45 adopts the first dialogue segment in the first dialogue information as the correct dialogue segment. If the first topic word is not a specific related word and the second topic word is a specific related word, the dialogue segment adoption unit 45 adopts the second dialogue segment in the second dialogue information as the correct dialogue segment. For example, the first topic word, the second topic word, and the specific related word are read out from the storage unit. Then, the arithmetic unit is made to perform a process of determining whether the first topic word matches the specific related word. In addition, the arithmetic unit is made to perform a process of determining whether the second topic word matches the specific related word. At this time, the arithmetic unit is made to perform an arithmetic process of determining whether the first topic word is the same as the second topic word. And when the arithmetic unit determines that the first topic word is a specific related word and the second topic word is not a specific related word, the first dialogue segment in the first dialogue information is adopted as the correct dialogue segment, and the result is stored in the storage unit. On the other hand, when the first topic word is not a specific related word and the second topic word is a specific related word, the second dialogue segment in the second dialogue information is adopted as the correct dialogue segment and stored in the storage unit. In this way, the first dialogue segment or the second dialogue segment can be adopted as the correct dialogue segment.
[0095] The dialogue segment adoption unit 45 can also be configured as follows. That is, the dialogue segment adoption unit 45 compares the number of first topic words that are specific related words with the number of second topic words that are specific related words, and adopts the first dialogue segment as the correct dialogue segment when the former is larger, and adopts the second dialogue segment as the correct dialogue segment when the latter is larger.
[0096] For example, all topic words associated with the terms in the first dialogue and the terms in the second dialogue can be read out, the number of topic words that match the specific related word among the read multiple topic words can be measured, and the dialogue segment with a larger number is adopted as the correct dialogue segment. In addition, coefficients can be respectively added to the specific related words, and a higher score can be given when a specific related word with a higher degree of association with the presentation material matches a topic word, and the dialogue segment with a higher score is adopted as the correct dialogue segment.
[0097] Next, an example (embodiment) of a method for obtaining a dialogue segment will be described. A preferred mode of the speech analysis system 1 is that the first speech analysis terminal 3 further includes a time storage unit 51 for storing a time or a moment. In this system, the first dialogue information includes the words included in the dialogue and the time associated with each word. The first dialogue segment acquisition unit 33 analyzes the dialogue segment using the time information of each word. For example, when the silent state continues for a certain period of time or more after the speech has continued for a certain period of time, it can be said that the dialogue segment has changed. If there is a time interval between words, it can be known that the dialogue segment has changed. In this case, for example, the storage unit of the computer causes the first dialogue storage unit to store the first dialogue information, and causes the time storage unit 51 to correspondingly store the time of each piece of information regarding the first dialogue information. Then, for example, when analyzing the first dialogue information, the first analysis unit 11 can read out the time of each dialogue information, thereby obtaining the time interval. Then, the threshold value stored in the storage unit is read out, and the read threshold value is compared with the obtained time interval. If the time interval is greater than the threshold value, it is determined that it is a dialogue segment. Additionally, it is preferred that the second speech analysis terminal 5 also has a second time storage unit for storing a time or a moment. Thus, by comparing the times of the dialogues, the correspondence relationship between each segment of the first dialogue information and each segment of the second dialogue information can be grasped.
[0098] A preferred mode of the speech analysis system 1 is that the first speech analysis terminal 3 further includes a frequency analysis unit 53 for analyzing the frequency of the speech included in the dialogue. In this system, the first dialogue information includes the words included in the dialogue and the frequency of the speech associated with each word. The first dialogue segment acquisition unit 33 analyzes the dialogue segment using the frequency of each word. When the pitch of the sound changes, it can be known that the speaker has changed. Therefore, by analyzing the frequency of the sound of each word, it can be known that the dialogue segment has changed. In this case, the frequency information of the speech is also stored in the storage unit in association with each piece of information included in the dialogue information. The first analysis unit 11 reads out the frequency information stored in the storage unit, calculates the change in frequency, and thereby calculates the dialogue segment. Additionally, it can also be that the storage unit pre-stores the terms that form a dialogue segment, and when the dialogue information includes the terms that form the dialogue segment, it is determined that it is a dialogue segment. Examples of such terms that form a dialogue segment are "Ah.", "What?", "Right?", "Of.", "Well.", "Huh?", "Right?", "Ah.", "Well...".
[0099] A preferred mode of the speech analysis system 1 is that the related words stored in the related word storage unit 15 include presenter-related words and listener-related words. The first dialogue segment acquisition unit 33 analyzes the dialogue segment using the presenter-related words and listener-related words included in the dialogue information.
[0100] The terms used by the presenter and the terms used by the listener in the speech are different. Therefore, the respective terms can be used to analyze the dialogue segment.
[0101] The specific related word reading unit 35 is a structural element for extracting related words related to specific presentation materials included in the first dialogue information and the second dialogue information.
[0102] For example, associated with the name (location) and page number of a certain presentation material, the storage unit stores "diabetes", "new drug", "XYZ", "ABC" (names of other therapeutic drugs), "blood glucose level", "side effects", "blood sugar", "glaucoma", "retinopathy", "insulin", "DC Pharmaceutical", "appendix". Therefore, the specific related word reading unit 35 reads out the related words related to these specific presentation materials from the storage unit. Then, an arithmetic process is performed to check whether the terms included in the first dialogue information are consistent with the related words. Then, the consistent related words are stored in the storage unit together with the dialogue information and the segment number.
[0103] For example, the first dialogue information consists of 2 dialogue segments. In the first dialogue segment "Next, an explanation will be given about XYZ, a new drug for diabetes.", there are 3 related words: "diabetes", "new drug", and "XYZ". On the other hand, there are no related words in the second dialogue segment of the first dialogue information. Regarding the first dialogue segment of the first dialogue information, the first voice analysis terminal 3 stores, for example, the related words "diabetes", "new drug", and "XYZ" and the value 3. Additionally, regarding this dialogue segment, it is also possible to store only the value 3, or only the related words. The same applies to the second dialogue segment or the next second dialogue information.
[0104] In the first dialogue segment of the second dialogue information, that is, "Next, an explanation will be given about XYZ, a new drug for fish soup birds.", the related word "XYZ" is included. On the other hand, in the second dialogue segment of the second dialogue information, that is, "Can this new drug lower the blood glucose level?", the related word "blood glucose level" is included.
[0105] The system of this method is such that the second voice analysis terminal is also a terminal that can obtain accurate dialogue segments in the same way as the first voice analysis terminal. Therefore, the processing of each structural element is the same as in the above method.
[0106] One method described in this specification relates to a server-client system. In this case, for example, it is also possible to make the first mobile terminal have a display unit 17, and the server undertakes the functions of any one or two or more of the first term analysis unit 7, the first dialogue storage unit 9, the first analysis unit 11, the presentation storage unit 13, the related word storage unit 15, the topic word storage unit 19, and the dialogue information receiving unit 25.
[0107] One aspect described in this specification relates to a program. This program is a program for causing a computer or a processor of a computer to function as a first language analysis unit 7, a first dialogue storage unit 9, a first analysis unit 11, a presentation storage unit 13, a related word storage unit 15, a display unit 17, a topic word storage unit 19, and a dialogue information receiving unit 25. This program can be a program for installing the above-described various systems. This program can also be in the form of an application program installed in a mobile terminal.
[0108] One aspect described in this specification relates to a computer-readable information storage medium storing the above program. Examples of the information storage medium are CD-ROM, DVD, floppy disk, memory card, and memory stick.
[0109] Figure 2 It is a flowchart showing a processing example of a speech analysis system. Figure 3 It is a conceptual diagram showing a processing example of a speech analysis system. The above program is installed in two mobile terminals. One terminal is, for example, an MR's laptop computer, and the other mobile terminal is a smartphone, which is placed near a doctor as the other party to easily collect the other party's speech. The application program for installing the above program is installed in the laptop computer or the smartphone.
[0110] Selection process of presentation materials (S101)
[0111] MR opens a certain PPT (PowerPoint: registered trademark) saved in the laptop computer or read from a server. Then, information about the selection of this PPT (PowerPoint: registered trademark) is input into the computer.
[0112] Display process of presentation materials (S102)
[0113] Pages of presentation materials made with this PPT (PowerPoint: registered trademark) are displayed on the display unit of the laptop computer. On the other hand, pages of PPT (PowerPoint: registered trademark) are also displayed on the display unit of the smartphone.
[0114] Reading process of related words of presentation materials (S103)
[0115] On the other hand, specific related words associated with the presentation materials made with PPT (registered trademark) are read out from the storage unit. Examples of the read specific related words are "diabetes", "new drug", "XYZ", "ABC" (names of other therapeutic drugs), "blood glucose level", "side effect", "blood sugar", "glaucoma", "retinopathy", "insulin", "DC Pharmaceutical", "attachment". The read specific related words are appropriately temporarily stored in the storage unit.
[0116] Dialogue based on presentation materials (S104)
[0117] Associated with the displayed materials, a dialogue takes place between the MR and the physician. The dialogue can be a demonstration or an explanation. Examples of the dialogue are "Next, an explanation will be given about XYZ, a new drug for diabetes." and "Can this new drug lower blood sugar levels?" ( Figure 3 ).
[0118] First dialogue information acquisition process (S105)
[0119] The laptop records the dialogue and inputs it into the computer. Then, the words contained in the dialogue are analyzed to obtain the first dialogue information. An example of the first dialogue information before analysis is "jie xia lai dui zuo wei guan yu tangniao bing de xin yao ai ke si wai zei jin xing shuo ming,gai xin yao nengjiang di xue tang zhi ma?". The laptop is set on the MR side and can collect the MR's voice well. The dialogue information is stored in the storage unit.
[0120] First dialogue analysis process (S106)
[0121] For example, the first dialogue information after analysis is the dialogue statement "Next, an explanation will be given about XYZ, a new drug for diabetes. Can this new drug lower blood sugar levels?". Then, the analyzed dialogue statement is stored in the storage unit. Additionally, the dialogue fragments of this first dialogue information can also be analyzed. In this case, examples of the dialogue fragments are "Next, an explanation will be given about XYZ, a new drug for diabetes." and "Can this new drug lower blood sugar levels?". The dialogue fragments can also be analyzed in subsequent processes.
[0122] Second dialogue information acquisition process (S107)
[0123] Conversations are also input into and stored in the smartphone. Then, the smartphone also analyzes the conversations through the launched application. An example of the second conversation information is "Next, explain the new drug XYZ for diabetes. Can this new drug lower blood sugar levels?" There are differences between the laptop and the smartphone in terms of the set location and the direction of collecting voices, etc. Therefore, even when analyzing the same conversation, there are differences in the conversations analyzed between the laptop (the first voice analysis terminal) and the smartphone (the second voice analysis terminal). This process is usually carried out simultaneously with the first conversation information acquisition process (S105).
[0124] The second conversation analysis process (S108)
[0125] The second conversation information is also analyzed on the smartphone side. An example of the second conversation information is "Next, explain the new drug XYZ for diabetes. Can this new drug lower blood sugar levels?" At this time, it is also possible to analyze the conversation fragments. The second conversations obtained by analyzing the conversation fragments are "Next, explain the new drug XYZ for diabetes." and "Can this new drug lower blood sugar levels?". The second conversation information is also appropriately stored in the storage unit. In addition, the analysis of the second conversation information can be carried out by the laptop (the first voice analysis terminal) or by the server.
[0126] The second conversation information sending process (S109)
[0127] The second conversation information is, for example, sent from the smartphone to the laptop. Then, the laptop (the first voice analysis terminal 3) receives the second conversation information sent from the smartphone (the second voice analysis terminal 5).
[0128] The conversation fragment acquisition process (S110)
[0129] It is also possible to analyze the conversation fragments in the first conversation information and the second conversation information to obtain one or more conversation fragments. It is also possible to analyze the conversation fragments at each terminal. On the other hand, it is preferable to centrally analyze the conversation fragments in the conversation information recorded by the two terminals using the laptop (the first voice analysis terminal) to obtain the corresponding conversation fragments between the first conversation information and the second conversation information. In this case, the conversation times of the respective conversation fragments of the first conversation information and the respective conversation fragments of the second conversation information should be approximately the same. Therefore, it is preferable to use a timing mechanism to correspond to each fragment. In this way, it is possible to divide the first conversation information into fragments and also obtain the respective conversation fragments of the corresponding second conversation fragments.
[0130] The first dialogue segment acquisition unit 33 may also analyze the dialogue segments in the second dialogue information to obtain one or more dialogue segments.
[0131] The first dialogue information is analyzed as
[0132] "Next, XYZ, a new drug for diabetes, will be described."
[0133] The dialogue statement "Can this new drug lower blood sugar levels?"
[0134] The second dialogue information is analyzed as
[0135] "Next, XYZ, a new drug for soup birds and combination, will be described."
[0136] The dialogue statement "Can this new drug lower blood sugar levels?"
[0137] Dialogue segment selection process (S111)
[0138] Extract the terms "diabetes", "new drug", and "XYZ" from the first dialogue. Among them, the terms in the first dialogue are the terms in the dialogue included in "Next, XYZ, a new drug for diabetes, will be described." which is the first dialogue segment.
[0139] Then, use the terms "diabetes", "new drug", and "XYZ" in the first dialogue to read out the topic terms stored in association with these terms from the topic term storage unit 19. Additionally, here, extract "XYZ", which is the common topic term among the three dialogue terms, as the first topic term.
[0140] Extract the terms "soup birds", "combination", "related to fish", "new drug for combination", and "XYZ" from the second dialogue. Among them, the terms in the second dialogue are the terms in the dialogue included in "Next, XYZ, a new drug for soup birds and combination, will be described." which is the second dialogue segment. Use the terms "soup birds", "combination", "related to fish", "new drug for combination", and "XYZ" in the second dialogue to read out the topic terms stored in association with them from the topic term storage unit 19. Additionally, here, extract "Bible", which is the topic term with the most common associations among the five dialogue terms, as the second topic term.
[0141] Comparing "XYZ" which is the first topic word with "diabetes", "new drug", "XYZ", "ABC" (the names of other therapeutic drugs), "blood glucose level", "side effects", "blood sugar", "glaucoma", "retinopathy", "insulin", "DC Pharmaceutical", and "appendix" which are specific related words, it can be seen that the first topic word is a specific related word. On the other hand, comparing "Bible" which is the second topic word with specific related words, it can be seen that the second topic word is not a specific related word. Using this result, "Next, an explanation will be given of XYZ which is a new drug for diabetes." which is the first dialogue segment is adopted as the correct dialogue segment. Similarly, for the second dialogue segment immediately following, it is also judged which one is correct. The consecutive dialogue segments thus adopted are stored in the storage unit.
[0142] The consecutive dialogue segments are, "Next, an explanation will be given of XYZ which is a new drug for diabetes." "Can this new drug lower the blood glucose level?"
[0143] The above is an example of the processing. Different processing from the above processing can also be performed to adopt the correct dialogue segment.
[0144] Figure 4 It is a conceptual diagram showing a second processing example of the speech analysis system. In this example, the dialogue segments are analyzed in the second speech analysis terminal, and the second dialogue information obtained from the analysis of the dialogue segments is sent to the first speech analysis terminal. In this example, in order to avoid inconsistency of the dialogue segments, it is preferable to store time information in association with each dialogue segment, and send each dialogue segment together with the time information from the second speech analysis terminal to the first speech analysis terminal. Then, in the first speech analysis terminal, it is possible to make the dialogue segments included in the first dialogue information and the dialogue segments included in the second dialogue information consistent.
[0145] Figure 5 It is a conceptual diagram showing a processing example different from the above processing example of the speech analysis system. In this example, the second speech analysis terminal collects speech, sends the digitized dialogue information to the first speech analysis terminal, and the first speech analysis terminal performs various analyses. In addition, it is also possible to analyze the correct dialogue segments not only by the first speech analysis terminal but also by the second speech analysis terminal, but this is not particularly illustrated.
[0146] [Industrial Applicability]
[0147] This system can be used as a voice analysis device. In particular, it is considered that voice analysis devices such as Google speaker (registered trademark) will become more popular in the future. In addition, it is envisioned that voice analysis devices will also be installed in terminals near users such as smartphones and mobile terminals. For example, it is also envisioned that in a voice analysis device, it is difficult to record the user's voice because noise other than the user's voice is recorded. On the other hand, even in such a case, it is possible that the terminal near the user has appropriately recorded the user's voice. Thus, by recording voice information with the terminal near the user and sharing the voice information with the voice analysis device, voice can be analyzed with high accuracy.
[0148] Description of reference numerals
[0149] 1: Voice analysis system; 3: First voice analysis terminal; 5: Second voice analysis terminal; 7: First language analysis unit; 9: First dialogue storage unit; 11: First analysis unit; 13: Demonstration storage unit; 15: Related word storage unit; 17: Display unit; 19: Topic word storage unit; 21: Second language analysis unit; 23: Second dialogue storage unit; 25: Dialogue information receiving unit; 31: Specific demonstration information acquisition unit; 33: First dialogue segment acquisition unit; 35: Specific related word reading unit; 37: First in-dialogue term extraction unit; 39: First topic word extraction unit; 41: Second in-dialogue term extraction unit; 43: Second topic word extraction unit; 45: Dialogue segment adoption unit.
Claims
1. A speech analysis method, which uses a speech analysis system including a first speech analysis terminal (3) and a second speech analysis terminal (5). The first speech analysis terminal (3) performs a first term analysis process, which refers to analyzing the words included in the dialogue to obtain first dialogue information. The second speech analysis terminal (5) performs a second term analysis process, which refers to analyzing the words included in the dialogue to obtain second dialogue information. The speech analysis system performs a specific presentation information acquisition process, a first dialogue segment acquisition process, a specific related word reading process, a term extraction process in the first dialogue, a first topic word extraction process, a term extraction process in the second dialogue, a second topic word extraction process, and a dialogue segment adoption process. In the specific presentation information acquisition process, information about the selection of a specific presentation material is received, and the specific presentation material is one of multiple presentation materials. In the first dialogue segment acquisition process, the dialogue segments in the first dialogue information are analyzed to obtain one or more dialogue segments. In the specific related word reading process, a specific related word is read, and the specific related word is a related word associated with the specific presentation material. In the term extraction process in the first dialogue, terms in the first dialogue are extracted, and the terms in the first dialogue are terms in the dialogue included in the first dialogue segment, and the first dialogue segment is one of the dialogue segments in the first dialogue information. In the first topic word extraction process, a first topic word is extracted, and the first topic word is a topic word associated with the terms in the first dialogue. In the term extraction process in the second dialogue, terms in the second dialogue are extracted, and the terms in the second dialogue are terms in the dialogue included in the second dialogue segment, and the second dialogue segment is the dialogue segment corresponding to the first dialogue segment in the second dialogue information. In the second topic word extraction process, a second topic word is extracted, and the second topic word is a topic word associated with the terms in the second dialogue. In the dialogue segment adoption process, the relationship between the first topic word and the specific related word and the relationship between the second topic word and the specific related word are used to adopt the first dialogue segment or the second dialogue segment as the correct dialogue segment.
2. The voice analysis method according to claim 1, characterized in that , The dialogue segment adoption process is as follows: when the first topic word and the second topic word are different, when the first topic word is the specific related word and the second topic word is not the specific related word, the first dialogue segment in the first dialogue information is adopted as the correct dialogue segment; when the first topic word is not the specific related word and the second topic word is the specific related word, the second dialogue segment in the second dialogue information is adopted as the correct dialogue segment.
3. The voice analysis method according to claim 1, characterized in that , The dialogue segment adoption process is as follows: compare the number of cases where the first topic word is the specific related word and the number of cases where the second topic word is the specific related word. When the former is more, the first dialogue segment is adopted as the correct dialogue segment; when the latter is more, the second dialogue segment is adopted as the correct dialogue segment.
4. The voice analysis method according to claim 1, wherein , The first voice analysis terminal (3) has a time storage unit (51), and the time storage unit (51) is used to store the time. The first conversation information includes the words included in the conversation and the times associated with each word. The first conversation segment acquisition process uses the time information of each word to analyze the conversation segment.
5. The voice analysis method according to claim 1, characterized in that , The first voice analysis terminal (3) has a frequency analysis unit (53), and the frequency analysis unit (53) analyzes the frequency of the voice included in the conversation. The first conversation information includes the words included in the conversation and the frequencies of the voices associated with each word. The first conversation segment acquisition process uses the frequencies of each word to analyze the conversation segment.
Citation Information
Patent Citations
Recognition sharing supporting method, discussion structuring supporting method, condition grasping supporting method, graphic thinking power development supporting method, collaboration environment construction supporting method, question and answer supporting method and proceedings preparation supporting method
JP2002259635A
Presentation evaluation device, presentation evaluation system, presentation evaluation program and control method of presentation evaluation device
JP2017224052A
Voice recognition system and voice processing system
CN1920948A
Conversation screening program, conversation screening device, and conversation screening method
JP2014102513A