Voice analysis system, voice analysis program, and voice analysis method
The voice analysis system addresses the limitation of analyzing combined conversation content by identifying and grouping conversational elements, enhancing understanding and improving sales skills through detailed analysis and display of conversational expressions.
Patent Information
- Application Number
- JP2025061488
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-12-01
- Estimated Expiration
- 2045-04-02
AI Technical Summary
Existing voice analysis systems only determine the positivity or negativity of individual keywords in conversations, failing to analyze the combined content effectively.
A voice analysis system that identifies multiple conversation elements and groups them into sections based on conversational expressions, using generative artificial intelligence to analyze and evaluate staff conversations at events, providing insights into conversational flow and skills.
Enables accurate understanding of conversation content and improves sales skills by analyzing and displaying conversational elements and expressions, facilitating learning from high-performing staff members.
Smart Images

Figure 0007777840000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a voice analysis system, a voice analysis program, and a voice analysis method. [Background technology]
[0002] Conventionally, in order to improve a user's communication ability or communication technique, a voice analysis is performed on the user's speech and the analysis result is fed back to the user. An example of such a technique is proposed in, for example, Patent Document 1.
[0003] For example, Patent Document 1 describes a method of analyzing user voice data to extract keywords from the user voice data, and then referring to a keyword classification database in which words used in conversation are classified into positive and negative phrases and stored as positive and negative words, determining whether the extracted keywords are positive or negative words, and the results of the determination are compiled into an evaluation sheet and fed back to the user. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2023-533 Summary of the Invention [Problem to be solved by the invention]
[0005] In order to understand the meaning of a conversation between people, it is necessary to use multiple pieces of content in a composite manner. However, Patent Document 1 was only able to determine the positive or negative of individual keywords contained in the conversation, and was unable to perform conversation analysis that uses the combined content contained in the conversation.
[0006] In view of the above problems, an object of the present invention is to provide a technology that can accurately grasp the content of a user's conversation. [Means for solving the problem]
[0007] In order to solve the above problems, the present invention provides a voice analysis system that analyzes the content of conversations between staff at an event, the voice analysis system comprising a conversation identification unit and a development identification unit, wherein the conversation identification unit identifies multiple conversation elements into which the content of the conversations between staff and attendees is divided based on conversational voice data including the content of conversations between staff and attendees, and the development identification unit identifies conversation sections in a series of conversations into which the multiple conversation elements are grouped based on conversational expressions contained in the multiple conversation elements.
[0008] In addition, in order to solve the above-mentioned problems, the present invention provides a voice analysis program that analyzes the content of conversations between staff at an event, wherein the voice analysis program causes a computer to function as a conversation identification unit and a development identification unit, wherein the conversation identification unit identifies multiple conversation elements into which the content of the staff's conversation is divided based on conversational voice data including the content of the conversation between staff and attendees, and the development identification unit identifies conversation sections in a series of conversation into which the multiple conversation elements are grouped based on conversational expressions included in the multiple conversation elements.
[0009] In addition, in order to solve the above-mentioned problems, the present invention is a voice analysis method for analyzing the content of conversations between staff at an event, in which a computer performs the following processes: identifying multiple conversation elements into which the content of conversations between staff and visitors is divided, based on conversational voice data including the content of conversations between staff and visitors; and identifying conversation sections in a series of conversations into which the multiple conversation elements are grouped, based on conversational expressions contained in the multiple conversation elements.
[0010] This configuration makes it possible to analyze staff conversations at an event. Specifically, by grouping multiple conversation elements contained in staff conversations, it is possible to easily understand the flow of staff conversations.
[0011] In a preferred embodiment of the present invention, the speech analysis system further includes an expression identification unit, wherein the series of conversational contents of the staff member includes a first conversation section and a second conversation section, and the expression identification unit inputs the conversational expressions included in the first conversation section to the second conversation section into a generation artificial intelligence and identifies, from the conversational expressions included in the first conversation section to the second conversation section, linking expressions that are conversational expressions that connect the first conversation section and the second conversation section.
[0012] In a preferred embodiment of the present invention, the database has an expression identification unit that inputs conversational expressions included in the first conversation section to the second conversation section into a generation artificial intelligence and identifies the linking expression for each combination of the first conversation section and the second conversation section.
[0013] This configuration makes it possible to identify conversational expressions (linking expressions) that are used to smoothly connect conversation sections. This allows other staff members to improve their sales skills by learning from the linking expressions of staff members with high sales performance.
[0014] In a preferred embodiment of the present invention, the speech analysis system further includes a display processing unit, which displays and processes a graph connecting the first conversation section and the second conversation section as nodes and the linking expressions as edges based on the conversation order of the first conversation section, the second conversation section, the linking expressions, and the plurality of conversation elements.
[0015] This structure makes it easy to understand the connections between multiple conversation sections and linking expressions, making it easier to analyze the conversations of staff with high sales skills and helping to improve the sales skills of other staff.
[0016] In a preferred embodiment of the present invention, the speech analysis system further includes a linking unit that links and registers a plurality of conversational expressions based on a combination of the plurality of conversational expressions included in each conversation section.
[0017] By using this configuration, it is possible to easily grasp conversational expressions that are used together in a conversation.
[0018] In a preferred embodiment of the present invention, the speech analysis system further includes an evaluation unit that inputs the plurality of conversational elements into a generation artificial intelligence and processes to generate a staff evaluation based on the plurality of conversational expressions contained in the plurality of conversational elements.
[0019] With this configuration, it is possible to evaluate the conversational expressions used by staff members in their conversations.
[0020] In a preferred embodiment of the present invention, the speech analysis system further includes an evaluation unit that inputs the plurality of conversation elements into a generation artificial intelligence to generate and process a staff evaluation based on the conversation section and the linking expression.
[0021] With this configuration, it is possible to evaluate the flow of the staff's conversation.
[0022] In a preferred embodiment of the present invention, the conversation identifying unit identifies the plurality of conversation elements based on the most frequently included piece of voice data among the plurality of voice data included in the conversation voice data.
[0023] This configuration makes it possible to identify the conversational elements of staff members to be analyzed without having to register their voices in advance, which means that even if there is a sudden change in staff members, the content of their conversations can be analyzed without any problems. [Effects of the Invention]
[0024] The present invention has an effect of providing a technique that makes it possible to accurately grasp the content of a user's conversation. [Brief explanation of the drawings]
[0025] [Figure 1] 1 is a block diagram showing a configuration of a system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a block diagram of a hardware configuration of a system according to the present invention. [Figure 3] FIG. 2 is a block diagram of a functional configuration according to an embodiment of the present invention. [Figure 4] 1 is an example of a processing flowchart according to an embodiment of the present invention. [Figure 5] FIG. 10 is a diagram showing an example of a key phrase analysis screen according to an embodiment of the present invention. [Figure 6] FIG. 10 is a diagram showing an example of a deployment analysis screen according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0026] The present invention will be described in more detail below with reference to the accompanying drawings, in which preferred embodiments are shown, but which may be embodied in many different forms and are not limited to the embodiments set forth herein.
[0027] For example, although the configuration, operation, etc. of a voice analysis system are described in this embodiment, similar effects can be achieved by an executed method, device, computer program, etc. Furthermore, the program may be stored on a recording medium. By using this recording medium, the program can be installed on, for example, a computer, thereby configuring a voice analysis device or a voice analysis system. Here, the recording medium storing the program may be a non-transitory recording medium such as a CD-ROM.
[0028] <1. Overview of the present invention> The present invention relates to a system for analyzing conversations that take place at an event. The present invention analyzes staff conversations by recording conversations between event staff (hereinafter simply referred to as staff) and event attendees (hereinafter simply referred to as attendees) at an event venue and acquiring conversational audio data. In this embodiment, conversational audio data containing the voices of the staff and attendees is first input into a generation artificial intelligence (AI), which extracts the content of the staff conversation (hereinafter referred to as staff conversation content) from the conversational audio data. Then, multiple conversation elements obtained by dividing the staff conversation content are extracted, and conversation sections, each grouping the multiple conversation elements, are identified based on the conversational expressions contained in the multiple conversation elements. The staff conversation content is then analyzed based on the conversation order and conversational expressions of the multiple conversational sections contained in the staff conversation content, thereby evaluating the staff's conversational skills.
[0029] The conversational voice data in this embodiment is a series of recorded voice data of one or more visitors talking to one staff member. The staff conversation content is voice data in which only the staff member's conversation is extracted from the series of voice data, or text data based on the voice data. The conversational voice data may include voice data of two or more staff members.
[0030] A conversation element is the smallest unit of text data of an utterance included in the staff conversation content (the content uttered by a speaker in one breath), but it may also be a section in a conversation where one speaker speaks, or a unit of text data separated by a breath or change in intonation. A conversation expression is a word or phrase included in the staff conversation content.
[0031] A conversation section is a series of text data consisting of multiple conversation elements from a conversation element containing an expansion key phrase, which is a conversational expression for dividing a conversation section, to a conversation element immediately before the conversation element in which another expansion key phrase appears.
[0032] Furthermore, the generative artificial intelligence is a model that learns on its own based on a huge amount of data, and generates output data based on the input of text data or voice data related to the data to be output. For example, text is generated as output data. Here, the output text may include program code such as a script or text data written in a specific format such as CSV. In this embodiment, the generative artificial intelligence generates output data based on the analysis results of the staff conversation content in response to input conversation voice data.
[0033] In this embodiment, the generative artificial intelligence is a large-scale language model such as a GPT (Generative Pretrained Transformer), but there are no restrictions on its scale or model structure, and it may be other language models that combine one or more models, such as a convolutional neural network (CNN); a recurrent neural network (RNN) such as an LSTM (Long short-term memory) or a GRU (Gated recurrent unit); or a Transformer model.
[0034] Furthermore, although an exhibition will be described as an event in this embodiment, the event is not limited to this and may be a short-term event or a permanent event.
[0035] <1.1. System Configuration> Fig. 1 is a block diagram showing the configuration of a system of the present invention. As shown in Fig. 1, a speech analysis system 0 includes a speech analysis device 1 and a user terminal device 3, and is configured to be able to communicate via a communication network NW. The communication network NW in the present invention is an IP (Internet Protocol) network, but there is no restriction on the type of communication protocol, and there is also no restriction on the type or scale of the network.
[0036] A general-purpose server computer, a personal computer, or the like can be used as the voice analysis device 1. A smartphone, a tablet terminal, a personal computer, a wearable device, or the like can be used as the user terminal device 3. The voice analysis device 1 may also be configured with multiple computers that are capable of sending and receiving information via a communication network NW or another network.
[0037] <1.2. Hardware configuration of the present invention> 2 is a block diagram of the hardware configuration of the voice analysis system 0. As shown in FIG. 2(a), the server 10 (voice analysis device 1) includes a processing unit 101, a storage unit 102, and a communication unit 103.
[0038] The processing unit 101 has one or more processors such as a CPU capable of executing an instruction set, and controls the overall operation and processing of the voice analysis device 1 by executing the voice analysis program according to the present invention, an OS, and other applications. The storage unit 102 has a volatile memory such as a RAM capable of storing an instruction set, and a non-volatile recording medium such as an HDD or SSD capable of recording an OS, a voice analysis program according to the present invention, and the like. The communication unit 103 has a communication interface device with the communication network NW, and controls communication with the communication network NW to input and output information.
[0039] As shown in FIG. 2(b), the terminal device 9 (user terminal device 3) includes a processing unit 91, a storage unit 92, a communication unit 93, an input unit 94, and an output unit 95.
[0040] The processing unit 91 has one or more processors such as a CPU that can execute an instruction set, and controls the overall operation and processing of the terminal device 9 by executing an OS and other applications. The storage unit 92 has a volatile memory such as RAM that can store an instruction set, and a non-volatile recording medium such as an HDD or SSD that can record an OS, voice data of conversations between visitors and staff, and the like. The communication unit 93 has a communication interface device for connecting to a network, and controls communication with the communication network NW to input and output information. The input unit 94 has input devices capable of input processing, such as a keyboard, a touch panel, and a microphone for acquiring conversation voice data. The output unit 95 has a display device capable of display processing, such as a display.
[0041] <1.3. System Functional Configuration> Fig. 3 is a block diagram of the functional configuration of the speech analysis device 1. As shown in Fig. 3, the speech analysis device 1 includes an acquisition unit 11, a conversation identification unit 12, a development identification unit 13, an expression identification unit 14, a negotiation identification unit 15, a linking unit 16, an evaluation unit 17, a display processing unit 18, and a database 2. This is a specific implementation of information processing by software (stored in a storage unit 102) using hardware (processing unit 101).
[0042] The system configuration in this embodiment is a so-called server-client type, in which the user terminal device 3 (client) receives the processing results performed by the speech analysis device 1 (server) in response to a request from the client. Alternatively, it may be a so-called standalone type, in which a speech analysis program is started on the client terminal. In this case, the user terminal device 3 may include some or all of the functional components (units) included in the speech analysis device 1. For example, the user terminal device 3 may include an acquisition unit 11, a conversation identification unit 12, a development identification unit 13, an expression identification unit 14, a business negotiation identification unit 15, a linking unit 16, an evaluation unit 17, a display processing unit 18, etc., and the speech analysis device 1 may be a cloud storage that stores the analysis results, etc.
[0043] <1.3.1. Acquisition part 11> The acquisition unit 11 acquires conversational voice data. The acquisition unit 11 acquires conversational voice data including the voice of the staff member to be analyzed from the user terminal device 3 that recorded the conversational voice data.
[0044] <1.3.2. Conversation Identification Unit 12> The conversation identification unit 12 identifies the content of the staff conversation. The conversation identification unit 12 identifies the content of the staff conversation based on conversational voice data including the content of the conversation between the staff and the visitors.
[0045] The conversation identifying unit 12 also identifies staff conversation elements. Based on the identified staff conversation content, the conversation identifying unit 12 identifies a plurality of conversation elements into which the staff conversation content is divided.
[0046] <1.3.3. Deployment specific part 13> The development identification unit 13 identifies a conversation section based on conversational expressions included in a plurality of conversational elements. The development identification unit 13 identifies a conversation section in a series of conversations in which the plurality of conversational elements are grouped.
[0047] The development identification unit 13 also identifies a development type. The development identification unit 13 identifies a development type, which is a type of conversation section, based on conversational expressions included in the conversation section. Here, examples of development types include introduction, main point, example, conclusion, conjunction, and others.
[0048] <1.3.4. Expression identification part 14> The expression identification unit 14 identifies conversational expressions included in the staff conversation elements. The expression identification unit 14 identifies conversational key phrases, which are important phrases in business negotiations, from the conversational expressions included in the staff conversation elements. The determination of whether a phrase is important will be described later.
[0049] The expression specification unit 14 also specifies the conversation type of the conversation expression based on the conversation expression included in the staff conversation element. The expression specification unit 14 specifies the conversation type of the conversation key phrase based on the specified conversation key phrase.
[0050] In addition to or instead of conversational key phrases, conversational keywords, which are important words in business negotiations, may be used as conversational expressions.
[0051] Furthermore, the expression identification unit 14 identifies a linking expression that connects two conversation sections from among the conversation expressions included in the staff conversation elements. For the first conversation section and the second conversation section identified by the development identification unit 13, the expression identification unit 14 identifies a linking expression that connects the first conversation section and the second conversation section from among the conversation expressions included from the first conversation section to the second conversation section.
[0052] <1.3.5.Business Negotiation Specification Section 15> The business negotiation identifying unit 15 identifies one business negotiation between a staff member and a visitor. The business negotiation identifying unit 15 identifies one business negotiation based on the development type of the conversation section.
[0053] <1.3.6. Linking section 16> The linking unit 16 links and registers a plurality of conversational expressions included in the staff conversation content. The linking unit 16 links and registers a plurality of conversational key phrases based on a combination of a plurality of conversational key phrases included in each conversation section.
[0054] <1.3.7. Evaluation section 17> The evaluation unit 17 generates a staff evaluation based on a plurality of conversational expressions contained in a plurality of conversational elements. The evaluation unit 17 inputs the plurality of conversational elements contained in the staff conversation content to the generation artificial intelligence, and generates a staff evaluation based on a plurality of conversational key phrases as a staff evaluation based on a plurality of conversational expressions contained in the plurality of conversational elements.
[0055] In addition, the evaluation unit 17 inputs multiple conversation elements contained in the staff conversation content into the generation artificial intelligence, and generates and processes a staff evaluation based on multiple conversation sections and linking expressions as a staff evaluation based on the multiple conversation expressions contained in the multiple conversation elements.
[0056] <1.3.8. Display Processing Unit 18> The display processing unit 18 performs display processing on the analysis results of the conversation content data, and causes the display processing results to be displayed on the user terminal device 3. An example of displaying the analysis results will be described later.
[0057] <1.4. Processing Flowchart> A voice analysis method using the voice analysis system 0 will be described with reference to Fig. 4. Fig. 4 is a flowchart showing the process in which the voice analysis device 1 acquires conversational voice data including the voices of a visitor and one staff member to be analyzed, identifies the content of the staff member conversation, performs various analyses based on the content of the staff member conversation, and displays the analysis results.
[0058] <1.4.1. Acquisition of conversation voice data> First, in step S1 (hereinafter, "step SX" will be abbreviated simply as "SX"), the acquisition unit 11 acquires conversational voice data. In this embodiment, the user terminal device 3 accepts input of an event name from a staff member, acquires conversational voice data including the content of the conversation between the staff member and attendees at the event venue, and stores the conversational voice data and the recording date in the storage unit 92. Then, the acquisition unit 11 acquires the conversational voice data, the event name, and the recording date.
[0059] In this embodiment, the acquisition unit 11 acquires conversational voice data via the communication network NW, but may also acquire the data via a recording medium on which the conversational voice data is recorded.
[0060] <1.4.2. Identifying staff conversation content> In S2, the conversation identification unit 12 identifies the staff conversation content. In this embodiment, the conversation identification unit 12 inputs the conversation voice data acquired in S1 and an instruction to identify the voice data that is most frequently included in the conversation voice data to the generation artificial intelligence, and based on the voice data that is most frequently included among the multiple voice data included in the conversation voice data, identifies the voice data that is most frequently included or text data based on the voice data as the staff conversation content.
[0061] In this embodiment, the conversation identification unit 12 identifies the staff conversation content of one staff member based on the voice data that is most frequently included in the conversation voice data. On the other hand, if there are two staff members to be analyzed, the conversation identification unit 12 may input the conversation voice data, instructions to identify the voice data that is most frequently included in the conversation voice data, and the number of staff members to be analyzed to the generation artificial intelligence, and identify the staff conversation content for the number of staff members based on the voice data up to the input number of staff members, based on the voice data that is most frequently included among the voice data that is most frequently included in the conversation voice data.
[0062] Furthermore, the conversation identification unit 12 may identify the content of the staff conversation based on the pitch, formants, and tempo of the voice included in the conversation voice data.
[0063] <1.4.3. Identifying Conversational Elements> In S3, the conversation identification unit 12 identifies conversation elements. In this embodiment, the conversation identification unit 12 inputs the staff conversation content identified in S2 into the generating artificial intelligence, thereby identifying multiple conversation elements that make up the staff conversation content. Specifically, by inputting the staff conversation content into the generating artificial intelligence, the conversation identification unit 12 performs a morphological analysis of the staff conversation content and identifies multiple conversation elements based on locations where the audio is interrupted, based on audio continuity such as whether or not there is breathing sound from the staff, whether or not fillers are inserted, etc.
[0064] <1.4.4. Identifying the conversation section> In S4, the development identification unit 13 identifies a conversation section. In this embodiment, the development identification unit 13 inputs a plurality of conversation elements to the generation artificial intelligence, and identifies the conversation section based on development key phrases for dividing the conversation section from among the conversation expressions included in the plurality of conversation elements.
[0065] Specifically, the expansion identification unit 13 identifies, as a conversation section, a group of conversation elements from a conversation element including an expansion keyphrase to a conversation element immediately preceding the conversation element including a different expansion keyphrase, using a predetermined phrase for dividing the conversation section as an expansion keyphrase. Based on the expansion keyphrase included in one conversation section, the expansion identification unit 13 identifies the expansion type of the conversation section including the expansion keyphrase. Examples of predetermined phrases for dividing the conversation sections include "company overview," "company introduction," "product introduction," "product proposal," and "answer."
[0066] The development identification unit 13 may identify a conversation section by grouping together multiple conversation elements that include the same development key phrase from among a series of multiple conversation elements. The development identification unit 13 may also identify a development type based on a phrase that summarizes multiple conversation elements included in a conversation section. The development identification unit 13 may also identify a conversation section by using development keywords for dividing conversation sections in addition to or instead of development key phrases from among conversational expressions included in multiple conversation elements.
[0067] <1.4.5. Identifying linking expressions> In S5, the expression identification unit 14 identifies a linking expression. In this embodiment, the expression identification unit 14 inputs conversational expressions included in the first conversation section to the second conversation section to the generation artificial intelligence, and identifies a linking expression from among the conversational expressions included in the first conversation section to the second conversation section.
[0068] Specifically, the development identification unit 13 generates conversation section labels by inputting the conversation sections into the generation artificial intelligence, and the database 2 stores a linking expression index indicating a score for determining that a conversation expression is a linking expression for each combination of a first conversation section label and a second conversation section label. The expression identification unit 14 identifies a linking expression for each combination of the first and second conversation sections from among the conversation expressions included in the first to second conversation sections, based on the labels of the first and second conversation sections generated by inputting the conversation expressions included in the first to second conversation sections into the generation artificial intelligence, and the linking expression index. In other words, a conversation expression that is likely to be a linking expression is identified from among a plurality of conversation expressions included in the first conversation section and the second conversation section using the linking expression index.
[0069] Here, the conversational expressions contained in the first conversation section through the second conversation section include conversational expressions contained in either the first conversation section or the second conversation section, and conversational expressions contained between the first conversation section and the second conversation section in the series of conversational content of the staff. Note that the conversational expressions contained between the first conversation section and the second conversation section may be another conversation section (third conversation section) or may be conversational elements.
[0070] Furthermore, the label of a conversation section is a summary (heading or phrase) that expresses the content of the conversation section in one word, but it may also be a development type.
[0071] In addition, the database 2 may store a linking expression index that indicates the score for each conversation expression as a linking expression, and in this case, the expression identification unit 14 may identify a linking expression from among the conversation expressions included in the first conversation section to the second conversation section based on the linking expression index and the conversation expressions included in the first conversation section to the second conversation section.
[0072] <1.4.6. Identifying business opportunities> In S6, the business negotiation identification unit 15 identifies one business negotiation from the staff conversation content. In this embodiment, based on the development type identified in S4, the business negotiation identification unit 15 identifies a conversation section whose development type is the start of the business negotiation (e.g., introduction) as a starting conversation section, and further identifies a conversation section whose development type is the end of the business negotiation (e.g., conclusion) as an ending conversation section. Then, the business negotiation identification unit 15 identifies multiple conversation sections from the starting conversation section to the ending conversation section in the series of staff conversation content as one group, and identifies this group as a business negotiation.
[0073] <1.4.7. Identifying conversational key phrases> In S7, the expression identification unit 14 identifies conversational key phrases. In this embodiment, the expression identification unit 14 inputs a plurality of conversational expressions included in the conversation elements identified in S3 to the generation artificial intelligence, and identifies conversational key phrases from the plurality of conversational expressions based on the appearance frequency of the conversational expressions.
[0074] Specifically, the expression identification unit 14 identifies conversational key phrases from the multiple conversational expressions included in the staff conversation content identified in S2 based on the appearance frequency of each conversational expression included in the staff conversation content identified in S2 and the appearance frequency of each conversational expression included in the staff conversation content identified in S2 in the multiple staff conversation contents to be analyzed.More specifically, the expression identification unit 14 substitutes the proportion of the appearance frequency of each conversational expression among all conversational expressions included in the staff conversation content identified in S2 and the proportion of the number of staff conversation contents in which each conversational expression included in the staff conversation content identified in S2 appears in the total number of the multiple staff conversation contents to be analyzed into a multiplicative calculation formula, and identifies conversational key phrases depending on the magnitude of the value of each conversational expression obtained as a result of the calculation (so-called TF-IDF value).
[0075] Furthermore, the expression identification unit 14 identifies conversational key phrases by inputting the context of the staff conversation content to the generating artificial intelligence in addition to or instead of the appearance frequency of each conversational expression. Specifically, the expression identification unit 14 identifies conversational key phrases by inputting combinations of visitor conversation elements (e.g., questions) and staff conversation elements (e.g., replies to the questions) as the context of the staff conversation content to the generating artificial intelligence. This makes it possible to determine whether the staff's response to the visitor's conversation was appropriate.
[0076] Furthermore, the expression identification unit 14 identifies the conversation type of the identified conversation key phrase. Specifically, the expression identification unit 14 inputs the identified conversation key phrase and multiple conversation elements included in the business negotiation identified in S6 into the generation artificial intelligence, thereby identifying the conversation type of the conversation key phrase in the business negotiation. In other words, even for the same conversation key phrase, different conversation types are identified depending on the content of the business negotiation (for example, the conversation type is identified from the entire conversation for homonymous words).
[0077] <1.4.8. Linking conversational key phrases> In S8, the linking unit 16 links and registers the conversational key phrases. In this embodiment, the linking unit 16 links and registers the multiple conversational key phrases based on the frequency of use of the combination of multiple conversational key phrases included in each conversation section. Specifically, the linking unit 16 identifies the frequency of use of the combination of the first conversational key phrase and the second conversational key phrase based on the ratio of the frequency at which the second conversational key phrase appears after the first conversational key phrase in the frequency of appearance of the first conversational key phrase (so-called conditional probability), and links and registers the multiple conversational key phrases according to the frequency of use.
[0078] <1.4.9. Creation of staff evaluations> In S9, the evaluation unit 17 generates a staff evaluation based on the analysis results. In this embodiment, the evaluation unit 17 inputs the conversational key phrases and conversation types of the conversational key phrases identified in S7, the combinations of conversational key phrases registered in S8, and model staff characteristics into the generation artificial intelligence, thereby generating a key phrase evaluation for one business negotiation as a staff evaluation. Here, the model staff characteristics may be, for example, knowledge characteristics related to ideal knowledge, response characteristics related to ideal response ability, and phrase characteristics related to ideally used key phrases. In addition, the key phrase evaluation may include, for example, an evaluation of the conversational key phrases and conversation types used by the staff, an evaluation of the number of such conversation types, an evaluation of the number of combinations of conversational key phrases (e.g., an evaluation of the patterns of conversational key phrases), an evaluation of differences between the model staff characteristics and the identified conversational key phrases, and improvement proposals based on the analysis results.
[0079] Furthermore, in S9, the evaluation unit 17 generates a development evaluation of the business negotiation in 1 as a staff evaluation by inputting the conversation section identified in S4, the development type of the conversation section, and the linking expression registered in S5 into the generating artificial intelligence. Here, the development evaluation may include, for example, an evaluation of the linking expression (e.g., a technique used to attract the audience), an evaluation of the combination of the development key phrase and the linking expression (e.g., an evaluation of the clarity and consistency of the topic transition), an evaluation of the development type (e.g., an evaluation of whether examples are used), and improvement suggestions based on the analysis results.
[0080] Furthermore, the evaluation unit 17 generates an expansion evaluation based on a plurality of conversation sections linked via linking expressions and each expanded key phrase in the conversation sections. Specifically, the evaluation unit 17 inputs the plurality of conversation sections linked via linking expressions and each expanded key phrase in the conversation sections to the generation artificial intelligence, and calculates an expanded phrase relevance indicating the degree of relevance of each expanded key phrase as the expansion evaluation. This makes it possible to grasp, for example, whether highly relevant keywords are used in a conversation section whose development type is "example" compared to a conversation section whose development type is "main point." In other words, it is easy to grasp whether the main points are clearly reinforced by the examples.
[0081] In a preferred embodiment of the present invention, the evaluation unit 17 generates staff evaluations using the conversation speed based on the attendee's voice data included in the conversational voice data and the conversation speed based on the staff's voice data included in the conversational voice data. Specifically, the evaluation unit 17 inputs the conversational voice data into a generating artificial intelligence to identify the attendee's conversation speed and the staff's conversation speed for each business negotiation, and generates staff evaluations that are higher the closer these conversation speeds are to each other.
[0082] <1.4.10. Displaying analysis results> In S10, the display processing unit 18 displays the analysis results. In this embodiment, the display processing unit 18 sorts the data by the event name and recording date acquired in S1, and receives a designation of the analyzed conversation voice data, thereby displaying the analysis results of the conversation voice data acquired in S3 to S9.
[0083] Figure 5 shows an example of a key phrase analysis screen display for conversational key phrases in multiple identified business negotiations. The key phrase analysis screen W1 has a key phrase display section W11 and a key phrase evaluation display section W12. The key phrase display section W11 displays a graph in which conversational key phrases are used as nodes and connected for each conversation section, based on the conversational key phrases that are linked to each other and registered for each conversation section. The key phrase evaluation display section W12 displays the key phrase evaluations generated in S9.
[0084] The size of the nodes in the key phrase display section W11 is adjusted based on the frequency of appearance of each conversational key phrase. The appearance (color, shape, etc.) of the nodes is adjusted based on the conversation type (in the illustrated example, business field, product, client, etc.) of each conversational key phrase. The thickness of the edges is adjusted based on the relevance of the nodes.
[0085] Here, the relevance of a node is a numerical value indicating the degree of relevance between multiple linked conversational key phrases, and is calculated by the linking unit 16. Specifically, the linking unit 16 inputs the staff details in one identified business negotiation and multiple conversational key phrases linked for each conversation section into the generating artificial intelligence, and calculates the conversational phrase relevance between the conversational key phrases for each conversation section as the relevance of the node. This allows the relevance of the same two conversational key phrases to be calculated differently depending on the conversation section (the context of the staff conversation), making it possible to grasp the relevance of more appropriate conversational key phrases in a conversation.
[0086] 6 shows an example of a display of a development analysis screen for the conversation sections in the identified multiple business negotiations. The development analysis screen W2 has a conversation flow display section W21 and a development evaluation display section W22. Based on the first conversation section, the second conversation section, linking expressions, and the conversation order of the multiple conversation elements, the display processing unit 18 displays a graph connecting the first conversation section and the second conversation section as nodes and the linking expressions as edges.
[0087] In the illustrated example, the conversation flow display unit W21 generates titles based on multiple conversation elements contained in each conversation section by the development determination unit 13, and assigns these titles to nodes corresponding to each conversation section. Furthermore, linking expression phrases are assigned to edges corresponding to linking expressions connecting the conversation sections. The conversation section "Introduction to the Exhibition" is displayed as the development type "Introduction (Starting Point of the Business Negotiation)," followed by the conversation section "Design Company" as the development type "Main Point." These conversation sections are displayed connected by an edge leading from "Introduction to the Exhibition" to "Design Company." Furthermore, the development evaluation display unit W22 displays the development evaluation generated in S9. Although not shown in the figure, a graph composed of multiple nodes and edges may be displayed for multiple business negotiations. If a node with a common title exists among the nodes constituting the graphs for each business negotiation, the node may be used as a common node to display the graph.
[0088] 5 and 6, the analysis results for one staff member are displayed, but the analysis results for two or more staff members may be displayed side by side on the same screen. In this embodiment, the analysis results for multiple negotiations conducted by one staff member are displayed on the same screen. On the other hand, the analysis results for multiple negotiations conducted by one staff member may be displayed in a switchable manner.
[0089] Furthermore, the analysis results of the business negotiation with the highest negotiation points based on the development evaluation among multiple business negotiations conducted by one staff member may be displayed. Specifically, the evaluation unit 17 calculates the negotiation points for each business negotiation conducted by one staff member by inputting the combination of multiple conversation sections and linking expressions in the generated development evaluation into the generation artificial intelligence. For example, the evaluation unit 17 calculates the negotiation points so that the negotiation points are highest for a business negotiation in which multiple conversation sections from the starting conversation section to the ending conversation section are connected in series without branching.
[0090] As described above, by executing the processes S1 to S10 in embodiment 1, the conversations of each staff member can be analyzed in detail, and by referring to the content of the conversations of staff members with high sales skills, the sales skills of the entire staff can be improved.
[0091] In this embodiment, the conversation elements are identified after the staff conversation content is identified, but the staff conversation content may be identified after the conversation elements are identified.
[0092] Furthermore, in this embodiment, the display processing refers to a process in which the display processing unit 18 executes a process of generating information necessary for display, and transmits the generated information to the terminal device 9, thereby causing the terminal device 9 to display the generated information. On the other hand, the display processing may also be a process in which the display processing unit 18 transmits a processing command to the terminal device 9 to generate information necessary for display, thereby causing the terminal device 9 to generate information necessary for display, and display the generated information. Furthermore, in the case where the display processing unit 18 is provided in the terminal device 9 (in the case of a stand-alone type), the display processing may also be a process in which the display processing unit 18 executes a process of generating necessary information, transmits the generated information to the output unit 95 of the terminal device 9, and causes the output unit 95 to display the generated information.
[0093] In addition, the generation process in this embodiment involves sending an analysis request including at least the information to be analyzed to a generation AI external to this system and obtaining the generated analysis results. Alternatively, the system may have a generation AI, and the generation AI may analyze the information to be analyzed. [Explanation of symbols]
[0094] 0: Voice analysis system 1: Voice analysis device 2: Database 3: User terminal device 100: Server 101: Processing section 102: Storage section 103: Communications Department 9: Terminal device 91: Processing section 92: Storage section 93: Communications Department 94: Input section 95: Output section 11: Acquisition part 12: Conversation identification part 13: Deployment specific part 14:Expression specific part 15: Negotiation Specific Department 16: Attachment section 17: Evaluation section 18: Display processing section NW: Communication network W1: Key phrase analysis screen W11: Key phrase display section W12: Keyphrase evaluation display section W2: Deployment analysis screen W21: Conversation flow display section W22: Deployment evaluation display section
Claims
1. A voice analysis system for analyzing the content of conversations of staff at an event, The speech analysis system includes a conversation identification unit, a development identification unit, and an expression identification unit, the conversation identification unit identifies a plurality of conversation elements into which the content of the staff's conversation is divided based on conversational voice data including the content of the conversation between the staff and the attendee; the development identification unit identifies a conversation section in a series of conversations in which the plurality of conversation elements are grouped, based on conversational expressions included in the plurality of conversation elements; The series of conversation contents of the staff member includes a first conversation section, a second conversation section, and the conversation elements between the first conversation section and the second conversation section that do not belong to the first conversation section or the second conversation section; The expression identification unit identifies linking expressions from conversation expressions included in the conversation elements between the first conversation section and the second conversation section using the labels of the first conversation section and the labels of the second conversation section obtained by inputting instructions to a generation artificial intelligence to identify the conversation expressions included in the first conversation section to the second conversation section and the labels of the conversation sections, and linking expression indicators for identifying linking expressions, which are conversation expressions connecting the conversation sections, for each combination of the labels of the first conversation section and the labels of the second conversation section. Speech analysis system.
2. The speech analysis system further includes a display processing unit, The display processing unit processes a graph connecting the first conversation section and the second conversation section as nodes and the linking expressions as edges based on the first conversation section, the second conversation section, the linking expressions, and the conversation order of the plurality of conversation elements to display the graph. The speech analysis system of claim 1 .
3. The voice analysis system further includes a linking unit, The linking unit identifies a frequency of use of a combination of a first conversation expression and a second conversation expression for a plurality of conversation expressions included in each conversation section based on a ratio of a frequency at which a second conversation expression appears after a first conversation expression in the frequency of appearance of the first conversation expression, and links and registers the plurality of conversation expressions according to the degree of the frequency of use. The speech analysis system of claim 1 .
4. The speech analysis system further comprises an evaluation unit, The evaluation unit inputs the conversational expressions, the combinations of the conversational expressions, the conversation types of the conversational expressions, the model staff characteristics, and instructions for generating evaluations of conversational expressions in business negotiations into a generation artificial intelligence, and generates evaluations for the plurality of conversational expressions included in the plurality of conversational elements. The speech analysis system of claim 3 .
5. The speech analysis system further comprises an evaluation unit, The evaluation unit inputs the plurality of conversation sections, the development types of the conversation sections, the linking expressions, and instructions for generating an evaluation of the development of the conversation in the business negotiation into a generation artificial intelligence, and performs processing to generate an evaluation of a combination of the development key phrases and the linking expressions for dividing the conversation sections. The speech analysis system of claim 1 .
6. The conversation identification unit inputs the conversation voice data and an instruction to identify the voice data that is most frequently included among the plurality of voice data included in the conversation voice data to a generation artificial intelligence, thereby identifying the voice data that is most frequently included among the plurality of voice data included in the conversation voice data, and identifying text data based on the voice data as the plurality of conversation elements. The speech analysis system of claim 1 .
7. A voice analysis program that analyzes the content of conversations between staff at an event, the speech analysis program causes a computer to function as a conversation identification unit, a development identification unit, and an expression identification unit; the conversation identification unit identifies a plurality of conversation elements into which the content of the staff's conversation is divided based on conversational voice data including the content of the conversation between the staff and the attendee; the development identification unit identifies a conversation section in a series of conversations in which the plurality of conversation elements are grouped, based on conversational expressions included in the plurality of conversation elements; The series of conversation contents of the staff member includes a first conversation section, a second conversation section, and the conversation elements between the first conversation section and the second conversation section that do not belong to the first conversation section or the second conversation section; The expression identification unit identifies linking expressions from conversation expressions included in the conversation elements between the first conversation section and the second conversation section using the labels of the first conversation section and the labels of the second conversation section obtained by inputting instructions to a generation artificial intelligence to identify the conversation expressions included in the first conversation section to the second conversation section and the labels of the conversation sections, and linking expression indicators for identifying linking expressions, which are conversation expressions connecting the conversation sections, for each combination of the labels of the first conversation section and the labels of the second conversation section. Voice analysis program.
8. A voice analysis method for analyzing the content of conversations between staff at an event, comprising: The computer A process for identifying multiple conversation elements into which the staff's conversation content is divided based on conversation audio data including the conversation content between the staff and visitors; A process of identifying a conversation section in a series of conversations in which the plurality of conversation elements are grouped based on conversational expressions included in the plurality of conversation elements; The series of conversation contents of the staff member includes a first conversation section, a second conversation section, and the conversation elements between the first conversation section and the second conversation section that do not belong to the first conversation section or the second conversation section; a process of identifying linking expressions from conversational expressions included in the conversation elements between the first conversation section and the second conversation section, using the labels of the first conversation section and the labels of the second conversation section obtained by inputting instructions to a generation artificial intelligence to identify the conversational expressions included in the first conversation section to the second conversation section and the labels of the conversation sections, and linking expression indices for identifying linking expressions, which are conversational expressions connecting the conversation sections, for each combination of the labels of the first conversation section and the labels of the second conversation section; A speech analysis method that performs
Citation Information
Patent Citations
Talk support system and talk support method
JP2019197293A
Evaluation system, evaluation method, and computer program
JP2020160336A
Text analysis apparatus, method, and program
JP2023064114A
Communication review / feedback system
JP2023000533A