Conversation monitoring program, conversation monitoring device, conversation monitoring system, and conversation monitoring method

The conversation monitoring system addresses false detections by using audio feature scoring and tag embedding to accurately assess conversations, enhancing precision and evidence recording.

JP7893542B1Active Publication Date: 2026-07-22PLAYERS GATE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
PLAYERS GATE CO LTD
Filing Date
2026-05-15
Publication Date
2026-07-22

AI Technical Summary

Technical Problem

Existing conversation monitoring systems rely solely on text-based sentiment analysis, leading to false detections of normal conversations as abnormal and vice versa, failing to accurately monitor conversations.

Method used

A conversation monitoring system that utilizes feature scoring, tag embedding, and analysis models to assess audio features like volume and pitch, filters ambient sounds, and adjusts analysis based on thresholds, ensuring accurate monitoring by reducing false positives.

Benefits of technology

The system effectively reduces false detections and monitors conversations with high precision by incorporating audio features and intelligent filtering, providing detailed analysis and evidence recording.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893542000001_ABST
    Figure 0007893542000001_ABST
Patent Text Reader

Abstract

This invention provides a conversation monitoring program, a conversation monitoring device, a conversation monitoring system, and a conversation monitoring method that reduce false positives in automated conversation monitoring and enable highly accurate automated conversation monitoring. [Solution] The computer functions as a feature scoring unit 52 that scores the feature quantities of conversational audio, a feature determination unit 53 that determines whether or not to analyze the conversational audio based on the feature quantities, a conversational text acquisition unit 56 that acquires the transcribed conversational text, a tag-embedded text creation unit 59 that creates tag-embedded text by embedding event tags indicating speech changes into the conversational text, an analysis result acquisition unit 60 that acquires the analysis results from the conversational analysis model, a recording data storage unit 62 that acquires and saves conversational audio with abnormal conversational quality, and an analysis result transmission unit 63 that sends the analysis results to a management server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a conversation monitoring program, a conversation monitoring device, a conversation monitoring system, and a conversation monitoring method for monitoring conversations during customer service, telephone response, or conversations in the workplace, etc.

Background Art

[0002] Conventionally, technologies for monitoring conversations have been proposed. For example, Japanese Patent Laid-Open No. 2026-019057 discloses a system including means for collecting voice data in a workplace in real time, means for converting the voice data into text, means for analyzing the text-converted data and performing sentiment analysis, means for determining discomfort based on the sentiment analysis result, and means for storing related data and reporting it to the compliance department when the discomfort exceeds a specific threshold (Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the system described in Patent Document 1, sentiment analysis and discomfort are determined only based on the text-converted data obtained by converting voice data into text. Therefore, there is a risk of overlooking conversations that have no problem in the text content but are actually abnormal, or conversely, misdetecting conversations that seem to have a problem when only looking at the text content but are actually normal.

[0005] The present invention has been made to solve such problems, and an object thereof is to provide a conversation monitoring program, a conversation monitoring device, a conversation monitoring system, and a conversation monitoring method that can reduce false detections in automatic conversation monitoring and automatically monitor conversations with high accuracy.

[0006] Furthermore, the description of the problems mentioned above does not preclude the existence of other problems. Also, one aspect of the present invention does not need to solve all of the problems mentioned above. Moreover, it is possible to extract problems other than those mentioned above from the description, drawings, and claims. [Means for solving the problem]

[0007] The conversation monitoring program according to the present invention solves the problem of reducing false detections in automatic conversation monitoring and automatically monitoring conversations with high accuracy, and is a conversation monitoring program for monitoring conversations, comprising: a feature scoring unit that acquires a predetermined amount of conversation audio from an audio buffer that temporarily records the conversation audio of the conversation and scores the feature quantities of the conversation audio; a feature determination unit that determines whether or not to continue analyzing the conversation audio based on the feature quantities; a conversation text acquisition unit that acquires conversation text transcribed from the conversation audio by a speech recognition engine; a tag embedding text creation unit that creates tag embedding text in which event tags indicating sound changes that occurred in the conversation audio are embedded at the locations where the sound changes occurred in the conversation text; an analysis result acquisition unit that acquires analysis results regarding the quality of the conversation, which have been analyzed based on the tag embedding text by a conversation analysis model; a recording data storage unit that, if there is an abnormality in the quality of the conversation based on the analysis results, acquires a predetermined amount of conversation audio from the audio buffer and saves it in a recording data storage unit; and an analysis result transmission unit that transmits the analysis results to a management server, with the computer functioning as such.

[0008] Furthermore, in one aspect of the present invention, in order to solve the problem of suppressing unnecessary analysis of speech that is prone to false detection with a single feature quantity and improving analysis accuracy, the feature quantity includes a volume score calculated based on sudden fluctuations in volume, sudden occurrence of loud noises, long periods of silence, and sustained occurrence of loud noises; a pitch score calculated based on the magnitude of pitch fluctuations and sudden fluctuations in pitch; and a speech score obtained by summing the volume score and the pitch score. The feature quantity determination unit may continue the analysis of the conversational speech if any of the speech score, volume score, and pitch score are above a predetermined threshold.

[0009] Furthermore, in one aspect of the present invention, in order to solve the problem of analyzing even normal, peaceful conversations without omission while keeping analysis costs down, the computer may function as a waiting time adjustment unit that adjusts the waiting time before starting the analysis based on the previous analysis results if the feature determination unit determines that any one of the voice score, volume score, and pitch score is below a predetermined threshold.

[0010] Furthermore, in one aspect of the present invention, in order to define specific speech changes at the timing when a conversation is spoken and to enable a conversation analysis model to understand what was happening in context, the event tag may include at least one of the following: a volume increase tag indicating a sudden increase in volume, a sustained loud volume tag indicating that loud volume was sustained, a pitch increase tag indicating a sudden increase in pitch, a silence tag indicating a silent period, and a tempo acceleration tag indicating a sudden increase in conversation speed.

[0011] Furthermore, in one aspect of the present invention, in order to solve the problem of reducing API call costs and communication load by excluding conversational voices that are presumed to be ambient sounds from the analysis and analyzing only the voices that need to be analyzed, a computer may function as an unnecessary sound determination unit that determines whether or not the conversational voices are voices that do not need to be analyzed, and the unnecessary sound determination unit may not continue the analysis if the conversational voices satisfy any one of the following conditions (1) to (3), and may continue the analysis if the conversational voices do not satisfy any of the following conditions (1) to (3). (1) The average volume of the conversational audio is below the silence threshold. (2) The duration of the spoken speech is less than the threshold for determining if it is spoken. (3) The human voice ratio, which is the time ratio of human voices included in the conversational audio, is less than the human voice detection threshold.

[0012] Furthermore, in one aspect of the present invention, in order to solve the problem of reducing API call costs and communication load by excluding fragmented conversational text derived from ambient sounds from the analysis, a computer may be used as a text content determination unit to determine whether the content of the conversational text is normal or not based on the number of words, the number of characters, and the percentage of Japanese characters contained in the conversational text.

[0013] Furthermore, in one aspect of the present invention, in order to solve the problem of detecting similar duplicate statements and effectively eliminating slightly different repeated statements or re-analysis / duplicate recording of the same conversation, a computer may be used as a similar text determination unit to determine whether the conversation text is similar to conversation texts that have been analyzed in the past, and to continue analyzing only conversation texts that are not similar, while not continuing to analyze similar conversation texts.

[0014] Furthermore, in one aspect of the present invention, in order to solve the problem of obtaining more detailed analysis results when the quality of the transcription is low or when there is uncertainty in the analysis and judgment of the conversation analysis model, a computer may be used as a detailed analysis determination unit to determine whether or not to have the conversation analysis model perform a detailed analysis regarding the quality of the conversation based on the reliability of the conversation text and the confidence level of the analysis results.

[0015] Furthermore, in one aspect of the present invention, in order to solve the problem of immediately detecting emergencies such as sudden loud noises and saving conversational audio as evidence with the highest priority, if the voice score is above the emergency threshold and the human voice ratio, which is the time ratio of human voices included in the conversational audio, is above the human voice determination threshold, the recording data storage unit may immediately start recording the conversational audio and continue analyzing the conversational audio.

[0016] The conversation monitoring device according to the present invention solves the problem of reducing misanalysis in conversation monitoring and monitoring conversations with high accuracy, and comprises: a feature scoring unit that acquires a predetermined amount of conversation audio from an audio buffer that temporarily records the conversation audio and scores the feature quantities of the conversation audio; a feature determination unit that determines whether or not to continue analyzing the conversation audio based on the feature quantities; a conversation text acquisition unit that acquires conversation text transcribed from the conversation audio by a speech recognition engine; a tag embedding text creation unit that creates tag embedding text in which event tags indicating sound changes that occurred in the conversation audio are embedded at the locations where the sound changes occurred in the conversation text; an analysis result acquisition unit that acquires analysis results regarding the quality of the conversation, which have been analyzed by a conversation analysis model based on the tag embedding text; a recording data storage unit that, if there is an abnormality in the quality of the conversation based on the analysis results, acquires a predetermined amount of conversation audio from the audio buffer and stores it in a recording data storage unit; and an analysis result transmission unit that transmits the analysis results to a management server.

[0017] The conversation monitoring system according to the present invention, in order to reduce misanalysis in conversation monitoring and solve the problem of monitoring conversations with high accuracy, comprises a conversation monitoring device, a speech recognition engine for transcribing the conversation text, a conversation analysis model for analyzing the quality of the conversation audio based on the tag-embedded text, and a management server for storing the analysis results.

[0018] The conversation monitoring method according to the present invention is a conversation monitoring method for monitoring conversations in order to reduce misanalysis in conversation monitoring and solve the problem of monitoring conversations with high precision. The method includes: obtaining conversation audio for a predetermined time from an audio buffer that temporarily records the conversation audio of the conversation; a feature quantity scoring step of scoring the feature quantity of the conversation audio; a feature quantity determination step of determining whether to continue analyzing the conversation audio based on the feature quantity; a conversation text acquisition step of obtaining conversation text transcribed from the conversation audio by a speech recognition engine; a tag embedding text creation step of creating tag-embedded text in which an event tag indicating a voice change occurring in the conversation audio is embedded at the occurrence position of the voice change in the conversation text; an analysis result acquisition step of obtaining an analysis result regarding the quality of the conversation analyzed based on the tag-embedded text by a conversation analysis model; a recorded data storage step of obtaining conversation audio for a predetermined retrospective time from the audio buffer and storing it in a recorded data storage unit when there is an abnormality in the quality of the conversation based on the analysis result; and an analysis result transmission step of transmitting the analysis result to a management server.

Effect of the Invention

[0019] According to the present invention, false detection in automatic conversation monitoring can be reduced, and conversations can be automatically monitored with high precision.

Brief Description of the Drawings

[0020] [Figure 1] It is a block diagram showing an embodiment of a conversation monitoring program, a conversation monitoring device, and a conversation monitoring system according to the present invention. [Figure 2] It is the first half of a flowchart showing a conversation monitoring method executed by the conversation monitoring program and the conversation monitoring device of the present embodiment. [Figure 3] It is the second half of a flowchart showing a conversation monitoring method executed by the conversation monitoring program and the conversation monitoring device of the present embodiment. [Figure 4] It is a flowchart showing the emergency processing of the present embodiment.

Best Mode for Carrying Out the Invention

[0021] Hereinafter, an embodiment of a conversation monitoring program, a conversation monitoring device, a conversation monitoring system, and a conversation monitoring method according to the present invention will be described with reference to the drawings.

[0022] [1] System Configuration The conversation monitoring system 100 of the present embodiment is a system for automatically analyzing conversation audio collected in real time during customer service, telephone response, or in the workplace, etc., and evaluating the quality and risk of the conversation. In the present embodiment, as shown in FIG. 1, the conversation monitoring system 100 includes a conversation monitoring device 1 for monitoring conversations, a speech recognition engine 11 for performing speech recognition processing on conversation audio, a conversation analysis model 12 for analyzing the quality of conversations, a speech analysis engine 13 for performing speech analysis processing on conversation audio, and a management server 14 for aggregating and managing analysis results from the conversation monitoring device 1. Hereinafter, each configuration will be described.

[0023] Note that in the present embodiment, the speech recognition engine 11, the conversation analysis model 12, and the speech analysis engine 13 use external APIs (Application Programming Interfaces) provided as cloud services, but are not limited to this configuration, and may be constructed and used on-premises within an intranet or implemented within the conversation monitoring device 1. Also, in FIG. 1, only one conversation monitoring device 1 is shown, but it is also possible to simultaneously monitor a plurality of conversation monitoring devices 1.

[0024] [2] Conversation Monitoring Device 1 The conversation monitoring device 1 is configured by a computer such as a smartphone, a tablet terminal, or a personal computer, and automatically monitors conversations in real time. In the present embodiment, as shown in FIG. 1, the conversation monitoring device 1 mainly has a communication means 2, a voice input means 3, a storage means 4, and an arithmetic processing means 5. Hereinafter, each constituent means will be described.

[0025] Communication means 2 implements communication functions in the conversation monitoring device 1. In this embodiment, communication means 2 is composed of a communication module compatible with wireless LAN standards such as Wi-Fi (registered trademark), carrier network standards such as fifth-generation mobile communication network (5G), or wired Ethernet (registered trademark) standards. Communication means 2 transmits and receives various data between the speech recognition engine 11, conversation analysis model 12, speech analysis engine 13, and management server 14 via a communication network such as the Internet.

[0026] The voice input means 3 is for capturing the conversation of the monitored subject as conversational audio (electrical signals). In this embodiment, the voice input means 3 is composed of a microphone that collects sound in the space where the conversation takes place, or an audio interface that acquires audio signals from a telephone line or the like.

[0027] The storage means 4 stores various types of data and also functions as a working area when the arithmetic processing means 5 performs arithmetic processing. In this embodiment, the storage means 4 is composed of flash memory, solid-state drive (SSD), hard disk drive (HDD), ROM (Read Only Memory), RAM (Random Access Memory), etc., and as shown in Figure 1, it has a program storage unit 41, an audio buffer 42, a recording data storage unit 43, and an analysis result storage unit 44.

[0028] The program storage unit 41 has the conversation monitoring program 1a of this embodiment installed. The arithmetic processing means 5 then executes the conversation monitoring program 1a, thereby realizing the computer as the conversation monitoring device 1 with the various functional units described later. Note that the usage of the conversation monitoring program 1a is not limited to the above configuration. For example, the conversation monitoring program 1a may be stored on a non-temporary recording medium that can be read by a computer, such as a memory card or USB memory, and then read and executed from that recording medium.

[0029] The audio buffer 42 is for temporarily recording conversational audio input from the audio input means 3 in real time. In this embodiment, the audio buffer 42 is configured as a ring buffer, and for example, it always holds the most recent 300 seconds of conversational audio, and automatically overwrites and deletes older conversational audio. This allows the recording data storage unit 62, described later, to acquire conversational audio for a predetermined period of time prior to the time the abnormality occurred when an abnormality in conversational quality is detected.

[0030] The recording data storage unit 43 stores the conversational audio that has been selected to be recorded. In this embodiment, the recording data storage unit 43 stores the conversational audio encrypted by the recording data storage unit 62. The conversational audio stored in the recording data storage unit 43 is automatically deleted after analysis is complete and is also uploaded and stored separately in cloud storage.

[0031] The analysis result storage unit 44 stores intermediate data and final analysis results generated during the analysis of conversational audio. In this embodiment, the analysis result storage unit 44 stores the following: the characteristic quantities of conversational audio (volume score, pitch score, voice score), the conversational text transcribed from the conversational audio, the tagged text, the analysis results obtained from the conversational analysis model 12, and the analysis results obtained from the speech analysis engine 13.

[0032] The analysis results from the conversation analysis model 12 include a conversation quality score (e.g., score 1 (good), score 2 (stable), score 3 (tense), score 4 (high load), score 5 (needs review)) and a confidence level of the analysis results (e.g., 0.0 to 1.0). The analysis results from the speech analysis engine 13 include speaker separation results, sentiment analysis results, speech time ratios, and content safety.

[0033] The arithmetic processing unit 5 consists of a CPU (Central Processing Unit) and an NPU (Neural Processing Unit), and by executing the conversation monitoring program 1a installed in the program storage unit 41, it functions as shown in Figure 1: a conversation audio recording unit 51, a feature scoring unit 52, a feature determination unit 53, a waiting time adjustment unit 54, an unwanted sound determination unit 55, a conversation text acquisition unit 56, a text content determination unit 57, a similar text determination unit 58, a tag-embedded text creation unit 59, an analysis result acquisition unit 60, a detailed analysis determination unit 61, a recording data storage unit 62, and an analysis result transmission unit 63. Each of these functional units will be described in more detail below.

[0034] The conversation audio recording unit 51 is a functional unit that temporarily records conversation audio input from the audio input means 3 in the audio buffer 42. In this embodiment, while the conversation monitoring program 1a is running, the conversation audio recording unit 51 continuously samples and quantizes the conversation audio input from the audio input means 3 in real time and writes it to the audio buffer 42 as a digital signal.

[0035] The feature scoring unit 52 is a functional unit that periodically acquires a predetermined amount of conversational audio (e.g., 60 seconds) from the audio buffer 42 at predetermined analysis cycles (e.g., 60 seconds) and scores the feature quantities of the conversational audio. These feature quantities are used to screen whether the conversation is in an abnormal state (e.g., argument, shouting, complaint, excessive silence, etc.) based on its auditory characteristics, and in this embodiment, they include a volume score, a pitch score, and a voice score.

[0036] The volume score is a numerical representation of the volume characteristics of conversation, calculated based on sudden volume fluctuations, sudden occurrences of loud noises, prolonged silences, and sustained loud noises. In this embodiment, sudden volume fluctuations are scored when the volume fluctuation range exceeds a predetermined value and the average volume also exceeds a predetermined value. Sudden occurrences of loud noises are scored when the peak volume exceeds a predetermined value, the difference between the peak volume and the average volume exceeds a predetermined value, and the duration of the peak volume is greater than or equal to a predetermined value. Furthermore, prolonged silences are scored when the longest duration of silence is greater than or equal to a predetermined value and the number of silences is greater than or equal to a predetermined value. Finally, sustained loud noises are scored when the average volume exceeds a predetermined value and the duration of the loud noise is greater than or equal to a predetermined value. The final volume score is calculated by weighting and adding up each of these items.

[0037] The pitch score is a numerical representation of the characteristics of voice pitch, calculated based on the magnitude of pitch (fundamental frequency) fluctuations and sudden pitch changes. In particular, in conversations involving anger or excitement, pitch tends to fluctuate significantly or rise sharply, making the pitch score an important indicator for detecting abnormalities in conversation quality. In this embodiment, points are added for the magnitude of pitch fluctuations when the pitch standard deviation exceeds a predetermined value and the human voice ratio, which is the time ratio of human voice, exceeds a predetermined value. Points are added for sudden pitch changes when the number of pitch changes per second exceeds a predetermined value and the human voice ratio exceeds a predetermined value. The human voice ratio can be calculated, for example, from the estimated value of the fundamental frequency (F0) using the YIN algorithm.

[0038] The voice score is a comprehensive score obtained by adding the volume score and the pitch score. In this embodiment, the voice score is scored with a maximum of 100 points, which is the sum of the volume score (maximum 75 points) and the pitch score (maximum 25 points). However, the feature scoring method is not limited to the above, and any scoring method that can appropriately evaluate the characteristics of the voice in conversational speech is acceptable.

[0039] The feature determination unit 53 is a functional unit that determines whether or not to continue analyzing the conversational speech (text conversion by the speech recognition engine 11 or analysis by the conversation analysis model 12) based on the features calculated by the feature scoring unit 52. In this embodiment, the feature determination unit 53 continues analyzing the conversational speech if the voice score is above the voice threshold, the volume score is above the volume threshold, and the pitch score is above the pitch threshold. By making it a requirement that all of the voice score, volume score, and pitch score be above a predetermined threshold, unnecessary analysis of speech that is prone to misdetection with a single feature, such as loud noises (volume score only) or humming (pitch score only), is suppressed, and the analysis accuracy is improved.

[0040] In this embodiment, the feature determination unit 53 also determines whether emergency processing is required. Specifically, if the voice score is above an emergency threshold that is higher than the voice threshold, and the proportion of human voices included in the conversation voice is above a predetermined human voice determination threshold, it is strongly inferred that a highly urgent situation (e.g., violent verbal abuse, screaming) has occurred. In such cases, the feature determination unit 53 immediately instructs the recording data storage unit 62 to record the conversation voice and, in parallel, continues the analysis processing of the conversation voice.

[0041] The waiting time adjustment unit 54 is a functional unit that dynamically adjusts the waiting time until the next cycle of analysis begins, based on the determination result by the feature determination unit 53. Specifically, if the feature determination unit 53 determines that any one of the voice score, volume score, and pitch score is below a predetermined threshold, it is inferred that the conversation is peaceful. In this case, the waiting time adjustment unit 54 adjusts the time until the analysis begins, that is, the waiting time until the feature scoring unit 52 performs the scoring process for the next analysis cycle, based on the previous analysis result obtained by the analysis result acquisition unit 60.

[0042] For example, if the conversation quality score obtained as a result of the analysis is score 2 or lower (good / stable), there are no signs of abnormality in the conversation, so the waiting time is set to the maximum (e.g., 240 seconds). This reduces the number of calls to the speech recognition engine 11 and conversation analysis model 12, which incur costs each time they are called, thus keeping costs down. If the conversation quality score is score 3 (tense), there are signs in the conversation that require attention, so the waiting time is set to an intermediate value (e.g., 120 seconds). On the other hand, if the conversation quality score is score 4 or higher (high load / needs review), there is a possibility that a serious situation is continuing in the conversation, so the waiting time is set to the minimum (e.g., 60 seconds). This ensures that the conversation audio is analyzed frequently, maintaining a state where it can be recorded immediately.

[0043] The unwanted sound determination unit 55 is a functional unit that determines whether or not the conversational audio is audio that does not require analysis. In this embodiment, the unwanted sound determination unit 55 does not continue the analysis if the conversational audio satisfies any one of the following conditions (1) to (3), and continues the analysis if the conversational audio does not satisfy any of the following conditions (1) to (3). (1) The average volume of the conversation is below the silence threshold. (2) The duration of the spoken speech is less than the threshold for determining if it is spoken. (3) The human voice ratio, which is the proportion of time spent with human voices in the conversational audio, is below the human voice detection threshold.

[0044] If the above condition (1) is met, the conversation audio is almost silent and it is presumed that there is no conversation to monitor, so the analysis will not continue. As a silence detection threshold, for example, the average volume for each analysis cycle may be accumulated for the most recent 10 cycles, the lowest 10 percentile of which may be used as the noise floor of the audio input means 3, and this noise floor + 5dB may be used as the silence detection threshold.

[0045] If the above condition (2) is met, the sound is presumed to be an ambient sound such as coughing, door opening and closing (instantaneous impact), passing car noise (sustained low frequency), air conditioning or ventilation fan noise (steady noise), or keyboard or mouse operation sounds, and the analysis will not continue. The duration of voiced sounds can be detected, for example, using a Voice Activity Detection (VAD) algorithm that automatically distinguishes between human speech segments and silent / ambient sound segments from conversational audio.

[0046] If the above condition (3) is met, the sound is presumed to be a phone ringtone or notification sound, rather than a human voice, and the analysis will not continue. The proportion of human voices can be calculated, for example, from the estimated value of the fundamental frequency (F0) using the YIN algorithm. However, in noisy environments, the detection accuracy of the YIN algorithm decreases, so it is preferable to dynamically adjust the human voice detection threshold according to the SNR (signal-to-noise ratio) estimated from the noise floor and average volume. That is, the human voice detection threshold should be set higher in quiet environments and lower in noisy environments.

[0047] As described above, the unnecessary sound detection unit 55 automatically eliminates voices that do not need to be analyzed, thereby reducing unnecessary calls to the speech recognition engine 11 and conversation analysis model 12, and lowering costs. However, if the speech recognition engine 11 and conversation analysis model 12 are used on-premise or installed within the conversation monitoring device 1, there is no need to worry about the cost of API calls, and therefore the standby time adjustment unit 54 and unnecessary sound detection unit 55 do not need to be provided.

[0048] The conversation text acquisition unit 56 is a functional unit that acquires conversation text transcribed from conversation audio by the speech recognition engine 11. In this embodiment, the conversation text acquisition unit 56 transmits the conversation audio for one analysis cycle to the speech recognition engine 11. The speech recognition engine 11 then acquires the transcribed conversation text along with the confidence level of the conversation text, which is used in subsequent processing. The confidence level of the conversation text can be calculated, for example, from the logarithmic probability returned for each token (word / character) of the transcription result, and will decrease if the audio is unclear or there is a lot of noise.

[0049] Furthermore, the speech recognition engine 11 has a known problem in that if there is silence in the audio data to be recognized, it will infinitely repeat the preceding phrase. For this reason, it is preferable to remove the silent portion at the end of the conversation audio as a hallucination prevention (cause removal) process before sending the conversation audio to the speech recognition engine 11. Also, after obtaining the conversation text from the speech recognition engine 11, it is preferable to compress the same phrase to two repetitions if it is repeated three or more times as a hallucination removal (symptom correction) process.

[0050] Furthermore, as the speech recognition engine 11, a cloud-based speech recognition engine 11 such as the lightweight, high-speed model of the GPT-4(registered trademark) series (OpenAI Inc.) may be used, or an on-premise speech recognition engine 11 that has been tuned or customized for a specific industry may be used.

[0051] The text content determination unit 57 is a functional unit that determines whether the content of the conversation text is normal or not. In this embodiment, the text content determination unit 57 determines whether the content of the conversation text is normal or not based on the number of words, the number of characters, and the percentage of Japanese characters in the conversation text acquired by the conversation text acquisition unit 56, and excludes fragmented conversation text (BGM, in-store announcements, etc.) that originates from ambient sounds.

[0052] If the number of words in the conversation text is less than a predetermined value (e.g., 4 words), or if the number of characters in the conversation text is less than a predetermined value (e.g., 20 characters), it is presumed to be meaningless. Also, if the proportion of Japanese characters in the conversation text is less than a predetermined value (e.g., 60%), it is presumed to be meaningless in Japanese. Therefore, if any of these conditions are met, the text content determination unit 57 determines that the content of the conversation text is not normal and stops the analysis.

[0053] The similar text determination unit 58 is a functional unit that determines whether a conversation text is similar to a conversation text that has been analyzed in the past. In this embodiment, the similar text determination unit 58 determines whether the conversation text acquired by the conversation text acquisition unit 56 is similar to a conversation text that has been analyzed in the past and is stored in the analysis result storage unit 44. Based on the result of this determination, the analysis of similar conversation texts is not continued, and only the analysis of non-similar conversation texts is continued. As a result, even if they are not a perfect match, similar duplicate statements are detected, so slightly different repeated statements and re-analysis / duplicate recording of the same conversation are effectively eliminated.

[0054] The method for determining similarity between conversational texts is, for example, to break down each conversational text into consecutive pairs of two characters, and then determine whether the similarity (Bigram Jaccard similarity), calculated as the ratio of the number of common characters in each conversational text to the total number of characters, is above a predetermined threshold (e.g., 80%). For example, "おはよう" is broken down into {おは,はよ,よう}, and "おはよ" is broken down into {おは,はよ}. Therefore, the similarity between the two is 2 (common) / 3 (total) = 66.7% (below the threshold), and they are selected for analysis.

[0055] The tag-embedded text creation unit 59 is a functional unit that creates tag-embedded text by embedding event tags into the conversation text. In this embodiment, the event tags indicate speech changes that occurred in the conversation audio and include a volume increase tag indicating a sudden increase in volume, a high volume duration tag indicating that a high volume was sustained, a pitch increase tag indicating a sudden increase in pitch, a silence tag indicating a silent section, and a tempo acceleration tag indicating a sudden increase in conversation speed.

[0056] The tag-embedded text creation unit 59 then creates tag-embedded text by embedding the event tags mentioned above at the locations where speech changes occur in the conversation text. Specifically, the tag-embedded text creation unit 59 calculates the insertion position by multiplying the relative position (time ratio) of the occurrence time within the analysis cycle by the number of characters in the text, and inserts each event tag by snapping it to the nearest word boundary (space, punctuation, etc.) within 10 characters before or after. In addition, if there are multiple event tags, the unit prevents positional misalignment by inserting them sequentially from the end of the conversation text.

[0057] The following is an example of embedded tag text. Note that the text in parentheses indicates the event tag. "Why haven't you done this yet? (Volume increase: 52→78dB) (Pitch increase: 180→340Hz) How many times do I have to tell you?!" (Loud volume sustained: 75dB) (Silence: 12 seconds) "...I understand." "I told you last week, the deadline is today, why didn't you contact me? (Tempo acceleration: 2.1→4.8 words / second) Do you have any idea how much trouble this is causing me?!" (Loud volume sustained: 76dB)

[0058] As a result of the tag-embedded text described above, nonverbal nuances (such as anger, excitement, and impatience) that cannot be gleaned from the conversation text alone can be understood through the changes in voice at the time each line is spoken. Therefore, the conversation analysis model 12 can deeply understand the context of what was happening in the conversation and perform highly accurate analysis. In this embodiment, all five types of event tags are used, but at least one may be used. Furthermore, the event tags are not limited to the five types described above, and any tags that indicate changes in voice may be set as appropriate.

[0059] The analysis result acquisition unit 60 is a functional unit that acquires analysis results regarding the quality of conversations analyzed by the conversation analysis model 12. In this embodiment, the analysis result acquisition unit 60 transmits the feature quantities (voice score, volume score, and pitch score) scored by the feature quantity scoring unit 52 and the tag-embedded text created by the tag-embedded text creation unit 59 to the conversation analysis model 12. The unit then acquires the conversation quality score, the confidence level of the analysis results, etc., from the conversation analysis model 12 as analysis results.

[0060] In this embodiment, the conversation analysis model 12 is composed of a large-scale language model that excels at understanding and analyzing the meaning of text, such as Claude Haiku (Anthropic). The system prompts given to the conversation analysis model 12 are composed of a combination of parts, such as a role definition indicating that it will operate as a conversation quality analysis engine, criteria for determining the conversation quality score, important rules such as prohibiting a score of 3 or higher based solely on the voice score, input quality guidelines such as recommending a score of 1-2 for texts shorter than 20 characters, tag restrictions such as not determining the conversation quality score based solely on event tags, industry templates defining criteria set for each industry, and analysis options that dynamically switch the criteria based solely on the voice score, which are set in each industry template.

[0061] In this embodiment, the conversation quality score is evaluated comprehensively on a 5-point scale (score 1 to 5) based on the features of the conversational speech and the tagged embedded text. Furthermore, the criteria for determining the conversation quality score include a core axis that is always applied to all industries and a domain axis that is applied individually to each industry, with criteria for determining scores from 1 to 5 defined for each axis.

[0062] The core axes include, for example, an "aggressive behavior analysis" axis that evaluates the presence and severity of aggressive behavior such as insult, intimidation, threats, and discrimination among staff; an "emotional escalation" axis that comprehensively evaluates the progression of emotional tension and vocal tags (increased volume, pitch fluctuations) for both customers and staff; and an "staff response quality" axis that evaluates the appropriateness of listening, explanation, confirmation attitude, and presence or absence of overbearing behavior among staff.

[0063] Furthermore, as a domain axis, for example, in the case of a financial institution, multiple axes may be set, such as the "fraud analysis" axis, the "money laundering" axis, the "inappropriate sales (investment)" axis, the "conflict of interest" axis, the "excessive lending" axis, the "insider information" axis, the "psychological burden" axis, and the "inappropriate handling of personal information" axis.

[0064] Therefore, the conversation analysis model 12 refers to the industry template corresponding to the industry given in the role definition, extracts the judgment criteria for the core axis and the corresponding domain axis, and dynamically injects them into the system prompt. Then, it outputs the maximum value of the evaluated score for each axis as the final conversation quality score. In addition, the confidence level of the analysis result is a value that decreases in cases where the context is ambiguous or the judgment is on the borderline, and is output as a numerical value between 0.0 and 1.0, for example.

[0065] The detailed analysis determination unit 61 determines whether or not to have the conversation analysis model 12 perform a detailed analysis on the quality of the conversation, based on the reliability of the conversation text and the confidence level of the analysis results. In this embodiment, the detailed analysis determination unit 61 obtains the reliability of the conversation text obtained by the conversation text acquisition unit 56 and the confidence level of the analysis results obtained by the analysis result acquisition unit 60. If the reliability of the conversation text is above a predetermined threshold (e.g., 65%) and the confidence level of the analysis results is above a predetermined threshold (e.g., 0.75), the unit determines that the confidence level of the analysis is high and uses the analysis results from the conversation analysis model 12 as the final analysis results.

[0066] On the other hand, if the confidence level of the conversation text is below a predetermined threshold, the system determines that the quality of the transcription itself is low, resulting in low confidence in the analysis and requiring re-analysis. Similarly, if the confidence level of the analysis result is below a predetermined threshold, the system determines that the conversation analysis model 12 is unsure of its decision, resulting in low confidence in the analysis and requiring re-analysis. In other words, if at least one of the confidence level of the conversation text or the confidence level of the analysis is below a threshold, the detailed analysis determination unit 61 transmits the conversation audio to the speech analysis engine 13 to obtain the analysis result, and also transmits this analysis result to the conversation analysis model 12 as additional information to perform a detailed analysis.

[0067] In this embodiment, the speech analysis engine 13 is configured as a cloud service type speech analysis engine 13 such as AssemblyAI (AssemblyAI Inc.). By analyzing the conversational audio obtained from the detailed analysis and determination unit 61, the system outputs analysis results such as speaker separation results (information identifying the speaker of each statement in chronological order), sentiment analysis results (results of determining positive / negative / neutral sentiment on an utterance basis), speech time ratio (amount of speech and occupancy rate for each speaker), and safety of conversational content (presence or absence of sexual, violent, or discriminatory expressions).

[0068] In this embodiment, when the conversation analysis model 12 receives an instruction from the detailed analysis determination unit 61 to perform a detailed analysis, it recalculates the conversation quality score based on the additional information described above (analysis results of the speech analysis engine 13) and its own scoring indicators. Specifically, five scoring indicators are defined: the percentage of negative emotions, the number of overlaps between statements by different speakers, the number of long periods of silence, volume fluctuation (standard deviation), and average volume. A final conversation quality score (score 1 to 5) is then calculated according to the total score of each scoring indicator. This final conversation quality score is obtained by the analysis result acquisition unit 60 as the result of the detailed analysis.

[0069] In this embodiment, the conversation monitoring device 1 has a manual mode for manually recording conversation audio, a standard analysis mode for having the conversation analysis model 12 perform the analysis, and a detailed analysis mode for performing a detailed analysis when the confidence level of the analysis is low. However, the device is not limited to this configuration, and it is sufficient to have at least a standard analysis mode. Therefore, if there is no detailed analysis mode, it is not necessary to provide the detailed analysis determination unit 61.

[0070] The recording data storage unit 62 is a functional unit that stores conversation audio in which an abnormality in conversation quality has been detected. In this embodiment, the recording data storage unit 62 determines whether or not there is an abnormality in conversation quality based on the analysis results obtained by the analysis result acquisition unit 60. Specifically, if the conversation quality score is above a predetermined threshold (e.g., score 3), it is determined that there is some kind of abnormality in the conversation, and if the conversation quality score is below the predetermined threshold, it is determined that there is no abnormality. If an abnormality is found, the recording data storage unit 62 acquires conversation audio for a predetermined retrospective time from the audio buffer 42, encrypts it, and stores it in the recording data storage unit 43. This makes it possible to later check the overall picture of the conversation in which the problem occurred.

[0071] The retrospective recording time can be set arbitrarily, but it is preferable to set it to a time (e.g., 295 seconds) that is slightly less than the time equivalent to the total capacity of the audio buffer 42 (e.g., 300 seconds) by a margin (e.g., 5 seconds). This prevents corruption of the audio data because, as the audio buffer 42 is constantly being overwritten from the oldest area, the oldest area (margin portion) will not be read by the recording data storage unit 62. Furthermore, it is preferable to dynamically calculate the retrospective recording time by adding the lag (e.g., several seconds to tens of seconds) between the time of one analysis cycle (e.g., 60 seconds) and the actual start of recording. This ensures that the conversation audio at the moment the recording instruction is given is reliably recorded.

[0072] The analysis result transmission unit 63 is a functional unit that transmits various analysis results to the management server 14. In this embodiment, the transmitted analysis results include the conversation quality score and confidence level of the analysis results obtained by the analysis result acquisition unit 60, as well as features scored by the feature scoring unit 52, and conversation text and confidence level obtained by the conversation text acquisition unit 56.

[0073] In this embodiment, the management server 14 is configured with a real-time synchronous cloud database such as Cloud Firestore (Google Inc.), and centrally aggregates and stores analysis results transmitted from multiple conversation monitoring devices 1, immediately reflecting them on the web management dashboard. This web management dashboard displays the number of analysis items for the day, the number of alerts requiring attention (e.g., score of 3 or higher), the number of unaddressed items, a list of recent alerts, etc., allowing users to check the monitoring status of conversation audio in real time and view analysis reports for each alert. In addition, the management server 14 has a notification function that sends alerts to the administrator's chat tool or email when escalation is necessary.

[0074] [3] Effect Next, the operation of the conversation monitoring program 1a, conversation monitoring device 1, conversation monitoring system 100, and conversation monitoring method of this embodiment will be explained using the flowcharts in Figures 2 to 4. The following explanation will focus on the process when executed in detailed analysis mode.

[0075] When monitoring a conversation using the conversation monitoring system 100 of this embodiment, as shown in Figure 2, first, the conversation audio recording unit 51 continuously and temporarily records the conversation audio input from the audio input means 3 in the audio buffer 42 (Step S1: Conversation Audio Recording Step). Then, the feature scoring unit 52 periodically acquires a predetermined amount of conversation audio from the audio buffer 42 in a predetermined analysis cycle and scores the feature quantities of the conversation audio (Step S2: Feature Scoring Step). As a result, acoustic features that cannot be determined solely from the content of the conversation are quantified, improving the analysis accuracy of the conversation analysis model 12.

[0076] Next, the feature determination unit 53 determines whether the voice score is above the emergency threshold (step S3). If the result of this determination is that the voice score is above the emergency threshold (step S3: YES), the emergency processing shown in Figure 4 (step S4) is executed. Specifically, the feature determination unit 53 first determines whether the human voice ratio of the conversational voice is above the human voice determination threshold (step S41: human voice ratio determination step).

[0077] If the result of this determination is that the proportion of human voices is equal to or greater than the human voice detection threshold (step S41: NO), the recording data storage unit 62 immediately starts recording the conversational audio (step S42) and continues the analysis process described later (steps S13 to S20, S23) in parallel (step S43). This prioritizes the preservation of audio evidence in emergency situations such as sudden loud noises. On the other hand, if the proportion of human voices is less than the human voice detection threshold (step S41: YES), it is presumed that the audio originates from something other than human conversation, and the emergency processing is terminated.

[0078] Next, if the result of the judgment in step S3 is that the voice score is below the emergency threshold (step S3: NO), the feature determination unit 53 makes a judgment on the three features (steps S5-7: feature determination steps). Specifically, if the voice score is above the voice threshold (step S5: YES), the volume score is above the volume threshold (step S6: YES), and the pitch score is above the pitch threshold (step S7: YES), the process proceeds to step S10 to continue the analysis of the conversational voice. In this way, by making it a mandatory condition that all three features are above the threshold, unnecessary analysis of voices that are easily misdetected by a single feature, such as loud noises (volume score only) or humming (pitch score only), is suppressed, and the analysis accuracy is improved.

[0079] On the other hand, if any one of the three features is below the threshold (steps S5-S7: any is NO), the waiting time adjustment unit 54 adjusts the waiting time until the next analysis cycle according to the previous analysis results (step S8: waiting time adjustment step). If the waiting time has not elapsed (step S9: NO), the system waits until the next analysis cycle (step S25) and then returns to step S2. If the waiting time has elapsed, the system proceeds to step S10 and continues analyzing the conversational audio. This ensures that even normal, peaceful conversations are analyzed without omission while keeping analysis costs down.

[0080] In this embodiment, the standby time adjustment unit 54 sets the standby time to the maximum if the previous analysis result was good or stable, as there are no signs of abnormality in the conversation. This reduces the number of calls to the speech recognition engine 11 and the conversation analysis model 12, thereby lowering costs. If the previous analysis result was tense, the standby time is set to an intermediate value, as there are signs in the conversation that require attention. On the other hand, if the previous analysis result was high load or required review, the standby time is set to the minimum, as there is a possibility that a serious situation is continuing in the conversation. This ensures that the conversation audio is analyzed frequently, maintaining a state where it can be recorded immediately.

[0081] Next, the unwanted sound determination unit 55 determines whether or not the audio does not require analysis based on the three conditions (1) to (3) described above (steps S10 to S12: unwanted sound determination steps). Specifically, the unwanted sound determination unit 55 determines that the audio is unsuitable for analysis if the average volume of the conversational audio is below the silence determination threshold (step S10: YES), or if the voiced duration is below the voiced duration threshold (step S11: YES), or if the human voice ratio is below the human voice determination threshold (step S12: YES), and proceeds to step S24, which will be described later. On the other hand, if the average volume, voiced duration, and human voice ratio are all above the threshold (steps S10 to S12: all NO), the analysis continues and proceeds to step S13.

[0082] This means that conversational audio that is presumably ambient noise, such as near silence, coughing or door opening / closing sounds (momentary impact), passing car sounds (sustained low frequency), air conditioning or ventilation fan sounds (steady noise), keyboard and mouse operation sounds, and phone ringtones / notification sounds, will not be analyzed. Only audio that requires analysis will be analyzed, thus reducing API call costs and communication load.

[0083] Next, as shown in Figure 3, the conversation text acquisition unit 56 transmits the conversation audio to be analyzed to the speech recognition engine 11 and acquires the conversation text and its confidence level (Step S13: Conversation Text Acquisition Step). Then, based on the number of characters in the acquired conversation text, the percentage of Japanese characters, the match rate with the previous text, etc., the text content determination unit 57 determines whether the text content is normal or not (Step S14: Text Content Determination Step). If the result of this determination is that it is not normal (Step S14: NO), the analysis is stopped and the process proceeds to Step S24. As a result, fragmented conversation text originating from ambient sounds (BGM, in-store announcements, etc.) is excluded from the analysis, thus reducing API call costs and communication load. On the other hand, if the text content is determined to be normal (Step S14: YES), the process proceeds to Step S15.

[0084] Next, the similar text determination unit 58 determines whether the conversation text to be analyzed is similar to text that has been analyzed in the past (step S15: similar text determination step). If the determination is found to be similar (step S15: YES), the process proceeds to step S24. This effectively eliminates slightly different repeated statements, re-analysis of the same conversation, and duplicate recording. On the other hand, if the determination is found to be not similar (step S15: NO), the process proceeds to step S16.

[0085] Next, the tag-embedded text creation unit 59 creates tag-embedded text by embedding the aforementioned event tags into the conversation text (Step S16: Tag-embedded text creation step). With this tag-embedded text, the conversation analysis model 12 grasps nonverbal nuances (anger, excitement, impatience, etc.) through changes in speech at the time each line was spoken, which cannot be read from the conversation text alone. As a result, the context of what was happening in the conversation is deeply understood, reducing false positives in automated conversation monitoring.

[0086] Next, the analysis result acquisition unit 60 transmits the feature quantities scored in step S2 and the tag-embedded text created in step S16 to the conversation analysis model 12 (step S17), and obtains the conversation quality score and the confidence level of the analysis results as analysis results (step S18: analysis result acquisition step). As a result, the so-called AI (conversation analysis model 12) obtains a conversation quality score that comprehensively evaluates not only the text content but also the voice features, making it possible to automatically monitor conversations with high accuracy and providing useful reference information for administrators when making final decisions.

[0087] Next, in detailed analysis mode, the detailed analysis determination unit 61 determines whether the confidence level of the analysis is high or low (Step S19: Analysis Confidence Level Determination Step). If the result of this determination is that the confidence level of the analysis is low (Step S19: NO), the detailed analysis determination unit 61 transmits the conversational audio to the speech analysis engine 13 to obtain the analysis result, and also transmits this analysis result as additional information to the conversation analysis model 12 to perform a detailed analysis. Then, as an analysis result of the detailed analysis, the analysis result acquisition unit 60 obtains the final conversation quality score from the conversation analysis model 12 (Step S20: Detailed Analysis Result Acquisition Step).

[0088] This allows for more detailed analysis results when the conversation audio is unclear or noisy, resulting in low transcription quality, or when the context is ambiguous or the scoring is borderline, causing hesitation in the conversation analysis model 12's analysis and judgment. As a result, conversations are automatically monitored with higher accuracy, and more useful reference information is provided to administrators. On the other hand, if the confidence level of the analysis is high (step S19: YES), the process in step S20 is skipped, and the analysis results obtained in step S18 are used as is.

[0089] Next, the recording data storage unit 62 determines whether or not there is an abnormality in the quality of the conversation (step S21: conversation quality determination step). If the determination results in an abnormality in the quality of the conversation (step S21: YES), the recording data storage unit 62 acquires a predetermined amount of conversation audio from the audio buffer 42 and saves it to the recording data storage unit 43 (step S22: conversation audio recording step). This ensures that the conversation in question is saved as audio evidence, and that the overall picture of the conversation can be analyzed and confirmed later. On the other hand, if there is no abnormality in the quality of the conversation (step S21: NO), no recording of the conversation audio is performed, and the process proceeds to step S23.

[0090] Next, the analysis result transmission unit 63 transmits various analysis results (conversation quality score, confidence level of the analysis results, volume score, conversation text, confidence level of the conversation text, etc.) to the management server 14 (step S23: analysis result transmission step). As a result, the various analysis results are reflected in the web management dashboard in real time, and if escalation is necessary, an alert is notified to the administrator, enabling a quick and appropriate response to abnormal conversations.

[0091] Finally, the conversation monitoring program 1a determines whether or not there is an instruction to terminate monitoring (step S24: monitoring termination determination step). If there is no termination instruction (step S24: NO), it waits until the next analysis cycle (step S25), then returns to step S2 and repeats the process described above. On the other hand, if there is a termination instruction (step S24: YES), this process is terminated.

[0092] [4] Effects The conversation monitoring program 1a, conversation monitoring device 1, conversation monitoring system 100, and conversation monitoring method according to the present invention described above will produce the following effects. 1. By analyzing not only the text of conversations but also the features of the audio conversations, false positives in automated conversation monitoring can be reduced, enabling highly accurate automated monitoring of conversations. Furthermore, conversations with abnormal quality are automatically recorded retrospectively, ensuring that evidence of incidents is reliably and objectively preserved. This can then be used for future troubleshooting, compliance audits, and employee training. 2. By performing multi-stage filtering based on the features of the conversational audio, unnecessary analysis of audio that is prone to false positives with a single feature is suppressed, thereby improving analysis accuracy. 3. By adjusting the waiting time based on the results of the previous analysis, even normal, peaceful conversations can be analyzed without omission while keeping analysis costs down. 4. By providing physical characteristics of conversational audio as event tags, it is possible to define specific audio changes at the time the conversation was spoken, allowing the conversation analysis model 12 to understand what was happening, including the context. 5. By excluding conversational audio that is presumed to be ambient noise from the analysis, and analyzing only the audio that needs to be analyzed, API call costs and communication load can be reduced. 6. By excluding fragmented conversational text derived from ambient noise from the analysis, API call costs and communication load can be reduced. 7. It can detect similar duplicate statements and effectively eliminate slightly different repeated statements or re-analysis / duplicate recordings of the same conversation. 8. If the transcription quality is low, or if there is uncertainty in the analysis and judgment of the conversation analysis model 12, more detailed analysis results can be obtained. 9. It can immediately detect emergencies such as sudden loud noises and save the conversation audio as evidence with the highest priority.

[0093] Furthermore, the conversation monitoring program 1a, conversation monitoring device 1, conversation monitoring system 100, and conversation monitoring method according to the present invention are not limited to the embodiments described above and can be modified as appropriate. [Explanation of symbols]

[0094] 1. Conversation monitoring device 1a Conversation monitoring program 11. Speech Recognition Engine 12 Conversation Analysis Models 13. Voice Analysis Engine 14 Management Server 100 Conversation Monitoring System 2. Means of communication 3. Voice input means 4 Memory means 41 Program Storage Unit 42 Audio buffer 43 Recording data storage unit 44 Analysis result storage unit 5. Calculation processing means 51 Conversation Voice Recording Unit 52 Feature Analysis Section 53 Feature Analysis Unit 54 Standby time adjustment section 55 Unwanted sound determination section 56 Conversation Text Acquisition Unit 57 Text content determination unit 58 Similar Text Determination Unit 59 Tag Embedding Text Creation Section 60 Analysis result acquisition section 61 Detailed analysis judgment section 62 Recording data storage section 63 Analysis Result Transmission Unit

Claims

1. A conversation monitoring program for monitoring conversations, A feature scoring unit acquires a predetermined amount of conversational audio from an audio buffer that temporarily records the conversational audio of the aforementioned conversation, and scores the feature quantities of the conversational audio. A feature determination unit that determines whether or not to continue analyzing the conversational audio based on the aforementioned feature quantities, A conversation text acquisition unit that acquires conversation text transcribed from the conversation audio using a speech recognition engine, A tag-embedded text creation unit creates tag-embedded text by embedding event tags indicating the voice changes that occurred in the aforementioned conversation audio at the locations where the voice changes occurred in the aforementioned conversation text. An analysis result acquisition unit that acquires analysis results regarding the quality of the conversation, which are analyzed based on the tag-embedded text by the conversation analysis model, Based on the analysis results, if there is an abnormality in the quality of the conversation, the recording data storage unit acquires a predetermined amount of conversation audio from the audio buffer and stores it in the recording data storage unit. Analysis result transmission unit that transmits the aforementioned analysis results to the management server A conversation monitoring program that makes a computer function as such.

2. As the aforementioned feature quantities, A volume score is calculated based on sudden fluctuations in volume, sudden occurrences of loud noises, prolonged silences, and sustained occurrences of loud noises. The pitch score is calculated based on the magnitude of pitch fluctuations and sudden changes in pitch, The audio score obtained by adding the volume score and the pitch score, It has, The feature determination unit, The conversation monitoring program according to claim 1, which continues the analysis of the conversational audio if any of the voice score, volume score, and pitch score are above a predetermined threshold.

3. The conversation monitoring program according to claim 2, wherein if the feature determination unit determines that any one of the voice score, volume score, and pitch score is below a predetermined threshold, the computer functions as a waiting time adjustment unit that adjusts the waiting time until analysis is started based on the previous analysis results.

4. The conversation monitoring program according to claim 1, wherein the event tag includes at least one of the following: a volume increase tag indicating a sudden increase in volume, a sustained loud volume tag indicating that loud volume was sustained, a pitch increase tag indicating a sudden increase in pitch, a silence tag indicating a silent period, and a tempo acceleration tag indicating a sudden increase in conversation speed.

5. The computer is configured to function as a non-essential sound determination unit, which determines whether the aforementioned conversational audio is audio that does not require analysis. The conversation monitoring program according to claim 1, wherein the unwanted sound determination unit does not continue the analysis if the conversation voice satisfies any one of the following conditions (1) to (3), and continues the analysis if the conversation voice does not satisfy any of the following conditions (1) to (3); (1) The average volume of the conversational voice is below the silence threshold, (2) The duration of the spoken speech is less than the threshold for determining if it is spoken. (3) The human voice ratio, which is the time ratio of human voices included in the conversational audio, is less than the human voice detection threshold.

6. The conversation monitoring program according to claim 1, wherein a computer functions as a text content determination unit that determines whether the content of the conversation text is normal or not based on the number of words, the number of characters, and the percentage of Japanese characters contained in the conversation text.

7. The conversation monitoring program according to claim 1, wherein the computer functions as a similar text determination unit, determining whether the conversation text is similar to conversation texts that have been analyzed in the past, and not continuing the analysis of similar conversation texts, but continuing the analysis only of conversation texts that are not similar.

8. The conversation monitoring program according to claim 1, wherein the computer functions as a detailed analysis determination unit that determines whether or not to have the conversation analysis model perform a detailed analysis regarding the quality of the conversation, based on the reliability of the conversation text and the confidence level of the analysis results.

9. The conversation monitoring program according to claim 2, wherein if the voice score is above an emergency threshold and the human voice ratio, which is the time ratio of human voices included in the conversation audio, is above a human voice detection threshold, the recording data storage unit immediately starts recording the conversation audio and continues analyzing the conversation audio.

10. A conversation monitoring device for monitoring conversations, A feature scoring unit acquires a predetermined amount of conversational audio from an audio buffer that temporarily records the conversational audio of the aforementioned conversation, and scores the feature quantities of the conversational audio. A feature determination unit that determines whether or not to continue analyzing the conversational audio based on the aforementioned feature quantities, A conversation text acquisition unit that acquires conversation text transcribed from the conversation audio using a speech recognition engine, A tag-embedded text creation unit creates tag-embedded text by embedding event tags indicating the voice changes that occurred in the aforementioned conversation audio at the locations where the voice changes occurred in the aforementioned conversation text. An analysis result acquisition unit that acquires analysis results regarding the quality of the conversation, which are analyzed based on the tag-embedded text by the conversation analysis model, Based on the analysis results, if there is an abnormality in the quality of the conversation, the recording data storage unit acquires a predetermined amount of conversation audio from the audio buffer and stores it in the recording data storage unit. An analysis result transmission unit that transmits the aforementioned analysis results to a management server, A conversation monitoring device having the following features.

11. A conversation monitoring device according to claim 10, A speech recognition engine that transcribes the aforementioned conversation text, A conversation analysis model analyzes the quality of the conversational audio based on the tagged text, A management server that stores the aforementioned analysis results, A conversation monitoring system having the following features.

12. A conversation monitoring method for monitoring conversations, Computers A feature scoring step involves obtaining a predetermined amount of conversational audio from an audio buffer that temporarily records the conversational audio of the aforementioned conversation, and scoring the feature quantities of the conversational audio. A feature determination step that determines whether or not to continue analyzing the conversational audio based on the aforementioned features, A conversation text acquisition step involves obtaining conversation text transcribed from the conversation audio using a speech recognition engine, A tag-embedded text creation step, which creates tag-embedded text by embedding event tags indicating the voice changes that occurred in the aforementioned conversation audio at the locations where the voice changes occurred in the aforementioned conversation text, An analysis result acquisition step is to obtain analysis results regarding the quality of the conversation, which are analyzed based on the tag-embedded text by the conversation analysis model, Based on the analysis results, if there is an abnormality in the quality of the conversation, the recording data storage step involves acquiring a predetermined amount of conversation audio from the audio buffer and saving it to the recording data storage unit. The analysis results transmission step involves sending the aforementioned analysis results to a management server, A method for monitoring conversations.