Method and Apparatus for Context-Aware Audio Capture and Analysis

The system addresses fragmented voice communication workflows by integrating real-time speech-to-text conversion, hybrid summarization, and secure data management, enhancing accuracy and efficiency in capturing and managing voice data across diverse environments.

KR102993813B1Active Publication Date: 2026-07-21김인기
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
김인기
Filing Date
2025-07-02
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing voice communication systems lack integration and efficiency in capturing, transcribing, summarizing, and verifying voice data, leading to fragmented workflows and reduced accuracy due to segmental application of preprocessing techniques, which fail to adapt to acoustic changes during calls or conversations.

Method used

A system that automatically captures voice signals, performs real-time speech-to-text conversion, generates hybrid summaries, identifies key remarks, and synchronizes results on a central server, incorporating multi-language and speaker separation, noise reduction, and semantic analysis to enhance communication quality and work efficiency.

Benefits of technology

The system provides continuous, accurate transcription, summarization, and verification of voice data across various environments, reducing manual intervention and ensuring secure, reliable communication records with enhanced user convenience and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1020250088356
    Figure 1020250088356
Patent Text Reader

Abstract

The present invention relates to a method and apparatus for automatically capturing voice signals without separate user operation in various voice communication environments, such as conferences, calls, and general conversations, and for performing integrated tasks including real-time or post-hoc transcription (STT), hybrid summarization, identification of key remarks, comparative verification with pre-uploaded data, and synchronous storage on a remote server. The key steps consist of voice capture, STT, summarization, remark identification, verification of missing items, and synchronous storage, and each step ensures accuracy and reliability by including advanced preprocessing and analysis techniques such as multi-language support, speaker separation, noise removal, and OCR processing. The present invention significantly reduces the burden of communication recording and analysis tasks through an automated workflow, and resolves issues of missing information and work delays by providing real-time feedback and notification of missing items. In addition, it supports consistent data management among multiple terminals through central server synchronization and a permission-based security module, and simultaneously improves organizational collaboration efficiency and information reliability by enabling the secure storage and sharing of sensitive information.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a situation-aware audio recording and analysis system. Background Technology

[0002] In various voice communication situations, such as meetings, calls, one-on-one conversations, and group discussions, the recording and organization of discussion content was a major cause of the omission of important information and work inefficiency. Existing technologies relied primarily on simple recording functions or separate enterprise services; consequently, functions for capturing, converting, analyzing, and verifying voice data were fragmented, resulting in a lack of continuity in the overall workflow and limiting user convenience and efficiency due to the difficulty of providing real-time feedback.

[0003] Furthermore, as preprocessing techniques such as speaker separation, multilingual support, and noise reduction were applied segmentedly as individual modules during the speech-to-text conversion process, it was difficult to adapt to various acoustic environmental changes occurring during calls or conversations, resulting in a problem of reduced transcription and summarization accuracy. The problem to be solved

[0004] The present invention aims to provide a system capable of automatically capturing voice signals and performing integrated transcription, summarization, analysis, and verification in all voice communication situations, ranging from conferences to calls and general conversations, without additional user operation.

[0005] In particular, we aim to improve communication quality and work efficiency by identifying and notifying missing items throughout the conversation flow in real time through advanced speech preprocessing including multi-language and speaker separation functions, a hybrid summarization technique combining extractive and generative summarization, and an automatic verification function with pre-uploaded data.

[0006] In addition, the task is to maintain various communication records both inside and outside the organization consistently and securely by synchronizing and storing all analysis results on a central server and safely managing sensitive information through permission-based access control and data encryption. means of solving the problem

[0007] As one means according to the present invention, the following is disclosed.

[0008] A voice communication support method is disclosed as one means, comprising: (S1) a step of automatically capturing and recording a voice signal occurring in a voice communication environment; (S2) a step of automatically converting the captured voice data to speech-to-text (STT) in real-time or afterward to generate or store text containing timestamps for each speaker; (S3) a step of generating a conversation summary by performing a hybrid summary combining an extractive summary and a generative summary based on the generated text; (S4) a step of identifying key remarks by combining a keyword weighting calculation technique and a semantic analysis technique in the generated summary; (S5) a step of identifying and listing missing discussion items by performing natural language processing-based semantic matching when pre-uploaded reference material is confirmed; and (S6) a step of synchronizing and storing the identified summary, key remarks, and missing items on a remote server.

[0009] (S2) A voice communication support method is disclosed as one means, characterized by performing a multi-language support speech recognition and speaker separation function and a noise removal preprocessing technique in step (S2).

[0010] (S3) A voice communication support method is disclosed as one means, characterized by including a semantic-based control function that automatically recognizes the purpose of the conversation, the topic, and keywords to control the summary range and level of detail.

[0011] (S4) A voice communication support method is disclosed as one means, characterized by selecting key statements by combining a keyword weighting calculation technique, an embedding-based similarity analysis, and a user filtering function.

[0012] A voice communication support method is disclosed as one means, characterized by processing various types of reference materials such as documents, tables, and images in step (S5) using OCR and structured data mapping, calculating semantic matching based on natural language processing to identify missing items, and synchronously storing them together with metadata in step (S6).

[0013] A voice communication support device comprising a processor and memory configured to execute the above method, wherein the processor comprises a capture module that performs automatic voice capture,

[0014] A voice communication support device is disclosed, comprising an STT module for performing speech-to-text conversion, a summary module for performing summarization, an identification module for performing speech identification, a verification module for performing data verification, and a synchronization module for performing synchronization.

[0015] In addition, a voice communication support device is disclosed, characterized by further including a schedule control module that automatically schedules and starts a capture when a conversation begins by linking with an external schedule management service and automatically stops it when the conversation ends.

[0016] Additionally, a voice communication support device is disclosed, characterized by further including a tag management module that inputs tags via a user's touch or voice command and adjusts summary and notification priorities based on the tags.

[0017] In addition, a voice communication support device is disclosed, characterized by further including: a log analysis module that collects and analyzes user behavior logs before, during, and after conversation to generate engagement and concentration indicators; and a dashboard module that visualizes the generated indicators.

[0018] For example, if a user's 'focus metric' is consistently measured as low during a meeting, the dashboard can display a real-time warning icon and provide feedback by showing a graph of the user's focus trend after the meeting ends.

[0019] In addition, a voice communication support device is disclosed, characterized by including an authentication module and an encryption module for secure management and access control of stored data, and being configured to perform data decryption and retrieval according to user authority.

[0020] Specific details of other embodiments are included in the detailed description. Effects of the invention

[0021] According to the present invention, automatically captured voice data can be transcribed in real time without separate operation, and hybrid summarization, identification of key remarks, and verification of missing data items can be performed continuously, thereby significantly reducing the burden of recording and analysis tasks in all voice communication situations, such as meetings, calls, and conversations.

[0022] In addition, it provides summary and analysis results with high accuracy and comprehensibility through an STT module integrating speaker separation, multilingual support, and noise reduction technologies that consider various acoustic characteristics occurring in calls or general conversations, and a semantic control-based summarization module optimized for conversational purposes and context.

[0023] Furthermore, by securely storing, sharing, and managing records of meetings, calls, and conversations both inside and outside the organization through central server synchronization and permission-based security modules, it has the effect of simultaneously enhancing personal information protection and information reliability. Specific details for implementing the invention

[0024] The advantages and features of the invention and the methods for achieving them will become clear by referring to the embodiments described in detail below. However, the invention is not limited to the embodiments described herein and may be embodied in other forms. Rather, the embodiments introduced herein are provided to ensure that the disclosed content is thorough and complete and to ensure that the spirit of the invention is sufficiently conveyed to those skilled in the art.

[0025] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the invention. The singular expression includes the plural expression unless the context clearly indicates otherwise.

[0026] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.

[0027] The present invention relates to a voice communication support device or a method thereof. For example, the device described herein may be a device used by a user, which may be a wired or wireless device, and furthermore, may be a program or app installed on such a device. In addition, such a program or app, that is, such a device, may be executed by a processor (including a CPU, GPU, or a central server, etc.).

[0028] The present invention will be described in detail below.

[0029] For example, the present invention discloses a voice communication support method comprising: (S1) automatically capturing and recording a voice signal occurring in a voice communication environment; (S2) automatically performing speech-to-text conversion (STT) on the captured voice data in real-time or afterward to generate or store text containing timestamps for each speaker; (S3) generating a conversation summary by performing a hybrid summary combining an extractive summary and a generative summary based on the generated text; (S4) identifying key remarks by combining a keyword weighting calculation technique and a semantic analysis technique in the generated summary; (S5) identifying and listing missing discussion items by performing natural language processing-based semantic matching when pre-uploaded reference material is confirmed; and (S6) synchronizing and storing the identified summary, key remarks, and missing items on a remote server.

[0030] For example, the voice communication environment in (S1) includes all environments where voice signals are generated, such as conference rooms, office calls, smartphone calls, and voice chat apps, and can refer to situations where variations in acoustic conditions, such as ambient noise or echo, are allowed.

[0031] In addition, automatically capturing voice signals here includes recording raw audio streams such as WAV and PCM without loss using a microphone input interface or an OS-level audio API, and can be executed automatically in the background.

[0032] For example, it can uniformly capture the speech of five attendees using a laptop's built-in microphone in a conference room, record both sides of the voice (oneself and the other party) during a smartphone call and secure an isolated audio channel using a call application audio endpoint, or, in the case of talking to oneself, automatically detect when the user starts speaking without a button click, such as with a voice memo app, and start recording.

[0033] In addition, voice capture can support not only single-channel but also multi-channel inputs (e.g., stereo, multi-microphone arrays), which enables spatial location-based speaker identification and improved speech intelligibility. In this case, input intensity for each channel can be analyzed in real time to apply directional microphone functions or beamforming, allowing for the preferential capture of voice from a specific direction. For example, in an environment equipped with 360-degree conference microphones, channel separation can be performed based on the speaker's location, and audio for each speaker can be stored separately to improve the accuracy of subsequent STT and analysis.

[0034] (S2) STT is a process of converting 16kHZ sampled speech into text using a deep learning-based speech recognition engine, and timestamps can be recorded on a sentence-based or word-based basis.

[0035] Furthermore, in one embodiment of the present invention, an audio file automatically generated at the end of recording is linked to an STT module to continuously perform text transcription, and the transcription results are transmitted to a summary module to immediately generate a summary result, which is then automatically saved within the same folder. At this time, the saved filename is automatically named based on the date and time, and the result is configured so that the user can subsequently view it or output it in the form of a summarized PDF. Such a batch automation process of recording, transcription, summarization, and saving is effective in minimizing manual user intervention and supporting the structuring of repetitive communication.

[0036] For instance, in the case of a meeting on a specific topic, the recognition accuracy of technical terms or commonly used expressions can be achieved above a certain level based on a database that contains terms relevant to that topic, either embedded in a dictionary or retrieving them from an external source. For instance, in a hospital meeting, terms such as "myocardial infarction," "serum lipids," and "echo-Doppler" may appear; since general STT systems may misinterpret these terms, the recognition rate of these terms can be enhanced by utilizing a database—specifically a medical terminology dictionary—that is either embedded in a dictionary related to the field or retrieving it from an external source.

[0037] In addition, in the case of foreign languages, the system switches to a language model to transcribe the foreign language either in its original form or in a similar manner. For instance, if a meeting or conversation takes place in a foreign language, the system automatically switches to and applies the appropriate language model to accurately recognize or transcribe the foreign language. If the meeting is conducted in English or Japanese, the system automatically switches to an English or Japanese STT model to convert the text into the original language or, if necessary, to enable automatic translation.

[0038] Furthermore, post-processing reflecting the context of the conversation can be applied during the STT process, enabling the accurate identification of homonyms or the supplementation of missing subjects and objects. For example, a natural language post-processing algorithm can be integrated to clearly supplement the sentence by inferring what "that" refers to based on the preceding context in a sentence such as "Let's put that off until tomorrow." In addition, if slang, prohibited words, or sensitive words are included, automatic filtering or masking is performed according to user settings to support the protection of personal information and the maintenance of ethical content.

[0039] In (S3), the hybrid summary is a pipeline configuration for performing extractive summarization (selecting key sentences) and generative summarization (natural language reconstruction) in a sequential or cross-sectional manner, and can utilize pre-trained models such as BART and PEGASUS.

[0040] For instance, in a recording of an in-house meeting, the sentence "budget approval needed" can be extracted using an extract summary and restructured into a form like "The team decided to request budget approval by the end of this month" using a generate summary; similarly, in everyday conversations, the sentence "What should we have for lunch?" can be given low weight, while "Sharing data analysis results at the Monday meeting" can be treated as the top item in the summary.

[0041] In addition, the hybrid summarization process may include an alignment algorithm that recognizes the narrative structure (e.g., introduction-development-conclusion) and arranges summary sentences in a logical order. This enables the construction of summaries based on logical flow rather than chronological order, thereby enhancing the reader's understanding.

[0042] In addition, it supports changing summary styles according to user settings, allowing for customization into concise summaries for meeting minutes, polite summaries for emails, and action-oriented summaries for work instructions.

[0043] Finally, for the summarized results, a quality verification routine may be added to automatically evaluate quality using quantitative indicators such as BERTScore and ROUGE, and to induce user feedback if the result falls below a certain standard.

[0044] In the present invention, for summary quality evaluation and history management, meta-information such as the time of summary creation, the type of model used, and the length of the input transcription text can be stored for each summary result. This meta-information can be used as a reference point when regenerating summary results in the future and can be utilized to compare or manage multiple summary versions of the same transcription. For example, if a specific summary result is modified based on user feedback, timestamp and version information are recorded together to enable comparative analysis with the existing summary, thereby implementing a traceable flow of summary quality improvement.

[0045] Identifying key statements in (S4) involves combining text embeddings (e.g., BERT) with TF-based weights to calculate statement importance and including logic that considers sentences above a threshold as key statements.

[0046] For example, in the case of customer service calls, keywords such as "refund request" are filtered first, and high-severity remarks are highlighted separately in red, or sentences resulting from decisions during meetings, such as "design deadline extension," are automatically tagged as key remarks.

[0047] In addition, by performing intent classification based on the purpose of the statement (question, request, decision, information transmission, etc.), the importance of the same word can be reflected differently depending on the context.

[0048] For example, "The deadline is today" is of low importance if it is merely a message of information, but "Please extend the deadline" is considered a request or a decision and is identified as the key point.

[0049] In addition, a user-centric feedback loop can be implemented that incorporates sentences marked as important or repeatedly checked by the user into the learning process to assign weights to similar statements in the future.

[0050] Furthermore, by combining emotional signals (tension, excitement, anger, etc.) extracted from voice signals to identify emotionally sensitive remarks as auxiliary indicators, the precision of key remark selection can be improved.

[0051] In (S5), the comparison of reference materials refers to structuring the uploaded materials (PPT, PDF, DOCX, etc.) through OCR and DOM parsing, and determining whether a match exists by calculating the cosine similarity between keywords and the transcribed text for each slide and paragraph.

[0052] For instance, if the "SNS campaign schedule" is not mentioned company-wide in marketing strategy presentation materials, information such as the document number and title is listed as a missing item; similarly, if "safety guidelines" are not discussed in training materials, a notification is provided to verify emergency response procedures.

[0053] In addition, to improve the matching accuracy between data and utterances, omissions can be determined at the semantic level by combining simple keyword matching with the extraction of named entities (NER) within sentences and context-based embedding similarity comparisons.

[0054] For example, even if the phrase "SNS promotion" does not appear, if the statement "Instagram ad timing adjustment" exists, it can be determined to be related to the slide.

[0055] In addition, through UI interaction functions that allow users to manually verify, confirm, or exclude detected missing items, both the reliability and flexibility of the system's judgment can be secured.

[0056] In addition, text included in unstructured data such as tables and images is processed in parallel, and in cases where OCR processing is required, such as scanned images included in PDFs, pre-labeled layout-based structural recognition technology (Document Layout Analysis) can be combined to maintain accuracy.

[0057] Synchronous storage in (S6) summarizes the data in JSON format and sends the metadata to a central server via a RESTful API, and includes retry logic after local caching to account for response delays and offline mode.

[0058] For example, in a central server integration method, if a normal response (200 OK) is not received after an API call, the summary result may be stored in a local cache and resent at 5-minute intervals, or in a mobile offline method, when Wi-Fi is reconnected, the summary and metadata stored in all queues may be automatically synchronized.

[0059] In addition, during the synchronization process, version hash or timestamp-based version management strategies may be applied to prevent duplicate uploads or update conflicts of the same summary data.

[0060] For example, if summaries generated on different terminals conflict within the same meeting session, only the most recently updated item may be retained, or a conflict resolution interface may be provided that the user can select.

[0061] In addition, the stored data is transmitted via an HTTPS / TLS-based encrypted channel, and on the server side, user authentication token-based access control is implemented to prevent the leakage of sensitive information.

[0062] Meanwhile, in the event of a prolonged offline environment, policy-based transmission control is also possible to minimize unnecessary transmissions by managing the expiration time of data in the queue, or to select transmission targets based on data priority.

[0063] In a distributed server environment, synchronization efficiency can be improved through load balancing methods that consider latency between the local and cloud or location-based server allocation.

[0064] In addition, a voice communication support method is disclosed, characterized in that, in step (S2), it includes a multi-language support speech recognition and speaker separation function and a noise removal preprocessing technique.

[0065] Multi-language speech recognition applies a Language ID module to input voice streams to classify major languages, such as Korean, English, and Chinese, in real time. For example, if English and Korean code switching occurs during a meeting, each segment is routed to the corresponding language model to minimize recognition errors. Subsequently, precise post-mortem analysis enables more accurate transcription and summarization.

[0066] In other words, as the meeting progresses, the language of short segments (2–3 second intervals) is identified in real-time using a Language ID model, and routed to the appropriate STT engine. Subsequently, in the post-analysis stage, based on the timestamps and language tags generated during real-time recognition, precision models such as Whisper large-v3 for English segments and the Clova call Model for Korean segments are reapplied, and code-switching sections are re-recognized using a dual-pass decoding technique.

[0067] In addition, Voice Activity Detection (VAD) and clustering-based algorithms (e.g., x-vectors, UEM) are used to identify speakers for each voice segment. For example, in a four-person meeting, feature vectors for each speaker A, B, C, and D are learned, and the real-time stream is separated into segments and linked with timestamps. Even if two of the four meeting participants have similar voices, the algorithm can accurately separate and display the remarks of "Assistant Manager Kim" and "Manager Park" by combining each speaker's speaking speed, intonation, and pause patterns.

[0068] In addition to x-vector and UEM-based clustering, speaker segmentation can be performed in parallel with high-precision speaker embedding techniques based on ECAPA-TDNN or anchor-based classification techniques utilizing pre-registered voice samples. For video conferencing, a multimodal speaker recognition method using video-based lip-to-speech synchronization can also be applied. Furthermore, prosody-based segmentation or speaker role mapping based on speech role recognition can serve as useful complementary measures.

[0069] Meanwhile, to improve speaker separation accuracy, the embodiments of the present invention allow the number of participants in a conversation to be entered by the user prior to recording, and the initial conditions for speaker separation can be optimized by transmitting this value as an analysis parameter. For example, in a structure where minimum and maximum speaker count ranges can be set in a Pyannote-based model, if the user enters '3 people,' this can be immediately reflected as an analysis variable to constrain the number of clusters and improve initial classification accuracy. This method is a hybrid approach that compensates for the uncertainty of automatic speaker separation while minimizing the burden on the user, thereby increasing the reliability of STT and speech analysis.

[0070] Noise removal includes spectral gating and Wiener filters, while multilingual recognition refers to a method of selecting a language-specific model or using a multilingual model via a Language ID module. Furthermore, background noise (traffic, construction sites, etc.) is removed by combining a deep learning-based noise suppression model (Demucs Instrumental). For example, when recording a conversation in an airport lounge, the sound of an aircraft taking off and the conversational voice can be separated to cleanly extract only the voice.

[0071] In addition, a voice communication support method is disclosed, characterized by including a semantic-based control function that automatically recognizes the purpose of conversation, topic, and keywords in step (S3) and controls the summary range and level of detail.

[0072] The summary control function includes hyperparameter settings that dynamically adjust the summary length and the weight of key topics by reflecting user-specified keywords or topic tags.

[0073] For example, the conversation purpose analyzes the meeting name, calendar metadata, and user-specified tags (e.g., "Project Kick-off", "Customer Consultation") to automatically set the summary perspective (decision, schedule, task, etc.).

[0074] Topic and keyword extraction combines TF-IDF (Word Frequency-Inverse Document Frequency) and RAKE (Rapid Automatic Keyword Extraction) to extract major keywords, and performs topic filtering of the summary based on these. For example, if "budget approval" is identified as a major keyword, the summary narrows its scope to prioritize the inclusion of budget-related statements.

[0075] Controlling hyperparameters allows the user to specify or automatically adjust the summary length (min / max number of sentences) and level of detail. For example, it is set to recommend a maximum of 10 sentences for important strategy meetings and 3 sentences for 1:1 calls.

[0076] (S4) A voice communication support method is disclosed, characterized by selecting key statements by combining a keyword weighting calculation technique, an embedding-based similarity analysis, and a user filtering function.

[0077] The keyword weighting calculation technique assigns additional weights to important keywords pre-specified in a domain dictionary (e.g., medical terminology dictionary) based on TF-IDF values. However, the domain dictionary may be pre-injected or loaded from an external source.

[0078] Embedding-based similarity analysis embeds each sentence using Sentence-BERT and adjusts for importance by calculating cosine similarity with the representative sentence of the entire conversation (e.g., the first remark). For instance, in a startup pitch, remarks with high similarity to sentences centered on keywords such as "investment" and "business model" are considered key.

[0079] Custom filtering refers to prohibited words (credit card numbers, personal information, etc.), filtering of remarks of interest, and highlighted words (urgent, important), which enables the automatic detection of security and policy compliance remarks within the organization, or the inclusion or exclusion of specific remarks exceptionally or preferentially.

[0080] A voice communication support method is disclosed, characterized in that, in step (S5), various types of reference materials such as documents, tables, and images are processed by OCR and structured data mapping, a semantic match based on natural language processing is calculated to identify missing items, and in step (S6), they are stored in synchronization with metadata.

[0081] Using the Tesseract OCR engine, titles, table items, image captions, etc., of various reference material types are extracted as text, and chapter / slide numbers or element IDs are stored as metadata. Next, sentence-specific entities (e.g., product name, date) of the transcribed text are extracted using SpaCy NER, linked with document element metadata, and marked as missing if the cosine similarity between the transcribed sentence and the document element is below a specific threshold (e.g., 0.6).

[0082] For example, if the keyword "promotion schedule" in the slides of existing reference materials is not present in the text, it indicates that it is missing.

[0083] For example, the present invention discloses a voice communication support device comprising a processor and a memory configured to execute the above method, wherein the processor comprises a capture module for performing automatic voice capture, an STT module for performing speech-to-text conversion, a summary module for performing summary, an identification module for performing speech identification, a verification module for performing data verification, and a synchronization module for performing synchronization.

[0084] The above capture module includes a smartphone microphone driver and an audio service linkage component, and performs an automatic gain adjustment (AGC) function upon noise detection.

[0085] For example, Automatic Gain Control (AGC) dynamically adjusts the microphone gain by monitoring ambient noise levels in real time. For instance, since noise levels are low inside an office but high in an external environment, it automatically scales the gain from 20dB to 60dB.

[0086] In addition, FDN (Finite Impulse Response) filters are applied at the H / W and S / W levels to suppress echoes that feed back sound from the speaker to the microphone, and various audio codecs such as WAV, PCM, and AAC are supported, and the sampling rate (8kHz~48kHz) and bit depth (16 / 24-bit) are automatically detected.

[0087] The above STT module includes a pipeline that defines the preprocessing sequence (noise removal → language identification → speech recognition) and a GPU-accelerated inference engine.

[0088] Furthermore, the present invention is designed to be compatible with commercial speech recognition APIs such as Google STT and may include a function for automatically converting voice file formats considering device compatibility and text recognition rates. For example, the method involves automatically processing an original voice file recorded in an Android device environment to be converted and saved in a format (such as AAC or FLAC) that is optimally recognized by Google STT, and then transmitted to the corresponding API. This structure minimizes variations in speech recognition rates across platforms and is advantageous for flexible STT engine replacement and quality management.

[0089] For example, the preprocessing here includes removing noise, identifying the language, and normalizing the volume. Next, batch processing is performed using 1-second buffers in parallel on a GPU or NPU with a maximum latency of 200ms, and a model optimized for multilingual performance is used using a Conformer or Whisper-based model—in this case, the error rate (WER) is within 5% when using a Korean-English cross-model. Then, postprocessing is carried out to perform case restoration, punctuation insertion, and preprocessing based on a user dictionary.

[0090] For example, if a user says "next Tuesday at 3 PM meeting" in a situation with subway noise, the background noise is removed and the unclear pronunciation of 'Tuesday' is restored to provide a result that accurately transcribes it as "next Tuesday at 3 PM meeting".

[0091] The above summary module is implemented as an integrated summary service that includes a generation model call and clean-up logic after an initial sentence importance evaluation.

[0092] In particular, one embodiment of the present invention may include a structure that links with a GPT-based external summarization API to transmit transcribed text to an external model and receive and store the summary result. In the future, if the summary result is modified or new information is reflected, a function is also provided to synchronize the updated summary result by calling the summarization API again according to the modified transcription result. Through this, a bidirectional feedback structure is implemented in which changes to the text are reflected in the summary result in real time, and switching to various summarization models is flexibly supported according to changes in the GPT API or user requirements.

[0093] For example, the TextRank algorithm is used to extract the top 30% of important sentences (sentence importance evaluation), generative supplementation is performed using generative artificial intelligence (generative supplementation), and BERTScore is calculated to verify semantic consistency between summary sentences, and only outputs are those with a score of 0.8 or higher (consistency check). For example, if 8 sentences are input, they can be summarized into 3 sentences through this process.

[0094] The above identification module provides an embedding model, an algorithm for calculating weight-based importance scores, and a highlight display API to the user interface.

[0095] For example, 768-dimensional vectorization can be performed using Sentence-BERT, and the final importance can be scored by considering TF-IDF, cosine similarity, and domain weights, and then a threshold can be determined and linked to the UI.

[0096] The above verification module includes a document parser (OCR + text extractor) and a semantic similarity calculator (Transformer-based), and returns the verification results in the form of metadata.

[0097] For example, a JSON schema is generated using Tesseract OCR, and a hybrid score is calculated by combining the cosine similarity between sentences and elements and the Zakart index. Accordingly, for example, if the content of 2 slides is missing out of 10 PPT slides, a popup is automatically displayed in the UI.

[0098] The above synchronization module includes a background synchronization scheduler, a conflict resolution policy (version management), and encrypted transmission (HTTPS / TLS).

[0099] In addition, a voice communication support device is disclosed, characterized by further including a schedule control module that links with an external schedule management service to automatically schedule and start a capture when a conversation starts and automatically stop it when the conversation ends.

[0100] In addition, a voice communication support device is disclosed, characterized by further including a tag management module that inputs tags via a user's touch or voice command and adjusts summary and notification priorities based on the tags.

[0101] For example, if a user gives a voice command saying, "Make sure to include this part in the summary," the tag management module automatically assigns the #Key tag to the remark at that moment, and the summary module subsequently prioritizes reflecting remarks with that tag. This helps ensure that important content or action instructions during a meeting are not buried.

[0102] In addition, a voice communication support device is disclosed, characterized by further including: a log analysis module that collects and analyzes user behavior logs before, during, and after conversation to generate engagement and concentration indicators; and a dashboard module that visualizes the generated indicators.

[0103] For example, the log analysis module generates and / or analyzes detailed metrics such as the ratio of speaking volume, the frequency of speaker switching, and the ratio of silence intervals. Here, the ratio of speaking volume is the ratio of speaking time per user or relative to the total, and can display each participant's speaking volume as a percentage. Additionally, the frequency of speaker switching is the number of times the speaker changes during a conversation, indicating the degree of active interaction among team members, while the ratio of silence intervals is the proportion of time when voice is paused during a conversation, allowing for the evaluation of discussion focus and the flow of the debate.

[0104] The log analysis module can generate various detailed metrics based on user speech behavior in voice communications, such as meetings or calls. For example, Ratio of speaking time Speech refers to the proportion of the total conversation time that each participant speaks, and it is useful for identifying any bias in speaking among team members. For instance, in a 20-minute meeting, if User A speaks for 10 minutes and Users B and C each speak for 5 minutes, User A accounts for 50% of the total speaking time, while Users B and C each account for 25%. If the volume of speaking is concentrated on specific users in this way, it provides insights into the balance of the meeting or the level of participation.

[0105] In addition, the frequency of speaker switching measures the number of times the speaker changes during a conversation and serves as an indicator of how active interaction is among team members. For example, infrequent speaker switching in a meeting can be interpreted as one person doing most of the talking or the meeting being primarily one-sided. Conversely, frequent switching suggests that it was a discussion-centered meeting where diverse opinions were exchanged.

[0106] In addition, the silence interval ratio refers to the proportion of time during a conversation when no one speaks, and it is used to analyze the flow or level of engagement of a meeting. For example, if 100 seconds of silence occur during the entire meeting, this accounts for approximately 8% of the total time and is generally considered to be a highly focused meeting. However, if the proportion of silence is excessively high, one should also consider the possibility that inefficient factors, such as connection delays, a lack of organized thoughts, or a decline in interest, may have been involved.

[0107] These metrics are automatically calculated from voice data based on speech time, speaker identification, and silence detection, and can be visualized on a dashboard to provide intuitive feedback to the user.

[0108] In other words, these indicators are visualized in the Dash Mode module as bar graphs, line charts, or heatmaps, allowing users to intuitively identify participants with excessive speaking volume, periods of frequent discussion transitions, and periods of silence where concentration is low. Additionally, user convenience is further enhanced by adding interactive link features that allow users to check specific figures in the visualized data or go directly to the transcribed text of a specific time period when clicking on a specific section.

[0109] In addition, the log analysis and dashboard can include specific metrics and UI elements such as question / response ratios, pauses, sentiment change trends, response speed distributions, and participation random access panels.

[0110] For example, the question / response ratio, which represents the ratio of response utterances to the number of utterances detected and indicates how effectively participants communicated their questions, is displayed as a pie chart; when an item in the chart is selected, the IDs of the questioner and respondent are displayed in a popup.

[0111] Speech interruptions can be evaluated by counting the number of overlapping segments of speech (whether it is a discussion of issues or a one-sided briefing, etc.), visualizing the density over time with a heatmap, and expressing the severity with a color spectrum (green*red).

[0112] The trend of emotional changes can be visualized by plotting the results of the real-time emotion analysis module (positive / neutral / negative) on a timeline chart, allowing for the visual highlighting of the point where the meeting atmosphere shifted from positive to negative (e.g., changing the background color).

[0113] The response speed distribution displays the distribution of the average time taken from question utterance to response as a box plot, allowing for a quick visualization of response speed disparities within the team.

[0114] The participation random access panel places a "Listen to Random Speech" button on one side of the dashboard, allowing users to randomly play analyzed sample speeches by speaker, thereby vividly conveying key context in addition to the summary.

[0115] By utilizing these metrics and UI elements, it is possible to go beyond simple numerical presentation and support qualitative analysis of meetings, calls, and conversations, as well as user immersion, with greater detail.

[0116] In addition, user behavior and engagement patterns can be understood in a more multidimensional way through the following additional indicators and analysis items.

[0117] For example, through an analysis of variance on speech duration, if a specific user's speech is frequently interrupted for less than 10 seconds, it is marked as 'passive participation,' and if it lasts for a long time of 1 minute or more, it is marked as 'active participation,' and is visualized as a histogram. Outliers (such as speeches lasting 5 minutes or more) are distinguished by a separate notification.

[0118] Additionally, Conversation Turns measures how often speakers are exchanged per unit of time and demonstrates the interactivity of the conversation. If this value is low, it is considered one-way communication, and a notification indicating a need for improvement may be provided.

[0119] In addition, mutual mention frequency counts the number of times speakers mention each other's names or positions, and is used as an indicator of cross-reference and attention within the conversation. For example, if the phrase "As Manager Kim said" is detected, it is considered a two-way interaction.

[0120] In addition, silence detection identifies time slots without sound for a certain period, notifying users of meeting delays or stagnation. At this time, a time snapshot is displayed on the UI, allowing users to quickly jump to that point.

[0121] In addition, the density of key keyword mentions allows for the visual reconstruction of the main flow of the conversation by displaying a heatmap showing how often and at what points key terms related to the meeting topic (e.g., 'deadline', 'ROI', 'approval') were mentioned intensively.

[0122] Changes in nonverbal vocal characteristics—such as pitch, speech rate (wpm), and volume—are tracked in real-time to analyze emotional changes like stress, tension, and excitement. Example: If speech speed increases by more than 15% compared to usual and persists, a 'potential tension' alert is displayed.

[0123] In addition, the last 5-minute concentration report can detect 'concentration peaks' or 'relaxation periods' by synthesizing voice frequency, speaker replacement rate, keyword density, etc., in the period immediately before the end of the conversation, and evaluate the efficiency of drawing conclusions in the meeting.

[0124] These quantitative and qualitative analysis functions can be utilized as foundational data for post-conversation reviews and improving meeting culture, and can be effectively used to quantify the communication quality of companies and organizations.

[0125] In addition, a voice communication support device is disclosed, characterized by including an authentication module and an encryption module for the secure management and access control of stored data, and being configured to perform data decryption and retrieval according to user authority.

[0126] Accordingly, schedule control, tag management, log analysis, dashboards, authentication / encryption modules, etc., are separated into independent services by function and designed to be scalable through a microservices architecture.

[0127] Further explanation is provided regarding this. The following content may be performed by a processor, and the order of the work may be any one of (S1) to (S6) or between them, or may be added to other items.

[0128] For example, the voice communication environment refers to various situations such as meetings, calls, everyday conversations, and self-talk. By analyzing the focus of the application in use, network metadata (e.g., whether the Zoom process is running, call app packet headers), and microphone input volume patterns, it is possible to classify 'meetings', 'calls', 'everyday conversations', 'self-talk', etc. In this case, by automatically assigning context tags through a machine learning classifier that combines acoustic characteristics and device status, the process can be optimized without separate user settings (context awareness technology).

[0129] For example, it can also receive biosignals such as heart rate, movement patterns, and skin temperature from wearable devices like smartwatches to determine whether there is a conversation in a stressful situation or while moving, and this signal-based situational awareness helps to automatically adjust recording sensitivity or filtering strategies without the user having to speak.

[0130] For example, if a call starts with a sudden increase in heart rate, it can be automatically classified as a 'consultation' or 'customer complaint' situation, and the sentiment analysis module can be set to run first.

[0131] For example, if the microphone input is activated while Bluetooth earphones are connected to the user's smartphone, and background packet activity of the call app is detected simultaneously, the system can automatically classify the situation as 'in a call'.

[0132] In addition, by utilizing the smartphone's proximity and ambient light sensors, it is also possible to detect situations where voice input is received while the device is in a pocket or bag, such as by classifying it as an "unintended capture" and disabling recording. Combining physical sensor information with user behavior patterns in this way enables more precise context-based situational judgment, which is useful for reducing usability errors in both outdoor and indoor environments.

[0133] Furthermore, by analyzing users' location-based information and recurring communication patterns, situational awareness based on social context is also possible. For instance, if a user repeatedly converses with the same person at similar times in a specific location (e.g., a cafe or library), the system recognizes this as a 'regular informal conversation' and can automatically switch to providing keyword-based summaries or notifications instead of a full summary.

[0134] On the other hand, if a schedule management app and calendar event are scheduled within 5 minutes and the Zoom or Teams app is fully active, it is considered a 'meeting situation'. In this way, situation awareness is performed by integrating device status, app focus, network header analysis, etc.

[0135] Furthermore, to perform situational awareness more precisely, a method combining the following multi-layered indicators is useful.

[0136] For example, through background app usage logs and voice tone analysis, if a user continuously switches between Slack, Figma, PPT, Excel, etc., during a meeting, it is classified as a 'work-oriented meeting,' whereas if a joking intonation is frequently detected in the conversation, it is classified as an 'informal conversation.' This allows for the subdivision of meeting purposes even within the same speaker composition.

[0137] Alternatively, as a geofencing-based location tag, the user's location is identified based on GPS or Wi-Fi SSID data, and location information such as 'office', 'external meeting room', or 'cafe' is used as context-aware input. Even with the same speech pattern, recording sensitivity or security levels can be automatically adjusted depending on the location.

[0138] In addition, as an analysis of calendar event titles and participants, if the event title contains details related to the event through integration with programs such as Google Calendar or Outlook, or with a schedule management program, the situation is pre-classified by the corresponding keyword, and if the number of participants is 2 or fewer, it is automatically tagged as '1:1 communication', and if there are 3 or more, it is automatically tagged as 'group meeting'.

[0139] In addition, network type-based recognition detects whether Wi-Fi is connected, a VPN is used, or an LTE / 5G standalone connection is established, classifying the environment as a mobile meeting or a remote access environment requiring security. For example, it automatically strengthens sensitivity settings when connecting to public Wi-Fi.

[0140] In addition, speech pattern-based implicit tagging can subdivide the conversation into 'presentation,' 'summary,' and 'decision-making' stages and apply different STT and summarization algorithms to each when specific communication patterns (e.g., "What I am going to tell you now is," "to summarize," "I have a suggestion") are detected in succession.

[0141] In addition, utilizing the device's accelerometer and ambient light sensors, a situation where the smartphone remains continuously stationary and the surrounding brightness is constant is considered a 'static meeting environment,' while a situation where the user moves or lighting changes rapidly is classified as a 'conversation while on the move,' automatically enhancing sound quality correction and noise filtering functions.

[0142] In addition, as a context-based prediction for continuous conversations, if a conversation between the same users exists within the last hour and topic keywords overlap, it is recognized as a 'continuous conversation,' and the reference point for summarization and classification is maintained based on previous conversation records. This method enables the effective distinction of continuous meetings or subsequent calls without time interruptions.

[0143] Such multidimensional situational awareness techniques go beyond simple process detection or volume-aware methods to evolve into cognitive situational classification based on user intent and context, which can further enhance the responsiveness and accuracy of the system.

[0144] In addition, STT priority, summarization strategy, and security settings are dynamically changed based on recognized tags. For example, functional branching control is possible by lightweighting the summarization module for local processing and offline queuing in 'monologue' situations, executing the entire flow via server integration in 'meeting' situations, and activating the sentiment analysis module in 'call' situations (process branching by situation).

[0145] For example, the sentiment analysis module can be configured to always be activated when a customer-interaction-focused 'consultation' situation is detected, and to automatically generate a CRM ticket if a negative tone or complaint sentence is detected. On the other hand, for situations classified as 'monologue,' server uploads are withheld, and the workflow is changed so that the user manually reviews and approves the upload.

[0146] For example, if the battery level falls below a certain threshold or storage space is insufficient, the system automatically switches to a lightweight mode that increases the compression ratio or omits summarization to perform only keyword extraction, without compromising STT quality. This setting provides a stable user environment by flexibly adjusting the process based on the device status without requiring separate user intervention.

[0147] In addition, deep learning-based sound source separation models (Demucs, Conv-TasNet) are used to separate voice and ambient noise from the input stream. Specific noise patterns, such as construction sites or traffic noise, are detected through real-time spectrogram analysis, and only the voice signal can be extracted by combining an Adaptive Noise Cancellation (ANC) algorithm that suppresses the corresponding frequency band (voice-noise separation).

[0148] In particular, in environments with significant street noise, a noise suppression strategy is applied that automatically recognizes ambient noise as 'masking noise' and preserves only the user's speech. On the other hand, in classroom or seminar environments where white noise levels are constant, only the AGC (Automatic Gain Control) function is activated to evenly amplify the entire voice signal.

[0149] In addition, spatial speaker separation is performed using a multi-channel microphone array and beamforming technology, and after distinguishing each speaker using a Speaker Diarization algorithm (UEM-based clustering, x-vectors), conversational roles (presenter, responder, etc.) can be automatically tagged for each speaker through speaker profiling (multiple speaker separation and identification).

[0150] For example, the user who speaks the most and interacts with slides within the first five minutes of the meeting is automatically tagged as the 'main presenter,' while the user who intermittently repeats short sentences is classified as the 'questioner.' In this way, role-based tags can be directly utilized in the subsequent summarization stage to distinguish highlights by speaker or to structure meeting minutes.

[0151] In addition, it supports various resolutions ranging from sentence units to word units and applies dynamic segmentation strategies based on conversation speed or utterance length. For example, if the static time between utterances exceeds a certain threshold, a session separation timestamp is recorded, and short events (questions, responses) are timestamped on a word basis (timestamp criteria setting).

[0152] In addition, for example, in the extractive summarization stage, key sentences such as "10% increase in sales target for this quarter" are selected, and in the generative summarization stage, this is naturally restructured into "The team set the sales target for this quarter 10% higher than the previous year." In the present invention, these two stages are combined into a pipeline to simultaneously ensure readability and information completeness (hybrid summarization example).

[0153] For example, if there is a statement during a meeting such as "The target for this quarter is 12 billion won," the extraction summary includes the sentence as is, and the generation summary expresses it as "The sales target for this quarter has been set at 12 billion won" to naturally reflect it in the meeting minutes.

[0154] Keyword weights are calculated using TF-IDF, but their importance is adjusted by multiplying by weight correction values ​​based on conversational domain dictionaries (e.g., finance, medical). Semantic analysis utilizes cosine similarity between sentence embeddings (e.g., Sentence-BERT), considering statements with high similarity to topic-centered sentences as key. The threshold can be dynamically adjusted using user recommendation techniques (Top-K) or variable alpha thresholds (key statement identification method).

[0155] In addition, the uploaded document is structured into metadata for each element (slide number, paragraph ID) after OCR processing, and threshold-based matching candidates are selected when mapping similarity between each sentence of the transcribed text and metadata items. Finally, items with low matching scores are marked as 'missing items,' and a feedback loop is provided on the UI that allows users to directly check and modify them (reference material verification procedure).

[0156] For example, if the meeting materials included a slide titled "10% reduction in marketing budget" but the content was omitted from the meeting remarks, the system detects this, marks it as "missing compared to reference materials," and shows the user a suggestion for correction.

[0157] In addition to basic voice capture and analysis functions, the present invention can optionally combine various extension modules to enhance the user experience.

[0158] For example, by applying a real-time sentiment analysis module, the speaker's emotional state (positive, negative, stress, etc.) can be visualized in real time by analyzing voice tone, speed, and frequency fluctuations. It can also automatically detect sensitive remarks or signs of conflict and provide immediate notifications to administrators or users. This enables a rapid response to emergency situations that may occur during meetings or calls.

[0159] The results of the sentiment analysis are visualized on a time-axis graph in the form of 'stress index' or 'positive / negative index.' For instance, timeframes where a 'high tone' and 'rapid speech rate' are detected simultaneously can be highlighted in red as 'tension,' while sections where a 'slow tone' follows can be visualized in blue as 'depression or exhaustion.' This visualization helps managers or participants intuitively grasp the flow of emotions after a meeting.

[0160] In addition, the automatic action item extraction feature allows for the automatic identification and listing of 'To-Do' tasks and related responsibilities from summarized text. Utilizing Named Entity Recognition (NER) technology, tasks and responsible parties are distinguished, and the generated action items are integrated with calendars or task management systems (e.g., Jira, Asana) to be automatically registered, thereby supporting immediate transition to subsequent tasks.

[0161] For example, in the sentence "Assistant Manager Kim, please finish the report by next week," writing the report is automatically identified as an action item and Assistant Manager Kim as the responsible party; this information can be linked with the task management system and registered as an automatically generated To-Do item.

[0162] In a global collaboration environment, real-time multilingual translation and subtitle modules are highly effective. These modules translate captured audio into the target language and display it in real-time as subtitles, facilitating seamless communication by overcoming language barriers between multinational teams.

[0163] For example, if the speaker says in Korean, "The next meeting is on July 15 at 10 am," the caption "The next meeting is on July 15th at 10 am" is displayed at the bottom. This caption is recorded as a parallel table within the saved PDF file of the meeting minutes, and the user can click on the corresponding line of the caption to go directly to the time of the speech.

[0164] In the event of an unstable network environment, a lightweight model is provided that processes essential functions (STT, summarization, and key statement identification) on the local device through enhanced offline mode. Upon network recovery, data queued locally is automatically retransmitted to and merged with the server to maintain data continuity.

[0165] In addition, the system according to the present invention preserves original voice data (original sound) without loss, stores step-by-step transcription data such as rough transcription (machine recognition result) and paraphrased transcription (natural language refinement result) separately, and is configured to allow parallel verification or selective viewing through a user interface (UI).

[0166] For example, users can select a specific speech segment to directly play the original audio and compare and verify the initial transcription with the paraphrased translation. If an error is detected during this process, an interface is provided to request re-STT or re-summarization for the relevant segment, and the corrected result is managed as a separate version along with traceable logs. This original audio-centric multiple transcription result management function facilitates the verification and post-correction of transcription quality and provides a highly reliable foundation for speech analysis.

[0167] For example, if the connection is lost in mountainous areas or building basements, the lightweight STT and summarization engine within the device operates first, and the transcribed file is saved in the Opus compression format and placed in a background queue. Once the network is restored, automatic uploading begins based on the priority queue, and the file is automatically merged with the existing version on the server.

[0168] Furthermore, through knowledge-based integration and search functions linked to internal corporate wikis or databases, users can receive real-time recommendations for past meeting minutes or project documents. By instantly connecting keywords mentioned during conversations to relevant documents and providing them via pop-ups or links, users can be provided with immediate background knowledge.

[0169] In addition, as a privacy enhancement option, sensitive information (personally identifiable information, contract clauses, etc.) can be automatically blurred or voice modulation functions applied via keyword detection, and audit logs and anonymization functions for compliance with personal information protection regulations such as GDPR and CCPA can be added to simultaneously achieve the protection of sensitive information and legal compliance.

[0170] Furthermore, by integrating a user feedback loop, users can directly correct or supplement transcription errors or omissions within the app, and these corrections are reflected in the internal training database to continuously improve the performance of the STT and summarization models. For example, if frequently used technical terms or words with unclear pronunciation are incorrectly transcribed, users can directly correct them and register them in the dictionary, enabling automatic correction during the next similar utterance recognition.

[0171] Furthermore, in addition to the basic extension options presented above, the following additional features can be described in sentence form for more detailed usability improvements and environmental adaptation.

[0172] First, the environment-specific audio calibration function automatically adjusts microphone gain and EQ settings to match the characteristics of the recording location (indoors, outdoors, inside a vehicle, outdoor event venues, etc.). For example, for recording inside a vehicle, the gain and frequency response are optimized to suppress low-frequency engine noise and emphasize the human voice band.

[0173] Second, by applying user feedback loop integration, whenever users directly correct or provide feedback on enterprise errors or missing items within the app, those corrections are reflected in the model training data, enabling continuous improvement of enterprise and summary accuracy in the future.

[0174] For example, if the word 'shinjang' (height) is incorrectly transcribed as 'market' (market), and the user directly corrects the error within the app's UI, the word is automatically registered in the user dictionary and reflected in the STT model so that it can be correctly recognized for similar voice inputs in the future.

[0175] For example, if an incorrectly recognized technical term is stored in the user dictionary, it produces an improved result when the same term is recognized later.

[0176] Third, by introducing a privacy zone feature sensitive to specific locations or situations, recording is automatically paused based on user-designated areas (e.g., specific areas in a meeting room) or device proximity. This enables the simultaneous prevention of recording in sensitive areas and the protection of personal privacy. Additionally, when a user utilizes multiple voice-related apps (e.g., Zoom, KakaoTalk, phone, memo apps, etc.) concurrently, priority is automatically set based on the app currently in active focus and capture quality (microphone sensitivity, number of channels, etc.), and this priority can be changed in real-time based on context-sensing tags. For example, when Zoom switches to screen sharing mode, the audio from that app is prioritized, while the remaining apps switch to background capture mode.

[0177] Additionally, the present invention may include a function to automatically detect and recover from abnormal application behavior (e.g., app crash, frozen state, etc.) by collecting app status and device environment information in real time during recording. For example, if the app response is delayed by more than 3 seconds while voice capture is running, the state is detected as 'temporary suspension,' and the capture is automatically switched to background mode, or a recovery log is saved and a recovery notification is provided to the user to prevent data loss. Such a function is effective in mobile environments where device performance and app stability are critical.

[0178] For example, if a user leaves their smartphone at a specific location to attend a meeting, and a privacy zone named 'Meeting Room A' is designated, voice capture is automatically disabled within that area. This feature can prevent automatic recording in sensitive locations, such as private areas within meeting rooms, hospital examination rooms, and customer service consultation spaces.

[0179] Fourth, by periodically utilizing user-uploaded reference materials or conversation history as retraining data through model-customized updates, it is possible to maintain and manage STT, summarization, and identification models optimized for terms and styles specific to domains such as medical, legal, and engineering.

[0180] Fifth, through the performance monitoring dashboard, it visualizes the failure rate, processing delay, and resource usage (CPU, memory) of all modules from capture to synchronization in real time, and ensures system stability and availability by providing immediate notifications to operators in the event of abnormal signs.

[0181] Sixth, through security policy integration, it is possible to link with Data Loss Prevention (DLP) or Digital Rights Management (DRM) solutions to perform dynamic watermark insertion into recording files and summary results, file integrity checks, etc., thereby providing data leakage prevention and audit trail functions that comply with corporate security policies.

[0182] These sentence-based descriptions of additional functions emphasize the flexibility and scalability of the present invention and can be utilized as customized application and implementation guides tailored to actual usage environments.

[0184] In addition, regarding the user interface, the device may be implemented as follows according to the embodiment of the present invention. That is, the device according to the present invention may include a display unit or a display unit, wherein a user interface (though not previously specified) is provided so as to be operated by a user or output values, etc., by a processor are displayed.

[0185] For example, the 'Automatic Context Detection' option on the user interface (UI) is accessible from the menu panel at the top of the settings screen or conversation screen. When the user activates it, the system automatically analyzes and displays situations such as meetings, calls, self-talk, and everyday conversations in real time in the background. Detected situations are displayed in the form of real-time badges, visualized by adding intuitive colors and icons to text describing the situation, such as 'In a meeting,' 'On a call,' 'Self-talk,' or 'Conversation while on the go.'

[0186] In addition, the user interface is designed to allow integrated monitoring and control of context awareness, summary settings, speech highlighting, verification requests, synchronization status, and more.

[0187] For example, the currently detected situation is displayed at the top of the UI as a tag, such as 'meeting (meeting room, 4 people)', 'phone call (subway)', 'everyday conversation (cafe)', and the user can intuitively adjust the situation-specific voice capture sensitivity, enterprise accuracy priority, network usage policy, etc. by touching the tag.

[0188] In addition, summary strategies, STT model types, and battery saving modes are preset based on tags. When conditions change in real time (e.g., quiet office * noisy street), switching suggestions are provided to the user via an automatic pop-up, and settings can be applied with a single touch.

[0189] Through this, users can intuitively grasp the system's internal state and situational awareness, and obtain a UI-based control experience that allows for active response.

[0190] These context tags are stored together in the conversation history and linked with post-summary, search, and notification priority settings, and users can configure customized processes for each context tag on the UI. For example, it is possible to set STT accuracy to be the top priority in 'Meeting' situations and to perform only local saving in 'Self-talk' situations.

[0191] Additionally, if a user determines that a detected situation is incorrect, they can click the 'Request Correction' button in the UI to directly select or suggest a situation; this feedback is utilized as training data for model improvement. Users can also completely disable the automatic situation detection feature or selectively exclude specific situations (e.g., excluding self-talk detection). Furthermore, notification settings within the UI allow for the automatic execution of push notifications or summary generation requests upon the detection of specific situations.

[0192] This 'automatic situation detection' UI configuration is designed to ensure both reliability and efficiency in a voice-based communication environment by allowing users to actively check and control the system's recognition results.

[0193] In particular, the present invention is designed to facilitate user review and comparison by visualizing recorded text and GPT-based summary results in parallel on a real-time UI. For example, the system is configured to allow viewing of STT results and summaries on the same screen, enabling users to immediately check the correspondence between the original text and the summary for each speech segment. If necessary, users can request to re-summarize only the relevant segment or modify the text and reflect the changes. Such an interface strengthens user-centric feedback loops and contributes to increasing the precision of meeting minutes and record automation.

[0194] For example, the real-time text stream viewer is integrated with the STT module, enabling users to view transcription results, such as subtitles generated during a conversation, in a scrolling format in real time. In addition, through collaboration with the key speech identification module, important speech sentences are highlighted in bold or with colored backgrounds to make them highly visible, allowing users to instantly identify key points at a glance.

[0195] The summary dashboard panel is positioned on the side of the screen or in a separate tab, providing extractive and generative hybrid summary results in a card format. Key keywords are displayed as tags at the top of each card, and clicking the card title hyperlinks the user to the corresponding entry location, allowing for quick navigation and significantly improving readability and navigation convenience.

[0196] The timeline-based interactive player is positioned below the audio waveform along with timestamps, distinguishing important sentence segments detected by the speech identification module with color strips (e.g., red). When a user clicks a color strip, a summary of the corresponding segment or the original audio is played immediately, allowing for quick retrieval of desired information.

[0197] The missing item notification is integrated with the verification module to display a warning as a banner at the top of the app. When a user clicks, a list of missing items appears as a pop-up overlay. Within this pop-up, users can directly check each item or provide immediate feedback by pressing 'Confirm' or 'Delete' buttons, thereby simplifying the reference material verification process in a user-centric manner.

[0198] The Tag and Action Item widget is always displayed in the right sidebar, showing a field for users to enter tags along with an automatically generated To-Do list. Next to each list item are checkboxes and edit icons, allowing for convenient management of follow-up actions, such as changing completion status or assigning responsibility.

[0199] The Settings and Feedback screen includes options such as audio calibration, privacy zone settings, and user dictionary management, while the Feedback Loop section visualizes recent correction history and accuracy change graphs. A 'Correct' button is placed next to incorrectly recognized sentences, providing an intuitive interface that allows users to make immediate corrections.

[0200] The security access control UI integrates with the authentication / encryption module to guide users through the two-factor authentication (FIDO2 / WebAuthn) process via a popup layer when accessing sensitive data. For users without authorization, the corresponding UI element is grayed out and disabled, clearly indicating visually that access is denied.

[0201] This integrated UI approach is designed to maximize the user experience and simultaneously ensure intuitiveness and visibility by organically connecting various functions of the present invention on customer terminals.

[0202] Although various preferred embodiments of the present invention have been described above with some examples, the descriptions of various embodiments described in the "Specific details for carrying out the invention" section are merely illustrative, and those skilled in the art to which the present invention pertains will understand that the present invention can be modified in various ways or equivalent embodiments can be carried out based on the above description.

[0203] In addition, since the present invention can be implemented in various other forms, the present invention is not limited by the description above. The above description is provided merely to make the disclosure of the present invention complete and to fully inform those skilled in the art of the scope of the present invention, and it should be understood that the present invention is defined only by each claim of the claims.

Claims

Claim 1 A voice communication support method performed by a processor, comprising: (S1) classifying situations based on the application focus and microphone input volume patterns in a voice communication environment including a conference room, office call, smartphone call, voice chat app, daily conversation, or monologue; optimizing without separate user settings by automatically assigning situation tags by combining acoustic characteristics and device status through a machine learning classifier; utilizing proximity sensors and ambient light sensors to consider voice input when the device is in a pocket or bag as an unintended capture and disable recording; automatically capturing and recording voice signals generated in the voice communication environment in the form of raw audio streams using a microphone input interface, supporting single-channel or multi-channel input, analyzing input strength for each channel in real time to apply directional microphone functions or beamforming; performing automatic gain adjustment to dynamically adjust microphone gain by monitoring ambient noise levels in real time; applying an FDN filter to suppress echoes where speaker playback sound is fed back to the microphone input interface; automatically executing in the background to record without loss of voice signals; and performing schedule control linked with an external schedule management service to automatically schedule and start capture when a conversation starts and automatically stop when it ends. Step; (S2) removing noise from the recorded voice data using spectral gating and a Wiener filter, automatically switching to the corresponding language model when a foreign language is spoken, and supporting multi-language speech recognition by reapplying the model for each language based on timestamps and language tags during post-analysis, learning feature vectors for each speaker when there are multiple speakers, separating real-time recorded voice data by segment and linking it to timestamps, and receiving the number of conversation participants from the user in advance to optimize initial conditions for speaker separation and separate speakers.A step comprising: automatically performing speech-to-text (STT) in real-time or afterward; generating or storing transcribed text containing sentence or word-based timestamps; correcting the accuracy of technical term recognition by reflecting the purpose of the conversation, topic, keywords, and a database that can be injected into or retrieved from an external source; identifying homonyms and supplementing missing subjects and objects by reflecting the context of the conversation; and automatically filtering or masking if slang, prohibited words, or sensitive words are included; (S3) generating a conversation summary by performing a hybrid summary combining extractive and generative summaries based on the generated text; automatically recognizing the purpose of the conversation, topic, and keywords to control the scope and detail of the summary, wherein tags are input from the user's touch or voice command and the priority of the summary and notifications is adjusted based on the input tags; reflecting remarks marked as important by the user in the priority summary; setting the summary perspective according to the purpose of the conversation; extracting keywords; and filtering topics according to the extracted keywords; recognizing the development structure of the conversation according to an alignment algorithm to arrange the generated conversation summary in a logical order; and, according to the summary style according to the user's settings A step including a quality verification routine that provides a customized summary, evaluates quality by applying quantitative indicators including BERTScore or ROUGE to the generated summary, and induces feedback from the user if the quality is below a preset level, and stores metadata including the summary generation time, the type of model used, and the length of the transcribed text for each summary result for quality evaluation and history management, and transmits the transcribed text to an external summary model to receive and store the summary result, wherein if the summary result is modified or new information is reflected, the external model is called again to synchronize the updated summary result; (S4) a step including selecting key statements from the generated summary by combining a keyword weighting calculation technique, an embedding-based similarity analysis, and a user filtering function, andThe above keyword weight calculation technique calculates by multiplying the TF-IDF value by a weight correction value for important keywords pre-specified in a domain dictionary; the above embedding-based similarity analysis embeds each sentence using Sentence-BERT and corrects importance by calculating the similarity with the representative sentence of the entire conversation using cosine similarity; the above user-defined filtering automatically detects prohibited words, remarks of interest, emphasized words, and remarks regarding internal security and policy compliance, and allows for their inclusion or exclusion; and includes a user-centric feedback loop that learns sentences marked as important by the user or repeatedly confirmed to assign weights to similar remarks in the future; (S5) collects and analyzes user behavior logs before, during, and after the conversation to generate engagement and concentration indicators and visualizes the generated concentration indicators on a dashboard; the above concentration indicators include the ratio of speech volume, the frequency of speech switching, and the ratio of silence intervals; the above ratio of speech volume refers to the ratio of speaking time per user or relative to the total; the above frequency of speech switching measures the number of times the speaker is replaced during the conversation; and the above ratio of silence intervals is when the voice during the conversation A step including measuring the pause time ratio, wherein the concentration indicators are automatically calculated based on speech time, speaker identification, and silence detection from voice data, and moving to the transcribed text of the corresponding time period when a specific voice segment is clicked, and providing and visualizing the question-and-answer ratio, speech pauses, trends in emotional changes, response speed distribution, number of switching speakers, frequency of mutual mentions, density of core keyword mentions, changes in non-verbal voice characteristics, and the last 5 minutes concentration report; (S6) when pre-uploaded reference materials are verified, performing OCR and DOM parsing, calculating the cosine similarity between keywords in the reference materials and the transcribed text, and identifying and listing missing items as missing if they are below a preset threshold, and allowing the user to check the missing items,A method for supporting context-aware voice communication, comprising the steps of: (S7) synchronizing and storing summary, key remarks, and missing items on a server; including retry logic after local cache in case of response delay and offline mode; performing version hash or timestamp version management to prevent duplicate uploads or update conflicts of identical summary data; managing the expiration time of data in the queue to minimize unnecessary transmissions or selecting transmission targets according to data priority when the offline state continues; and applying two-factor authentication, and performing decryption and retrieval of data according to user authority through authentication and encryption of stored data. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete